Vehicle driving safety warning method and system

Optical flow and depth prediction are combined with FastFlowNet and MegaDepth neural networks, and the real-time and cost problems of vehicle driving safety warning systems are solved, real-time and accurate speed prediction and safety warning are achieved, independent of vehicle equipment and third-party authority.

CN114898328BActive Publication Date: 2025-07-11SHANGHAI JIAOTONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210298737.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-23
Publication Date
2025-07-11
Estimated Expiration
2042-03-23

AI Technical Summary

Technical Problem

In the prior art, vehicle driving safety warning systems rely on high-precision sensors and are costly, and the data exchange application scenarios at the remote control end are limited, so real-time warning and driving assistance cannot be achieved.

Method used

The image is obtained by using the on-board camera, and optical flow and depth prediction are carried out through the FastFlowNet and MegaDepth neural network models. Combined with feature fusion, real-time speed prediction is output and warning is issued in dangerous driving mode. It is uploaded to the cloud platform independently of the vehicle equipment.

Benefits of technology

It realizes real-time and accurate speed prediction and safety warning during vehicle operation, reduces costs, has third-party authority, fills the gap in market application, and provides drivers with real-time safety reminders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114898328B_ABST
    Figure CN114898328B_ABST
Patent Text Reader

Abstract

The present invention provides a vehicle driving safety warning method and system, including: Step S1: Obtain two adjacent frames of images as model inputs, stretch the images to conform to the camera parameters; Step S2: Use the modified FastFlowNet optical flow prediction neural network model, take two adjacent frames of images as inputs, and output an optical flow vector matrix; Step S3: Use the MegaDepth neural network model, take two adjacent frames of images as inputs, and output a depth matrix as the input of the feature fusion module in subsequent steps; Step S4: Perform feature fusion on the results of Step S2 and Step S3; Step S5: For the result of Step S4, take three sub-matrices of the same size; Step S6: For each pair of adjacent captured images obtained, output the predicted speed; Step S7: Output a warning message according to the prediction result of Step S6. The present invention can provide an analysis of vehicle data, road condition data, or obstacle types that is low-cost and independent of vehicle sensors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the cross - technical field of video processing and artificial intelligence. Specifically, it relates to a vehicle driving safety warning technology based on a fusion optical flow deep convolutional neural network, and particularly to a vehicle driving safety warning method and system. Background Art

[0002] With the rapid development of intelligent vehicles, artificial intelligence technology has broad application scenarios for improving people's safe driving coefficient. The driving recorders installed in vehicles only have the functions of recording videos and images, but lack intelligent analysis and application of these data. Generally, in the fields of road condition, vehicle speed, vehicle trajectory detection, etc., other vehicle - mounted sensors (such as acoustic wave, infrared sensors) are generally used to analyze the surrounding environment, obstacles, etc. To meet the usage requirements, the accuracy requirements of these sensors are very high, thus increasing the usage cost.

[0003] Since the convolutional neural network (CNN) was proposed, it has achieved the best results in related fields such as image recognition, segmentation, detection, and retrieval. Different from ordinary neural network models, the convolutional neural network extracts features by performing convolution operations on data, reducing the data dimension, and making the speed of model training and application much faster than that of general neural networks. The deep convolutional neural network (Deep CNN) further improves the network depth, width, and input image resolution on the basis of CNN, and can be used for applications in more complex scenarios. The network designed in this paper magnifies the network from the above - mentioned three dimensions on the basis of the existing CNN architecture, achieving a higher efficiency than the current mainstream networks in this application scenario.

[0004] Regarding the problem of Visual Odometry based on optical flow method or convolutional neural network, there are already some patents. Typically, such as the invention patent with the publication number CN112419411A, which discloses an implementation method of a visual odometer based on convolutional neural network and optical flow features. Two adjacent frames in the image sequence are input into the optical flow feature extraction network based on PWC-net, and the optical flow feature map is extracted by the optical flow feature extraction network; the obtained optical flow feature map is further feature-extracted by the convolutional neural network, and the mapping relationship between the optical flow feature map and the ground truth image is established, so as to estimate the relative pose between adjacent frame images; the relative pose in step two is transformed into an absolute pose to restore the original motion trajectory. However, this invention has relatively high requirements for the resolution and focal length conditions of the camera based on prediction only by the optical flow method without adding depth information. At the same time, the goal of this method is still to predict based on the existing image sequence, and it fails to achieve the functions of real-time prediction and early warning.

[0005] In addition, in the field of vehicle early warning, there is an invention patent with the publication number CN113269962A, which discloses a vehicle early warning and control system based on computer vision, including a system applied to the target vehicle, the first server, and the second server. Through data interaction between systems, data collection of the target vehicle and adjacent vehicles is carried out in the system of the target vehicle, data collection and analysis of other vehicles are carried out in the system of the first server, and data analysis of the target vehicle, adjacent vehicles, and other vehicles is carried out in the system of the second server to obtain early warning and route control instructions for providing driving assistance to the driver of the target vehicle. This invention applies the computer vision method to the field of vehicle early warning, but the application scenario of this invention is the early warning of the vehicle by the remote control end, which requires a large amount of data exchange between the vehicle and the server, and the application scenario is still limited. Summary of the Invention

[0006] In view of the deficiencies in the prior art, the present invention provides a vehicle driving safety early warning method and system.

[0007] According to a vehicle driving safety early warning method and system provided by the present invention, the solution is as follows:

[0008] In a first aspect, a vehicle driving safety early warning method is provided, and the method includes:

[0009] Step S1: Obtain two adjacent frames of images through an in-vehicle camera or a driving recorder as model inputs, and stretch the images to conform to the camera parameters;

[0010] Step S2: Use the modified FastFlowNet optical flow prediction neural network model. Take two adjacent frames of images as input and output an optical flow vector matrix, where each vector indicates the change in the illumination condition of the pixel point, which serves as the input for the feature fusion module in the subsequent steps;

[0011] Step S3: Use the MegaDepth neural network model. Take two adjacent frames of images as input and output a depth matrix, which serves as the input for the feature fusion module in the subsequent steps. Under the condition allowed by the computing device, this step S3 can run synchronously with step S2;

[0012] Step S4: Perform feature fusion on the results of step S2 and step S3;

[0013] Step S5: For the result of step S4, take three sub-matrices of the same size and merge them as the input for the real-time accurate speed prediction neural network in the subsequent steps;

[0014] Step S6: Use the designed speed prediction neural network for operation. For each pair of adjacent captured images obtained, output the predicted speed;

[0015] Step S7: Output a warning message according to the prediction result of step S6. If it meets the dangerous driving mode, output a warning.

[0016] Preferably, the step S1 includes: obtaining the video stream input of the camera. For every two adjacent frames of images, directly use them as the input for the subsequent module, and use the interpolation method to stretch or compress the resolution of the source RGB image into a 1266×370 RGB image.

[0017] Preferably, the step S2 includes: taking the two frames of images output by step S1 as input, calculating the optical flow vector matrix, and the output result is a vector field representing the change in the illumination condition of each pixel of the image, which is stored as a 2×1266×370 matrix.

[0018] Preferably, the design of the FastFlowNet optical flow prediction neural network model in the step S2 is as follows:

[0019] Step S2.1: Head enhanced pooling pyramid: The FastFlowNet model uses the head enhanced pooling pyramid for feature extraction. It fuses the higher-layer feature pyramid in step S1 with the lower-layer pooling pyramid in step S2 to combine the advantages of both; at the same time, add a convolutional layer on the high-level pyramid to strengthen the pyramid features through computational cost;

[0020] Step S2.2: Center Dense Dilated Correlation (CDDC) layer: Different from the structure of traditional convolutional layers, the model uses a center dense dilated correlation layer to densely sample in the central part of the image, while downsampling grid points in large motion areas, effectively reducing the computational amount while maintaining the model's cognitive radius; when performing stereo matching, the cost volume function of the model is constructed as follows:

[0021]

[0022] where c l (x, d) represents the cost volume function; l represents the l-th layer of the feature pyramid; x represents the vector pointing from the lower left corner of the image to the pixel point; d represents the vector pointing from the center of the image to the pixel point; represents the feature function of the image; represents the feature function of the warped image; N represents the dimension of the feature as the model input; r represents the search radius of the convolution;

[0023] Step S2.3: Aggregation block decoder: The compact cost volume constructed by CDDC reduces the maximum feature channels of the decoder from 128 to 96; three 96-channel convolutions are changed to group convolutions, and each decoder network contains three aggregation blocks with a group number of 3.

[0024] Preferably, the step S3 includes: using the RGB image obtained in step S1 as the input, performing speed prediction through the MegaDepth network, for each pixel point of the image, if the pixel does not represent the sky, output its relative depth, and if the pixel represents the sky, remove its depth information.

[0025] Preferably, the step S4 includes: inputting the matrices of step S2 and step S3, that is, the optical flow vector matrix OF i,j and the relative depth matrix D i,j into a convolutional layer with a convolutional kernel size of 3×3 for three layers. For the obtained results OF′ and D′, further perform feature fusion calculation, and the expression of the whole process of feature fusion calculation is as follows:

[0026]

[0027] where i represents the abscissa of the pixel point; j represents the ordinate of the pixel point; F′ 0,i,j represents the dimension indicating vertical direction movement in the fusion feature matrix; F′ 1,i,j represents the dimension indicating horizontal direction movement in the fusion feature matrix; OF′ 1,i,j represents the modified optical flow feature matrix calculated through the convolutional layer; D′i,j represents a modified depth feature matrix calculated by a convolutional layer;

[0028] For the obtained result, it is input into a transposed convolutional layer with a convolutional kernel size of 3×3 for three layers to obtain a feature matrix F i,j , which is used as the input of the prediction model.

[0029] Preferably, the step S5 includes: taking the result of step S4, that is, a sub-matrix of the feature matrix F. Specifically: taking three 200×60 two-dimensional matrices. For these three sub-matrices, the coordinates corresponding to the (0,0) coordinate in the original matrix are (113,200), (513,200), and (913,200), and F is discarded 0,i,j , and taking the F 1,i,j part as the input of the subsequent speed prediction network.

[0030] Preferably, the step S6 includes: inputting the output image of the previous step into a pre-trained convolutional neural network for prediction; the input of the speed prediction module network is the result of step S5, that is, a 3×200×66 matrix; the speed prediction module network includes 1 layer of normalization layer, 3 layers of convolutional layers with a convolutional kernel size of 5×5 and a stride of 2, 3 layers of convolutional layers with a convolutional kernel size of 3×3, and 4 layers of fully connected layers.

[0031] Preferably, the step S7 includes: analyzing the speed information of this frame and the previous frames. If the acceleration is too fast or the attitude change amplitude is large, a warning is issued; specifically: inputting the speed represented by this frame image and the speeds of the previous three frames into a judgment module. If its speed pattern conforms to the dangerous driving pattern, an alarm message is output, and finally the predicted speed is synchronously uploaded to the server of the cloud third-party platform.

[0032] In a second aspect, a vehicle driving safety warning system is provided. The system includes:

[0033] Module M1: Obtains two adjacent frame images as the model input through an in-vehicle camera or a driving recorder, and stretches the images to conform to the camera parameters;

[0034] Module M2: Uses a modified FastFlowNet optical flow prediction neural network model, takes two adjacent frame images as the input, and outputs an optical flow vector matrix, where each vector indicates the change in the illumination condition of the pixel point, which is used as the input of the feature fusion module in the subsequent steps;

[0035] Module M3: Uses a MegaDepth neural network model, takes two adjacent frame images as the input, and outputs a depth matrix, which is used as the input of the feature fusion module in the subsequent steps. Under the condition that the computing device allows, this module M3 can run synchronously with step module M2;

[0036] Module M4: Perform feature fusion on the results of Module M2 and Module M3;

[0037] Module M5: Take three sub-matrices of the same size from the results of Module M4, and after merging, use them as the input of the real-time and accurate speed prediction neural network in the subsequent steps;

[0038] Module M6: Perform operations using the designed speed prediction neural network, and for each adjacent two frames of camera images obtained, output the predicted speed;

[0039] Module M7: Output a warning message according to the prediction result of Module M6. If it meets the dangerous driving mode, output a warning.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] 1. Through the fusion of optical flow depth neural network training, the present invention intelligently detects comprehensive information such as vehicles, obstacles, pedestrians, lanes, driving speed, etc. in the front and side directions on the in-vehicle video dataset from the first-person perspective, gives real-time reminders to the driver, and ensures the safe driving of the vehicle. It has great pioneering significance for the intelligentization of automobile driving and the industrialization of safe driving AI assistants;

[0042] 2. At the same time, aiming at the current insufficient utilization of the data obtained by the driving recorder, and some existing technologies are mostly offline on databases such as Kitti and require the collection of the whole process of vehicle movement before prediction can be carried out. The real-time and accurate third-party speed prediction algorithm implemented by the present invention can output real-time speed prediction results during the running of the vehicle without relying on the vehicle's own equipment, and synchronize with the third-party cloud platform;

[0043] 3. The purpose of the present invention is to propose an effective vehicle safety warning system method through these data. This method can fill the application gap in the current market, is independent of other in-vehicle devices or sensors, and has the advantages of third-party authority, low cost, and good effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, purposes, and advantages of the present invention will become more obvious:

[0045] Figure 1 It is a flowchart of the method of the present invention;

[0046] Figure 2 It is a schematic diagram of the implementation result of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all belong to the protection scope of the present invention.

[0048] An embodiment of the present invention provides a vehicle driving safety warning method. Referring to Figure 1 and Figure 2 as shown, the method specifically includes the following steps:

[0049] Step S1: Obtain two adjacent frames of images through an in-vehicle camera or a driving recorder as the input of the model, and stretch the images to conform to the camera parameters.

[0050] Specifically, in this step S1, the video stream input of the camera is obtained. For every two adjacent frames of images, without further saving, they are directly used as the input of the subsequent module. On this basis, the resolution of the source RGB image is stretched or compressed to a 1266×370 RGB image using an interpolation method.

[0051] Step S2: Use the modified FastFlowNet optical flow prediction neural network model, with two adjacent frames of images as the input, and output a (dense) optical flow vector matrix, where each vector indicates the change in the illumination condition of the pixel point, as the input of the feature fusion module in the subsequent steps.

[0052] Specifically, in this step S2, taking the two frames of images output in step S1 as the input, calculate the (dense) optical flow vector matrix. The output result is a vector field representing the change in the illumination condition of each pixel of the image, and is stored as a 2×1266×370 matrix. In particular, different from general optical flow prediction models, in order to improve the calculation speed of optical flow prediction, the design of the FastFlowNet optical flow prediction neural network model is as follows:

[0053] Step S2.1: Head-enhanced pooling pyramid: The FastFlowNet model uses a head-enhanced pooling pyramid for feature extraction. It fuses the higher-level feature pyramid in step S1 with the lower-level pooling pyramid in step S2 to combine the advantages of both; at the same time, a convolutional layer is added to the high-level pyramid to strengthen the pyramid features with a relatively small additional computational cost;

[0054] Step S2.2: Center Dense Dilated Correlation (CDDC): Different from the traditional convolutional layer structure, the model adopts the center dense dilated convolutional layer, which densely samples in the central part of the image and downsamples the grid points in the large motion area, effectively reducing the computational complexity while maintaining the model's cognitive radius; when performing stereo matching, the cost volume function of the model is constructed as follows:

[0055]

[0056] where c l (x, d) represents the cost volume function; l represents the l-th layer of the feature pyramid; x represents the vector pointing from the lower left corner of the image to the pixel point; d represents the vector pointing from the center of the image to the pixel point; represents the feature function of the image; represents the feature function of the warped image; N represents the dimension of the feature as the model input; r represents the search radius of the convolution.

[0057] Step S2.3: Aggregation block decoder: The compact cost volume constructed by CDDC reduces the maximum feature channels of the decoder from 128 to 96 without affecting the model accuracy. Further, three 96-channel convolutions are changed to group convolutions, and each decoder network contains three aggregation blocks with a group number of 3, which can effectively reduce the model computational complexity within the tolerable accuracy degradation range.

[0058] Step S3: Use the MegaDepth neural network model, with two adjacent frames of images as input, and output the (relative) depth matrix as the input of the feature fusion module in the subsequent steps. In particular, under the condition allowed by the computing device, this step S3 can run synchronously with step S2.

[0059] Specifically in this step S3: Using the RGB image obtained in step S1 as the input, perform speed prediction through the MegaDepth network. For each pixel point of the image, if the pixel does not represent the sky, output its relative depth; if the pixel represents the sky, remove its depth information.

[0060] Step S4: Perform feature fusion on the results of steps S2 and S3. The model corrects the optical flow vector of each pixel with the depth information of a specific ratio. To reduce the computational complexity of the model and ensure the real-time performance of the prediction algorithm, we use the matrices of steps S2 and S3, that is, the optical flow vector matrix OF i,j and the relative depth matrix D i,jInput a convolutional layer with a 3×3 convolutional kernel size. For the obtained results OF′ and D′, perform feature fusion calculation. At the same time, noting that the optical flow vector is affected by the bumpy movement of the vehicle in the direction perpendicular to the ground, the model only retains the horizontal direction of the optical flow vector and sets the other dimension to zero and discards it. The expression for the entire process of feature fusion calculation is as follows:

[0061]

[0062] Among them, i represents the abscissa of the pixel point; j represents the ordinate of the pixel point; F′ 0,i,j represents the dimension indicating the vertical direction movement in the fusion feature matrix; F′ 1,i,j represents the dimension indicating the horizontal direction movement in the fusion feature matrix; OF′ 1,i,j represents the modified optical flow feature matrix calculated by the convolutional layer; D′ i,j represents the modified depth feature matrix calculated by the convolutional layer. For the obtained results, input them into a transposed convolutional layer with a 3×3 convolutional kernel size to obtain the feature matrix F i,j , which is used as the input of the prediction model.

[0063] Step S5: For the result of step S4, that is, the three-dimensional matrix, take three sub-matrices of the same size and merge them as the input of the real-time accurate speed prediction neural network in the subsequent steps. Take the sub-matrix of the result of step S4, that is, the feature matrix F. Specifically: take three 200×60 two-dimensional matrices. For these three sub-matrices, the coordinates corresponding to the (0,0) coordinate in the original matrix are (113,200), (513,200), and (913,200). Discard F 0,i,j , and take the F 1,i,j part as the input of the subsequent speed prediction network.

[0064] Step S6: Use the designed speed prediction neural network for operation, and output the predicted speed for each adjacent two-frame camera image. Input the output image of the previous step into a pre-trained convolutional neural network for prediction; the input of the speed prediction module network is the result of step S5, that is, a 3×200×66 matrix; the speed prediction module network includes 1 normalization layer, 3 convolutional layers with a convolutional kernel size of 5×5 and a stride of 2, 3 convolutional layers with a convolutional kernel size of 3×3, and 4 fully connected layers.

[0065] The learning rate optimization method of the model selects the Adam (Adaptive Moment Estimation) algorithm, which dynamically adjusts the learning rate of each parameter using the first-order moment estimation and second-order moment estimation of the gradient. The loss function selects L2 Loss to avoid overfitting, and its expression is as follows:

[0066]

[0067] Among them, J MSE represents the L2 Loss function; y i represents the speed output value of the speed prediction neural network; represents the true speed value of the benchmark.

[0068] Step S7: Output a warning message according to the prediction result of Step S6. If it meets the dangerous driving mode, a warning is output. This Step S7 is specifically as follows: Analyze the speed information of this frame and the previous frames, etc. If the acceleration is too fast or the attitude change range is large, a warning is given; specifically: Input the speed represented by this frame image and the speeds of the previous three frames into a judgment module. If its speed mode meets the dangerous driving mode (such as too high vehicle speed, too low vehicle speed, too fast speed change, etc.), an alarm message is output, and finally the predicted speed is synchronously uploaded to the server of the third-party cloud platform.

[0069] The embodiment of the present invention provides a method and system for vehicle driving safety warning. Aiming at the current insufficient utilization of the data obtained by the driving recorder, and some existing technologies are mostly offline on databases such as Kitti and need to collect the whole process of vehicle movement before prediction can be carried out. The real-time and accurate third-party speed prediction algorithm realized by the present invention can output real-time speed prediction results during the operation of the vehicle without relying on the vehicle's own equipment and synchronize with the third-party cloud platform. And the purpose of the present invention is to propose an effective method for the vehicle safety warning system through these data. This method can fill the application gap in the current market, is independent of other in-vehicle devices or sensors, and has the advantages of third-party authority, low cost, and good effect.

[0070] The present invention intelligently detects comprehensive information such as vehicles, obstacles, pedestrians, lanes, driving speed, etc. in the front and side directions on the in-vehicle video dataset from the first-person perspective by fusing the training of the optical flow depth neural network, and gives real-time reminders to the driver to ensure safe driving of the vehicle. The content of the present invention includes the establishment of the dataset, the training of the fused optical flow depth model, and the subsequent visualization demonstration effect work. It has great pioneering significance for the intelligentization of automobile driving and the industrialization of the safe driving AI assistant.

[0071] Those skilled in the art know that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same functions. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a kind of hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structures within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as either software modules for implementing the method or structures within the hardware component.

[0072] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A vehicle driving safety warning method, characterized in that, Including: Step S1: Obtain two adjacent frames of images through an in-vehicle camera or a driving recorder as the model input for subsequent applications, and stretch the images to conform to the camera parameters. Step S2: Use the modified FastFlowNet optical flow prediction neural network model, with two adjacent frames of images as the input, and output an optical flow vector matrix, where each vector indicates the change in the illumination condition of the pixel points, serving as the input to the feature fusion module in the subsequent steps. Step S3: Use the MegaDepth neural network model, with two adjacent frames of images as the input, and output a depth matrix, serving as the input to the feature fusion module in the subsequent steps. Under the condition allowed by the computing device, this step S3 can run synchronously with step S2. Step S4: Perform feature fusion on the results of step S2 and step S3. Step S5: For the result of step S4, that is, the sub-matrix of the feature matrix F, specifically: take three 200×60 two-dimensional matrices, take sub-matrices of the same size from these three parts, and merge them as the input to the real-time accurate speed prediction neural network in the subsequent steps. Step S6: Use the designed speed prediction neural network for operation, and output the predicted speed for each pair of adjacent captured images obtained. Step S7: Output a warning message according to the prediction result of step S6. If it meets the dangerous driving mode, output a warning. The design of the FastFlowNet optical flow prediction neural network model in step S2 is as follows: Step S2.1: Head-enhanced pooling pyramid: The FastFlowNet model uses a head-enhanced pooling pyramid for feature extraction. It fuses the feature pyramid with a higher number of layers and the pooling pyramid with a lower number of layers to combine the advantages of both; at the same time, add a convolutional layer on the high-level pyramid to strengthen the pyramid features through computational cost. Step S2.2: Center Dense Dilated Correlation (CDDC) layer: Different from the traditional convolutional layer structure, the model uses a center dense dilated correlation layer to perform dense sampling in the central part of the image, and at the same time downsample the grid points in the large motion area, effectively reducing the computational amount while maintaining the model's cognitive radius; when performing stereo matching, the cost volume function of the model is constructed as follows: Among them, c l (x, d) represents the cost volume function; l represents the l-th layer of the feature pyramid; x represents the vector pointing from the lower left corner of the image to the pixel point; d represents the vector pointing from the center of the image to the pixel point; represents the feature function of the image; represents the feature function of the warped image; N represents the dimension of the feature as the model input; r represents the search radius of the convolution; Step S2.3: Aggregation block decoder: The compact cost volume constructed by CDDC reduces the maximum feature channels of the decoder from 128 to 96; change three 96-channel convolutions to group convolutions, and each decoder network contains three aggregation blocks with a group number of 3. The matrices in steps S2 and S3, namely the optical flow vector matrix OF i,j and the relative depth matrix D i,j are input into a convolutional layer with a three-layer convolutional kernel size of 3×3. For the obtained results OF′ and D′, feature fusion calculation is then performed. The expression for the entire process of feature fusion calculation is as follows: Among them, i represents the abscissa of the pixel point; j represents the ordinate of the pixel point; F′ 0,i,j represents the dimension indicating the vertical direction movement in the fusion feature matrix; F′ 1,i,j represents the dimension indicating the horizontal direction movement in the fusion feature matrix; OF′ 1,i,j represents the modified optical flow feature matrix calculated by the convolutional layer; D′ i,j represents the modified depth feature matrix calculated by the convolutional layer; For the obtained results, input them into a transposed convolutional layer with a convolutional kernel size of 3×3 for three layers to obtain the feature matrix F i,j , which serves as the input to the prediction model; Step S6 includes: Input the output image of the previous step into a pre-trained convolutional neural network for prediction; the input of the speed prediction module network is the result of step S5, that is, a 3×200×66 matrix; the speed prediction module network includes 1 normalization layer, 3 convolutional layers with a convolution kernel size of 5×5 and a stride of 2, 3 convolutional layers with a convolution kernel size of 3×3, and 4 fully connected layers.

2. The vehicle driving safety warning method according to claim 1, wherein The step S1 includes: obtaining the video stream input of the camera, directly using every two adjacent frames of images as the input for subsequent modules, and stretching or compressing the resolution of the source RGB image to a 1266×370 RGB image using an interpolation method.

3. The vehicle driving safety warning method according to claim 1, wherein The step S2 includes: taking the two frames of images output from step S1 as the input, calculating the optical flow vector matrix, and the output result is a vector field representing the change in the illumination conditions of each pixel in the image, which is stored as a 2×1266×370 matrix.

4. The vehicle driving safety warning method according to claim 1, characterized in that, The step S3 includes: taking the RGB image obtained in step S1 as the input, performing speed prediction through the MegaDepth network. For each pixel point in the image, if the pixel does not represent the sky, its relative depth is output; if the pixel represents the sky, its depth information is removed.

5. The vehicle driving safety warning method according to claim 1, wherein The said step S5 includes: taking the result of step S4. For these three sub-matrices, the coordinates corresponding to their (0, 0) coordinates in the original matrix are (113, 200), (513, 200), and (913, 200), and discarding F 0,i,j , taking the F 1,i,j part as the input of the subsequent speed prediction network.

6. The vehicle driving safety warning method according to claim 1, characterized in that, The step S7 includes: analyzing the speed information of this frame and the previous frames. If the acceleration is too fast or the attitude change amplitude is large, a warning is issued. Specifically: input the speed represented by this frame of image and the speeds of the previous three frames into a judgment module. If its speed pattern conforms to the dangerous driving mode, an alarm message is output, and finally the predicted speed is synchronously uploaded to the server of the cloud third-party platform.

7. A vehicle driving safety warning system, characterized in that, Includes: Module M1: Obtaining two adjacent frames of images through an in-vehicle camera or a driving recorder as the input for the subsequent application model, and stretching the images to conform to the camera parameters. Module M2: Using the modified FastFlowNet optical flow prediction neural network model, taking two adjacent frames of images as the input, and outputting the optical flow vector matrix, where each vector indicates the change in the illumination conditions of the pixel points, as the input for the feature fusion module in the subsequent steps. Module M3: Using the MegaDepth neural network model, taking two adjacent frames of images as the input, and outputting the depth matrix, as the input for the feature fusion module in the subsequent steps. Under the condition allowed by the computing device, this module M3 can run synchronously with module M2. Module M4: Performing feature fusion on the results of module M2 and module M3. Module M5: For the result of module M4, that is, the sub-matrix of the feature matrix F. Specifically: taking three 200×60 two-dimensional matrices, taking sub-matrices of the same size from these three parts, and merging them as the input for the real-time accurate speed prediction neural network in the subsequent steps. Module M6: Performing operations using the designed speed prediction neural network, and outputting the predicted speed for every two adjacent frames of captured images obtained. Module M7: Outputting a warning message according to the prediction result of module M6. If it conforms to the dangerous driving mode, a warning is output. The design of the FastFlowNet optical flow prediction neural network model in module M2 is as follows: Module M2.1: Head enhanced pooling pyramid: The FastFlowNet model uses the head enhanced pooling pyramid for feature extraction. It fuses the higher-layer feature pyramid with the lower-layer pooling pyramid to combine the advantages of both. At the same time, a convolutional layer is added to the high-level pyramid to strengthen the pyramid features through computational cost. Module M2.2: Center Dense Dilated Convolution Layer (CDDC): Different from the traditional convolution layer structure, the model adopts the center dense dilated convolution layer to densely sample the central part of the image, and at the same time downsample the grid points in the large motion area, effectively reducing the computational amount while maintaining the model's cognitive radius; when performing stereo matching, the cost volume function of the model is constructed as follows: Among them, c l (x, d) represents the Cost Volume function; l represents the l-th layer of the feature pyramid; x represents the vector pointing from the lower left corner of the image to the pixel point; d represents the vector pointing from the center of the image to the pixel point; represents the feature function of the image; represents the feature function of the warped image; N represents the dimension of the feature as the model input; r represents the search radius of the convolution; Module M2.3: Aggregation Block Decoder: The compact cost volume constructed by CDDC reduces the maximum feature channels of the decoder from 128 to 96; changes the three 96-channel convolutions to group convolutions, and each decoder network contains three aggregation blocks with a group number of 3; The matrices of module M2 and module M3, namely the optical flow vector matrix OF i,j and the relative depth matrix D i,j are input into a convolutional layer with a three-layer convolutional kernel size of 3×3. For the obtained results OF′ and D′, feature fusion calculation is then performed. The expression for the entire process of feature fusion calculation is as follows: where, i represents the abscissa of the pixel; j represents the ordinate of the pixel; F′ 0,i,j represents the dimension indicating the vertical direction movement in the fusion feature matrix; F′ 1,i,j represents the dimension indicating the horizontal direction movement in the fusion feature matrix; OF′ 1,i,j represents the modified optical flow feature matrix calculated through the convolutional layer; D′ i,j represents the modified depth feature matrix calculated through the convolutional layer; For the obtained results, input them into a transposed convolutional layer with a three-layer convolutional kernel size of 3×3 to obtain the feature matrix F i,j , which serves as the input to the prediction model; The said module M6 includes: inputting the output image of the previous step into a pre-trained convolutional neural network for prediction; the input of the speed prediction module network is the result of step S5, that is, a 3×200×66 matrix; the speed prediction module network includes 1 normalization layer, 3 convolutional layers with a convolutional kernel size of 5×5 and a stride of 2, 3 convolutional layers with a convolutional kernel size of 3×3, and 4 fully connected layers.

Citation Information

Patent Citations

  • Implementation method based on convolutional neural network and optical flow characteristic visual odometer

    CN112419411A

  • Vehicle early warning and control system based on computer vision

    CN113269962A

  • Low power proximity-based presence detection using optical flow

    US20240404296A1