A dangerous behavior warning method based on the fusion of driver visual features
Through the visual multimodal feature fusion network, the driver's visual information is processed, and the driver's visual information is not used in the prior art is solved, and the driver's vision and dangerous behavior are accurately predicted, which improves the safety of the autonomous driving system.
Patent Information
- Application Number
- CN202411553043.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-11-01
AI Technical Summary
The existing driver driving behavior prediction model fails to effectively utilize driver visual information, resulting in insufficient prediction of driver vision direction and dangerous behavior, affecting the safety of the autonomous driving system.
Using a visual multimodal feature fusion network, by establishing a feature preprocessing model and a cross-modal feature fusion module, the feature data of different modes is fused by using depth separation convolution to generate a fusion feature map, and the precise processing and prediction of driver visual information is realized.
It improves the driver's eye direction and accurate prediction of dangerous driving behavior, and enhances the early warning efficiency of potential driving risks.
Smart Images

Figure CN119516519B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of visual feature perception, and in particular to a dangerous behavior warning method based on driver visual feature fusion. Background Art
[0002] With the application of artificial intelligence in intelligent transportation and autonomous driving, the construction of prediction models for driver driving behavior has become an important support for realizing traffic safety and efficient driving. These prediction models can better understand and predict the behavior of drivers during driving, thereby improving the performance and safety of autonomous driving systems.
[0003] However, the reference indicators of the current prediction models for driver driving behavior often only include the external environment parameters during vehicle driving and the operating parameters of the vehicle itself, and there is no model for predicting the driver's line of sight direction and driving behavior in the driving scenario. As the largest participant in the traffic system, the driver is the dominant factor affecting road traffic safety. 90% of the road traffic information required during driving comes from vision. Summary of the Invention
[0004] The purpose of this application is: to solve the above technical problems, this application provides a dangerous behavior warning method based on driver visual feature fusion, aiming to improve the processing efficiency and accuracy of driver visual features and the warning efficiency of driver dangerous behaviors.
[0005] In some embodiments of this application, based on a visual multi-modal feature fusion network, corresponding feature preprocessing feature models are established according to different visual modalities to process the feature data of three modalities, correct different modal feature data using the complementarity between different modalities, and through an embedded cross-modal feature fusion module, use depthwise separable convolution to fuse the feature data of different modalities to generate a fused feature map, realizing information fusion between different modalities and improving the accuracy of visual information understanding.
[0006] In some embodiments of this application, by establishing a time series, continuously collect the image data during the driver's driving process, and according to a preset visual processing model, achieve precise processing of the image data during the driver's driving process, improve the precise prediction of the driver's line of sight direction and dangerous driving behaviors, and thus improve the warning efficiency of potential driving risks.
[0007] In some embodiments of this application, a dangerous behavior warning method based on driver visual feature fusion is provided, including:
[0008] Generate a multi-modal data packet based on the collected multi-source image data. The multi-modal data packet includes: primary modal data, secondary modal data, and tertiary modal data;
[0009] Process all the data in the multi-modal data packet according to a preset fusion model, and generate a fusion feature image based on the fusion result;
[0010] Establish a visual perception model, and generate a driver's visual perception result based on the visual perception model and the fusion feature image;
[0011] Generate a corresponding warning instruction according to the driver's visual perception result.
[0012] In some embodiments of the present application, when generating the multi-modal data packet, it includes:
[0013] Obtain multi-source image data;
[0014] Establish a time series to form a time series sequence T, T=(t1, t2……t m ); where t i is the i-th moment, and m is the number of moments;
[0015] Successively set t i as the target moment;
[0016] Extract the RGB image features of the target moment in the multi-source image data according to the preprocessing feature model, and generate primary modal data based on the extraction result;
[0017] Extract the infrared features of the target moment in the multi-source image data according to the preprocessing feature model, and generate secondary modal data based on the extraction result;
[0018] Extract the depth features of the target moment in the multi-source image data according to the preprocessing feature model, and generate tertiary modal data based on the extraction result.
[0019] In some embodiments of the present application, when generating the fusion feature image according to the fusion result, it includes:
[0020] Establish a feature correction model;
[0021] Correct the primary modal data, secondary modal data, and tertiary modal data according to the feature correction model;
[0022] Generate a primary feature to be fused, a secondary feature to be fused, and a tertiary feature to be fused according to the correction result;
[0023] Perform linear embedding processing and feature fusion on the primary feature to be fused, the secondary feature to be fused, and the tertiary feature to be fused according to the depthwise separable convolution technology;
[0024] Generate the fusion feature image of the target moment;
[0025] Generate the fused feature images at each moment in sequence, and establish a sequence B of fused feature images, B = (b1, b2... b m ); where b i is the fused feature image at the i-th moment.
[0026] In some embodiments of the present application, when establishing the visual perception model, it includes:
[0027] Based on historical parameters, establish a plurality of sample training data packets, and establish a sequence W of sample training data packets, W = (w1, w2... w n ), where w i is the i-th sample training data packet; n is the number of sample training data packets;
[0028] Select w i in sequence as the target sample training data packet;
[0029] Generate a sequence I of face images according to the target sample training data packet, I = (I1, I2... I t );
[0030] The sequence I of face images generates a feature mapping sequence after feature extraction;
[0031] Construct a primary processing model according to the feature mapping sequence;
[0032] Generate a loss evaluation value of the primary processing model, and determine whether to generate a correction instruction according to the loss evaluation value;
[0033] Output the gaze prediction sub-model.
[0034] In some embodiments of the present application, when determining whether to generate a correction instruction, it includes:
[0035] Preset a first loss evaluation value threshold F1 and a second loss evaluation value threshold F2;
[0036] If f < F1, set the current primary processing model as the gaze prediction sub-model;
[0037] If F1 ≤ f < F2, generate a primary correction instruction, and set the next sample training data packet as the target sample data packet according to the primary correction instruction, and perform iterative optimization on the current primary processing model;
[0038] If f ≥ F2, generate a secondary correction instruction, update the sequence of sample training data packets according to the secondary correction instruction, re-select the target sample data packet according to the update result, and construct a primary processing model.
[0039] In some embodiments of the present application, when establishing the visual perception model, it further includes:
[0040] Generate multiple dangerous behavior data packets based on historical parameters;
[0041] Construct a behavior prediction sub-model according to all the dangerous behavior data packets;
[0042] Construct a visual perception model according to the behavior recognition sub-model and the line-of-sight prediction sub-model.
[0043] In some embodiments of the present application, when generating the driver's visual perception result, it includes:
[0044] Obtain the fused feature image sequence B;
[0045] Generate a line-of-sight deviation sequence H according to the fused feature image sequence B and the line-of-sight prediction sub-model, H=(h1, h2...h m ), where h i is the line-of-sight deviation at the i-th moment;
[0046] Generate a first reference evaluation value J1 according to the line-of-sight deviation sequence H;
[0047]
[0048] Among them, e1 is a preset first weight coefficient, e2 is a preset second weight coefficient, Q1 is a preset first fixed coefficient, Q2 is a preset second fixed coefficient, h' is the safety deviation, Y(i) is the first selection coefficient, if h i -h'≤0, Y(i)=0, if h i -h'>0, Y(i)=1, and ΔT is the time interval between adjacent moments.
[0049] In some embodiments of the present application, when generating the driver's visual perception result, it further includes:
[0050] Obtain the fused feature image data B;
[0051] Generate a behavior risk evaluation value sequence D according to the fused feature image sequence B and the behavior recognition sub-model, D=(d1, d2...d m ), where d i is the behavior risk evaluation value at the i-th moment;
[0052] Generate a second reference evaluation value J2 according to the behavior risk evaluation value sequence D
[0053]
[0054] Among them, e3 is a preset third weight coefficient, e4 is a preset fourth weight coefficient, Q3 is a preset third fixed coefficient, Q4 is a preset fourth fixed coefficient, d' is the safe behavior value, P(i) is the second selection coefficient, if d i-d' ≤ 0, P(i) = 0, if d i -d' > 0, P(i) = 1, where ΔT is the time interval between adjacent moments.
[0055] In some embodiments of the present application, corresponding warning instructions are generated according to the driver's visual perception results, including:
[0056] Generate a risk evaluation value k according to the first reference evaluation value J1 and the second reference evaluation value J2;
[0057] k = e5 * J1 + e6 * J2;
[0058] where e5 is a preset fifth weight coefficient and e6 is a preset sixth weight coefficient;
[0059] Set the risk evaluation value K as the driver's visual perception result, and preset the first risk evaluation value threshold K1 and the second risk evaluation value threshold K2;
[0060] If k ≤ K1, no warning instruction is generated;
[0061] If K1 < k ≤ K2, a first-level warning instruction is generated;
[0062] If k > K2, a second-level warning instruction is generated.
[0063] Compared with the prior art, the beneficial effects of a dangerous behavior warning method based on driver visual feature fusion in an embodiment of the present application are as follows:
[0064] Based on a visual multi-modal feature fusion network, corresponding feature preprocessing feature models are established according to different visual modalities, so as to process the feature data of three modalities, correct different modal feature data by using the complementarity between different modalities, and through an embedded cross-modal feature fusion module, use depthwise separable convolution to fuse the feature data of different modalities to generate a fused feature map, realizing information fusion between different modalities and improving the accuracy of visual information understanding.
[0065] By establishing a time series, continuously collect the image data during the driver's driving process, and according to a preset visual processing model, achieve precise processing of the image data during the driver's driving process, improve the precise prediction of the driver's line of sight direction and dangerous driving behavior, and thus improve the warning efficiency for potential driving risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a schematic flowchart of a dangerous behavior warning method based on driver visual feature fusion in a preferred embodiment of an embodiment of the present application. DETAILED DESCRIPTION
[0067] The following will further describe in detail the specific implementation manners of the present application with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0068] In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present application.
[0069] The terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "plurality" is two or more.
[0070] In the description of the present application, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0071] As Figure 1 shown, a method for warning of dangerous behaviors based on driver visual feature fusion in a preferred embodiment of an embodiment of the present application includes:
[0072] S101: Generate a multimodal data packet according to the collected multi-source image data. The multimodal data packet includes: primary modal data, secondary modal data, and tertiary modal data;
[0073] S102: Process all the data in the multimodal data packet according to a preset fusion model, and generate a fusion feature image according to the fusion result;
[0074] S103: Establish a visual perception model, and generate a driver visual perception result according to the visual perception model and the fusion feature image;
[0075] S104: Generate a corresponding warning instruction according to the driver visual perception result.
[0076] Specifically, when generating the multimodal data packet, it includes:
[0077] Obtain multi-source image data;
[0078] Based on the time series, establish a time series T, T = (t1, t2... t m ); where t i is the i-th moment, and m is the number of moments;
[0079] Successively set t i as the target moment;
[0080] Extract the RGB image features of the target moment in the multi-source image data according to the preprocessing feature model, and generate the first-level modal data according to the extraction results;
[0081] Extract the infrared features of the target moment in the multi-source image data according to the preprocessing feature model, and generate the second-level modal data according to the extraction results;
[0082] Extract the depth features of the target moment in the multi-source image data according to the preprocessing feature model, and generate the third-level modal data according to the extraction results.
[0083] Specifically, taking ViT as the backbone network, according to different visual modalities, establish a preprocessing model to process the data of the three modalities respectively, and embed a cross-modal feature fusion module after each layer of ViT to fuse the features of the data of each layer.
[0084] Specifically, when generating the fused feature image according to the fusion result, it includes:
[0085] Establish a feature correction model;
[0086] Correct the first-level modal data, second-level modal data, and third-level modal data according to the feature correction model;
[0087] Generate the first-level to-be-fused feature, second-level to-be-fused feature, and third-level to-be-fused feature according to the correction results;
[0088] Perform linear embedding processing and feature fusion on the first-level to-be-fused feature, second-level to-be-fused feature, and third-level to-be-fused feature according to the depthwise separable convolution technology;
[0089] Generate the fused feature image of the target moment;
[0090] Successively generate the fused feature images of each moment, and establish a fused feature image series B, B = (b1, b2... b m ); where b i is the fused feature image of the i-th moment.
[0091] Specifically, due to the complementarity of information between different modalities, when processing data of different modalities, it is divided into two steps. The first step is to correct different modal features by utilizing the complementarity between different modalities. The feature correction stage can be divided into the channel-wise stage and the spatial-wise stage. Through the feature correction stage, the features of different modalities can be supplemented and corrected. The second step is to design a cross-modal attention mechanism based on the principle of the attention mechanism. First, linearly embed the corrected feature maps of different modalities to generate corresponding Q, K, and V vectors, then exchange information between different modalities, and use depthwise separable convolution for feature fusion to generate the final fused feature image, realizing information fusion between different modalities and improving the accuracy of visual information understanding.
[0092] In a preferred embodiment of the present application, when establishing a visual perception model, it includes:
[0093] Establish a plurality of sample training data packets based on historical parameters, and establish a sequence of sample training data packets W, W=(w1, w2…w n ), where w i is the i-th sample training data packet; n is the number of sample training data packets;
[0094] Successively select w i as the target sample training data packet;
[0095] Generate a sequence of face images I according to the target sample training data packet, I=(I1, I2…I t );
[0096] The sequence of face images I generates a feature mapping sequence after feature extraction;
[0097] Construct a first-level processing model according to the feature mapping sequence;
[0098] Generate a loss evaluation value of the first-level processing model, and determine whether to generate a correction instruction according to the loss evaluation value;
[0099] Output a gaze prediction sub-model.
[0100] Specifically, the sample training data packet is a collected face video. Based on the gaze estimation model of ViT, learn the temporal information between consecutive frames from the video containing the face, thereby enhancing the accuracy and generalization performance of the model, and extract feature data of different modalities according to the preprocessing feature model, and generate a fused feature image of the corresponding training sample after correction and fusion.
[0101] Specifically, the loss evaluation value of the model can be set according to the average angular error between the predicted value and the true value. The larger the angular error, the larger the corresponding loss evaluation value.
[0102] Specifically, when determining whether to generate a correction instruction, it includes:
[0103] Preset a first loss evaluation value threshold F1 and a second loss evaluation value threshold F2;
[0104] If f < F1, set the current first-level processing model as the line-of-sight prediction sub-model;
[0105] If F1 ≤ f < F2, generate a first-level correction instruction, and set the next sample training data packet as the target sample data packet according to the first-level correction instruction, and perform iterative optimization on the current first-level processing model;
[0106] If f ≥ F2, generate a second-level correction instruction, update the sequence of sample training data packets according to the second-level correction instruction, re-select the target sample data packet according to the update result, and construct a first-level processing model.
[0107] Specifically, the first-level correction instruction means that the judgment accuracy of the current first-level processing model for the line of sight is relatively low, and it is necessary to use the remaining sample training data packets for iterative training to further improve the processing accuracy of the visual processing model.
[0108] Specifically, the second-level correction instruction means that the data feature extraction, calibration, and fusion in the current training sample data packet are relatively poor, resulting in deviation of the first-level processing model, and it is necessary to regenerate the corresponding sample training data.
[0109] It can be understood that in the above embodiments, the line-of-sight estimation model based on ViT learns the temporal information between consecutive frames from the video containing the face, thereby enhancing the accuracy and generalization performance of the model.
[0110] Specifically, when establishing the visual perception model, it also includes:
[0111] Generate multiple dangerous behavior data packets based on historical parameters;
[0112] Construct a behavior prediction sub-model according to all the dangerous behavior data packets;
[0113] Construct a visual perception model according to the behavior recognition sub-model and the line-of-sight prediction sub-model.
[0114] Specifically, according to the collected historical parameters, a corresponding behavior prediction sub-model is constructed through iterative training. The behavior prediction sub-model can identify whether there are dangerous behaviors of the driver in the collected images and evaluate the dangerous behaviors, so as to generate the behavior risk evaluation values at each moment. The behavior risk evaluation value indicates that there are more dangerous behaviors in the current driving behavior.
[0115] Specifically, dangerous behaviors refer to non-standard driving behaviors. The more dangerous behaviors there are, the greater the possibility of accidents during driving.
[0116] In the preferred embodiment of the present application, when generating the driver's visual perception result, it includes:
[0117] Obtain the sequence B of fused feature images;
[0118] Generate the sequence H of line-of-sight deviation degrees according to the sequence B of fused feature images and the line-of-sight prediction sub-model, H=(h1, h2... h m ), where h i is the line-of-sight deviation degree at the i-th moment;
[0119] Generate the first reference evaluation value J1 according to the sequence H of line-of-sight deviation degrees;
[0120]
[0121] Among them, e1 is a preset first weight coefficient, e2 is a preset second weight coefficient, Q1 is a preset first fixed coefficient, Q2 is a preset second fixed coefficient, h' is the safety deviation degree, Y(i) is the first selection coefficient. If h i -h'≤0, Y(i)=0. If h i -h'>0, Y(i)=1, and ΔT is the time interval between adjacent moments.
[0122] Specifically, when generating the driver's visual perception result, it further includes:
[0123] Obtain the fused feature image data B;
[0124] Generate the sequence D of behavior risk evaluation values according to the sequence B of fused feature images and the behavior recognition sub-model, D=(d1, d2... d m ), where d i is the behavior risk evaluation value at the i-th moment;
[0125] Generate the second reference evaluation value J2 according to the sequence D of behavior risk evaluation values
[0126]
[0127] Among them, e3 is a preset third weight coefficient, e4 is a preset fourth weight coefficient, Q3 is a preset third fixed coefficient, Q4 is a preset fourth fixed coefficient, d' is a safety behavior value, P(i) is a second selection coefficient. If d i -d'≤0, P(i) = 0. If d i -d'>0, P(i) = 1, and ΔT is the time interval between adjacent moments.
[0128] Specifically, corresponding warning instructions are generated according to the driver's visual perception results, including:
[0129] Generate a danger evaluation value k according to the first reference evaluation value J1 and the second reference evaluation value J2;
[0130] k = e5*J1 + e6*J2;
[0131] Among them, e5 is a preset fifth weight coefficient, and e6 is a preset sixth weight coefficient;
[0132] Set the danger evaluation value K as the driver's visual perception result, and preset the first danger evaluation value threshold K1 and the second danger evaluation value threshold K2;
[0133] If k≤K1, no warning instruction is generated;
[0134] If K1<k≤K2, generate a first-level warning instruction;
[0135] If k>K2, generate a second-level warning instruction.
[0136] Specifically, the first-level warning instruction means that the current driver has problems such as irregular driving or line-of-sight deviation, and the driver's driving behavior can be corrected through voice reminders. The second-level warning instruction means that when the driver has serious dangerous driving behaviors, it is necessary to give a warning immediately and take necessary measures to prevent driving accidents. By establishing a danger evaluation model, the warning efficiency for the driver's dangerous behaviors is improved.
[0137] Specifically, the greater the danger evaluation value, the greater the possibility that the current driver has driving risks, and risk warnings need to be given in a timely manner. By establishing a danger evaluation model, the warning efficiency for the driver's dangerous behaviors is improved.
[0138] According to the first concept of the present application, based on the visual multi-modal feature fusion network, corresponding feature preprocessing feature models are established according to different visual modalities to process the feature data of three modalities, correct the different modality feature data by using the complementarity between different modalities, and through embedding a cross-modal feature fusion module, use depthwise separable convolution to fuse the feature data of different modalities to generate a fused feature map. The information fusion between different modalities is realized, and the accuracy of visual information understanding is improved.
[0139] According to the second concept of the present application, by establishing a time series, continuously collect the image data during the driver's driving process, and according to the preset visual processing model, achieve precise processing of the image data during the driver's driving process, improve the precise prediction of the driver's line of sight direction and dangerous driving behaviors, thereby improving the early warning efficiency for potential driving risks.
[0140] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the technical principle of the present application, several improvements and replacements can also be made, and these improvements and replacements should also be regarded as the protection scope of the present application.
Claims
1. A warning method for dangerous behaviors based on the fusion of driver visual features, characterized in that, It includes: Generating a multimodal data packet based on the collected multi-source image data, where the multimodal data packet includes: primary modal data, secondary modal data, and tertiary modal data; Processing all the data in the multimodal data packet according to a preset fusion model, and generating a fused feature image based on the fusion result; Establishing a visual perception model, and generating a driver's visual perception result based on the visual perception model and the fused feature image; Generating a corresponding warning instruction according to the driver's visual perception result; When generating the multimodal data packet, it includes: Obtaining multi-source image data; Based on the time series, a time series sequence T is established, T = (t1, t2... t m ); where, t i is the i-th moment, and m is the number of moments; Set t sequentially i as the target time; Extracting the RGB image features of the target moment in the multi-source image data according to a preprocessing feature model, and generating primary modal data based on the extraction result; Extracting the infrared features of the target moment in the multi-source image data according to a preprocessing feature model, and generating secondary modal data based on the extraction result; Extracting the depth features of the target moment in the multi-source image data according to a preprocessing feature model, and generating tertiary modal data based on the extraction result; When generating a fused feature image based on the fusion result, it includes: Establishing a feature correction model; Correcting the primary modal data, secondary modal data, and tertiary modal data according to the feature correction model; Generating a primary feature to be fused, a secondary feature to be fused, and a tertiary feature to be fused according to the correction result; Performing linear embedding processing and feature fusion on the primary feature to be fused, secondary feature to be fused, and tertiary feature to be fused according to the depthwise separable convolution technique; Generating a fused feature image of the target moment; Generate the fused feature images at each moment in sequence, and establish a sequence B of fused feature images, B = (b1, b2... b m ); where b i is the fused feature image at the i-th moment.
2. The warning method for dangerous behaviors based on driver visual feature fusion according to claim 1, wherein When establishing a visual perception model, it includes: Establish a plurality of sample training data packets based on historical parameters, and establish a sequence of sample training data packets W, W=(w1, w2…w n ), where w i is the i-th sample training data packet; n is the number of sample training data packets; Select w in sequence i as the training data packet for the target sample; Generate a sequence of face images I according to the training data packet of the target sample, I = (I1, I2… I t ); The sequence of face images I generates a feature mapping sequence after feature extraction; Constructing a primary processing model according to the feature mapping sequence; Generating a loss evaluation value of the primary processing model, and judging whether to generate a correction instruction according to the loss evaluation value; Outputting a line-of-sight prediction sub-model.
3. The method for warning of dangerous behaviors based on driver visual feature fusion according to claim 2, wherein, When judging whether to generate a correction instruction, it includes: Presetting a first loss evaluation value threshold F1 and a second loss evaluation value threshold F2; If f < F1, setting the current primary processing model as the line-of-sight prediction sub-model; If F1 ≤ f < F2, generating a primary correction instruction, and setting the next sample training data packet as the target sample data packet according to the primary correction instruction, and iteratively optimizing the current primary processing model; If f ≥ F2, generating a secondary correction instruction, and updating the sequence of sample training data packets according to the secondary correction instruction, reselecting the target sample data packet according to the update result, and constructing a primary processing model; Where f is the loss evaluation value of the primary processing model.
4. The method for warning of dangerous behaviors based on driver visual feature fusion according to claim 3, wherein When establishing a visual perception model, it also includes: Generating multiple dangerous behavior data packets based on historical parameters; Constructing a behavior prediction sub-model according to all the dangerous behavior data packets; Constructing a visual perception model according to the behavior recognition sub-model and the line-of-sight prediction sub-model.
5. The dangerous behavior warning method based on driver visual feature fusion according to claim 4, wherein, When generating a driver's visual perception result, it includes: Obtaining a sequence of fused feature images B; Generate the line-of-sight deviation sequence H based on the fused feature image sequence B and the line-of-sight prediction sub-model, H = (h1, h2…h m ), where h i is the line-of-sight deviation at the i-th moment; Generating a first reference evaluation value J1 according to the sequence of line-of-sight deviation degrees H; Among them, e1 is a preset first weight coefficient, e2 is a preset second weight coefficient, Q1 is a preset first fixed coefficient, Q2 is a preset second fixed coefficient, h' is a safety deviation degree, Y(i) is a first selection coefficient. If h i -h' ≤ 0, Y(i) = 0. If h i -h' > 0, Y(i) = 1, and ΔT is the time interval between adjacent moments.
6. The warning method for dangerous behaviors based on driver visual feature fusion according to claim 5, characterized in that When generating a driver's visual perception result, it also includes: Obtaining a sequence of fused feature images B; Generate a sequence of behavioral risk evaluation values D based on the sequence of fused feature images B and the behavioral recognition sub-model, D = (d1, d2... d m ), where d i is the behavioral risk evaluation value at the i-th moment; Generating a second reference evaluation value J2 according to the sequence of behavior risk evaluation values D Among them, e3 is a preset third weight coefficient, e4 is a preset fourth weight coefficient, Q3 is a preset third fixed coefficient, Q4 is a preset fourth fixed coefficient, d' is a safety behavior value, P(i) is a second selection coefficient. If d i -d' ≤ 0, P(i) = 0. If d i -d' > 0, P(i) = 1. ΔT is the time interval between adjacent moments.
7. The method for warning of dangerous behaviors based on driver visual feature fusion according to claim 6, characterized in that, Generating a corresponding warning instruction according to the driver's visual perception result, including: Generate a danger evaluation value k based on a first reference evaluation value J1 and a second reference evaluation value J2; k = e5 * J1 + e6 * J2; where e5 is a preset fifth weight coefficient and e6 is a preset sixth weight coefficient; Set the danger evaluation value K as the driver's visual perception result, and preset a first danger evaluation value threshold K1 and a second danger evaluation value threshold K2; If k ≤ K1, no warning command is generated; If K1 < k ≤ K2, generate a first-level warning command; If k > K2, generate a second-level warning command.
Citation Information
Patent Citations
Visual perception detection method and system
CN118840633A
Scene classification prediction
US20200086879A1