Vehicle driving state detection method and device and storage medium

Through the video frame group segmentation and transfer learning strategy of the multi-stream fusion classification model, the problems of spatiotemporal feature fragmentation and computational resource waste in vehicle driving state detection are solved, and high-precision and efficient detection in complex road scenarios are achieved.

CN120526344AInactive Publication Date: 2025-08-22SHENZHEN DACHUAN SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510409172.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-08-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing vehicle driving state detection technology has problems such as spatiotemporal and spatial feature fragmentation, limited model generalization, waste of computing resources and poor migration adaptability, resulting in insufficient detection accuracy and low computing efficiency in complex road scenarios.

Method used

A multi-stream fusion classification model is adopted, including a spatial, spatial and temporal feature extraction module based on pre-trained deep convolutional networks. Through video frame group segmentation and transfer learning strategies, a multi-stream fusion classification model is constructed, and joint training and weighted fusion are carried out to generate the final detection results of the vehicle's driving state.

Benefits of technology

It significantly improves the accuracy of vehicle steering status recognition and speed estimation, enhances environmental adaptability in complex road scenarios, and reduces the amount of redundant calculations for video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526344A_ABST
    Figure CN120526344A_ABST
Patent Text Reader

Abstract

The invention relates to the field of traffic control, and provides a vehicle driving state detection method, which comprises the following steps: acquiring original video data of vehicle driving through video acquisition equipment; performing motion feature modeling on each frame group of the original video data to generate a motion feature image containing spatio-temporal information; a multi-stream fusion classification model is constructed, and the multi-stream fusion classification model comprises a spatial feature extraction module, a spatial-temporal feature extraction module and a time feature extraction module; inputting the motion feature image and the original video data into a multi-stream fusion classification model, and performing joint training on the multi-stream fusion classification model through a transfer learning strategy to obtain a trained multi-stream fusion classification model; and performing weighted fusion based on the classification probability value output by each independent stream in the trained multi-stream fusion classification model, and generating a final detection result of the vehicle driving state. According to the technical scheme of the invention, the steering, speed and other states of the vehicle in a complex road scene can be accurately estimated with low calculation cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of traffic control, and in particular to a vehicle driving state detection method, device and storage medium. Background Art

[0002] With the rapid development of urbanization, the number of vehicles on urban roads is increasing. Monitoring vehicle driving status, such as speed, direction, and road conditions, is a critical industry issue. This provides accurate data support and decision-making support for drivers, insurance companies, traffic police, and highway management departments, helping to reduce the incidence of road accidents and improve highway safety. Existing vehicle driving status detection technologies primarily rely on single-sensor data or single-frame image analysis. This approach suffers from technical drawbacks such as fragmented temporal and spatial features, limited model generalization, wasted computational resources, and poor migration adaptability. Summary of the Invention

[0003] The present application provides a vehicle driving state detection method, device and storage medium, which can accurately estimate the vehicle's steering, speed and other states in complex road scenarios at a low computing cost.

[0004] In one aspect, the present application provides a method for detecting a vehicle driving state, the method comprising:

[0005] Obtaining original video data of vehicle driving through video acquisition equipment;

[0006] Performing motion feature modeling on each frame group of the original video data to generate a motion feature image containing spatiotemporal information;

[0007] Constructing a multi-stream fusion classification model, wherein the multi-stream fusion classification model includes a spatial feature extraction module, a spatiotemporal feature extraction module, and a temporal feature extraction module based on a pre-trained deep convolutional network;

[0008] Inputting the motion feature image and the original video data into the multi-stream fusion classification model, and jointly training the multi-stream fusion classification model through a transfer learning strategy to obtain a trained multi-stream fusion classification model;

[0009] The classification probability values ​​output by each independent stream in the trained multi-stream fusion classification model are weightedly fused to generate a final detection result of the vehicle driving state.

[0010] On the other hand, the present application provides a vehicle driving state detection device, the device comprising:

[0011] An acquisition module is used to acquire original video data of the vehicle driving through a video acquisition device;

[0012] A first generating module is used to perform motion feature modeling on each frame group of the original video data to generate a motion feature image containing spatiotemporal information;

[0013] A construction module is used to construct a multi-stream fusion classification model, wherein the multi-stream fusion classification model includes a spatial feature extraction module, a spatiotemporal feature extraction module, and a temporal feature extraction module based on a pre-trained deep convolutional network;

[0014] a training module, configured to input the motion feature image and the original video data into the multi-stream fusion classification model, and jointly train the multi-stream fusion classification model through a transfer learning strategy to obtain a trained multi-stream fusion classification model;

[0015] The second generation module is used to perform weighted fusion based on the classification probability values ​​output by each independent stream in the trained multi-stream fusion classification model to generate a final detection result of the vehicle driving state.

[0016] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the technical solution of the vehicle driving state detection method described above are implemented.

[0017] In a fourth aspect, the present application provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the technical solution of the above-mentioned vehicle driving state detection method.

[0018] From the technical solutions provided in the present application, it can be seen that, on the one hand, by performing motion feature modeling on each frame group of the acquired original video data, the problem of spatiotemporal information separation in single-frame analysis is overcome, and the short-term trajectory characteristics and long-term trend characteristics of vehicle motion are effectively integrated, significantly improving the accuracy of steering state recognition and speed estimation; on the other hand, a multi-stream fusion classification model is constructed, including a spatial feature extraction module, a spatiotemporal feature extraction module and a temporal feature extraction module based on a pre-trained deep convolutional network, and the motion feature image and the original video data are input into the multi-stream fusion classification model, thereby realizing the complementary enhancement of spatial static features and motion dynamic features, and showing stronger environmental adaptability in complex road scenes; thirdly, the multi-stream fusion classification model is jointly trained by using a video frame group segmentation strategy combined with a transfer learning strategy, which can also significantly reduce the redundant computational amount of video processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 This is a flow chart of a vehicle driving state detection method provided by an embodiment of the present application;

[0021] Figure 2 Schematic diagram of the structure of the vehicle driving state detection device provided in an embodiment of the present application;

[0022] Figure 3 It is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0024] In this specification, adjectives such as first and second may be used only to distinguish one element or action from another element or action, without necessarily requiring or implying any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be construed as being limited to only one of the elements, components, or steps, but may be one or more of the elements, components, or steps, etc.

[0025] In this specification, for the convenience of description, the sizes of various parts shown in the drawings are not drawn according to the actual proportions.

[0026] With the rapid development of urbanization, the number of vehicles on urban roads is increasing. Monitoring vehicle driving status, such as speed, direction, and road conditions, is a key industry issue. This provides accurate data support and decision-making for drivers, insurance companies, traffic police, and highway management departments, helping to reduce the incidence of road accidents and improve highway safety. Existing vehicle driving status detection technology mainly relies on single sensor data or single-frame image analysis. This solution has the following main technical defects: 1) Split of spatiotemporal features, which is manifested in that existing technologies mostly adopt a frame-by-frame processing mode when processing video streams, and cannot effectively capture the continuous spatiotemporal correlation characteristics of vehicle motion, resulting in insufficient accuracy of dynamic detection such as steering recognition and speed estimation; 2) Limited model generalization, which is manifested in that existing deep learning models mostly adopt a single feature stream design, which makes it difficult to take into account both spatial morphological characteristics and motion trend characteristics at the same time, and is prone to misjudgment in complex road scenes; 3) Waste of computing resources, which is manifested in that when existing technologies directly process the original video sequence for the entire time period, redundant frame information leads to low computing efficiency, making it difficult to meet real-time detection needs; 4) Poor migration adaptability, which is manifested in that when the pre-trained model is directly applied to specific road scenes, the model degradation problem caused by differences in feature distribution has not been effectively solved.

[0027] In view of the above problems of the prior art, this application proposes a method for detecting vehicle driving status, the flow chart of which is shown in the attached figure. Figure 1 As shown, it mainly includes steps S101 to S105, which are detailed as follows:

[0028] Step S101: obtaining original video data of a vehicle traveling by means of a video acquisition device.

[0029] Taking into account that a vehicle may be in various complex environments while driving, such as bad weather, low illumination, etc., and different video acquisition devices have their own advantages in dealing with complex environments, therefore, in an embodiment of the present application, the video acquisition device can be a combination of multiple visual sensors, such as any combination of visible light cameras, infrared thermal imaging cameras, polarized light cameras, and millimeter wave radars. This is because visible light cameras can be used normally under conventional lighting conditions such as sunny days and cloudy days, while infrared thermal imaging cameras have advantages in low illumination environments such as at night and in tunnels. When the confidence level of video motion detection is low, the point cloud data obtained by the millimeter wave radar can be called to perform motion trajectory verification, and so on. Therefore, obtaining the original video data of vehicle driving through the video acquisition device can specifically be to use a visible light camera, an infrared thermal imaging camera and a polarized light camera to synchronously collect visible light, infrared and polarized three-channel video streams, and then perform time domain alignment on the data streams obtained by this channel; dynamically select the main channel based on the environmental visibility. For example, if the visibility is greater than 100m, select the visible light channel, that is, the video stream obtained by the visible light camera; when the visibility is between 50m and 100m, select the polarized light channel, that is, the video stream obtained by the infrared thermal imaging camera; in other cases, select the infrared channel, that is, the video stream obtained by the polarized light camera.

[0030] Step S102: Perform motion feature modeling on each frame group of the original video data to generate a motion feature image containing spatiotemporal information.

[0031] As mentioned above, the existing technology mostly adopts a frame-by-frame processing mode when processing video streams, which cannot effectively capture the continuous spatiotemporal correlation characteristics of vehicle motion, resulting in insufficient accuracy of dynamic detection such as steering recognition and speed estimation. In order to break through the two extreme limitations of single-frame or full-sequence processing in the existing technology and the spatiotemporal information fragmentation problem of single-frame analysis, the present application constructs spatiotemporal correlation features by frame group division, that is, motion feature modeling is performed on each frame group of the original video data. Specifically, as an embodiment of the present application, motion feature modeling for each frame group of the original video data can be achieved through steps S1021 to S1023, which are detailed as follows:

[0032] Step S1021: Divide the original video data into N consecutive frame groups.

[0033] When dividing the original video data into N consecutive frame groups, on the one hand, the number of frame group divisions N essentially defines the length of the time window for spatiotemporal feature extraction, which directly affects the temporal resolution (that is, the smaller N is, the shorter the window time span, and the stronger the ability to capture fast actions (such as emergency braking)) and motion continuity characterization (that is, the larger N is, the more it can characterize long-term motion trends); on the other hand, the number of frame group divisions N is closely related to the complexity of the calculation, that is, the time consumption of processing a single frame group is sublinearly related to N, and the memory consumption is linearly related to N. Therefore, in an embodiment of the present application, it is necessary to scientifically and reasonably determine the number of frame group divisions N. As an embodiment of the present application, the method for determining the number of frame group divisions N can be: testing the classification accuracy under different N values ​​through cross-validation; selecting an N value that takes into account both temporal resolution and computational efficiency based on the inflection point of the accuracy curve. Specifically, the method for determining the number of frame group divisions N includes the following steps S1 to S5:

[0034] Step S1: Divide the training data set into a training subset and a validation subset, for example, the training subset and the validation subset are equally divided.

[0035] Step S2: Traverse candidate N values ​​(e.g., N=5, 10, 15, ..., 30), and for each N value, perform the following: calculate the motion history image of each frame group under the continuous grouping of the N value; train the multi-stream fusion classification model; and record the classification accuracy of the verification set.

[0036] Step S3: Draw a curve showing the relationship between the N value and the classification accuracy.

[0037] Step S4: Calculate the second-order derivative of the relationship curve between the N value and the classification accuracy, and locate the inflection point, that is, the point where the slope changes the most.

[0038] Step S5: Select the N value corresponding to the inflection point as the optimal parameter.

[0039] Step S1022: calculating a motion history image for each of the N consecutive frame groups, wherein the pixel brightness in the motion history image is positively correlated with the time when the motion occurred.

[0040] The motion history image is calculated for each of the N consecutive frame groups as follows: for a continuous frame sequence in each of the N consecutive frame groups, a differential image of adjacent frames is calculated; and the differential image sequence is time-attenuated and superimposed to generate a motion history image, where the pixel brightness I(x, y, t) satisfies:

[0041] I(x,y,t)=max(τ-(tt last ),0)

[0042] Where τ is the time decay coefficient, t lastis the frame number of the last movement at the current position. From this, we can see that the pixel brightness in the motion history image is positively correlated with the time of motion, that is, the more recent the motion, the brighter it appears in the image.

[0043] Step S1023: Segment the motion history image into multiple sub-regions, and generate gradient histogram features based on the gradient magnitude and direction distribution of each sub-region.

[0044] On the one hand, directly using the raw pixels of the motion history image introduces illumination sensitivity issues. The gradient histogram eliminates the influence of illumination intensity differences on features by calculating the gradient direction of subregions. On the other hand, using only the motion history image without segmenting the subregions, global gradient statistics can blur local motion direction differences, making it impossible to distinguish between vehicle turning and pedestrian crossing motion patterns. Therefore, after calculating the motion history image, the motion history image can be segmented into multiple subregions, and gradient histogram features can be generated based on the gradient magnitude and direction distribution of each subregion. Specifically, the gradient histogram features can be generated by applying the Sobel operator in the horizontal and vertical directions to each of the multiple subregions, calculating the gradient magnitude and direction of each pixel; discretizing the gradient direction into four intervals of 0°, 45°, 90°, and 135°, and summing the gradient magnitude within each interval. Through these operations, the temporal encoding capability of the motion history image significantly reduces the error in vehicle steering angular velocity estimation by approximately 1 / 3, and the gradient direction statistics also improve the accuracy of vehicle motion direction recognition in rainy and foggy weather.

[0045] If the generated motion feature image is a grayscale image, a channel expansion operation can also be performed, that is, the grayscale image is repeated several times (for example, 3 times) in the channel dimension; a pseudo RGB image is generated and input into the pre-trained network.

[0046] Step S103: construct a multi-stream fusion classification model, wherein the multi-stream fusion classification model includes a spatial feature extraction module, a spatiotemporal feature extraction module, and a temporal feature extraction module based on a pre-trained deep convolutional network.

[0047] In an embodiment of the present application, the multi-stream fusion classification model includes a spatial feature extraction module, a spatiotemporal feature extraction module and a temporal feature extraction module based on a pre-trained deep convolutional network, which can extract spatial streams, spatiotemporal streams and temporal streams respectively for the original video data. Specifically, the spatial stream is the static spatial features extracted by the above-mentioned spatial feature extraction module after a single frame image in the original video (such as a randomly sampled sample frame) is input into the multi-stream fusion classification model. For example, the shape and position of the vehicle are extracted. The spatiotemporal stream is the spatiotemporal motion features extracted by the above-mentioned spatiotemporal feature extraction module after the block motion history image is input into the multi-stream fusion classification model. For example, the vehicle speed and direction are extracted. The temporal stream is the temporal dynamic features extracted by the above-mentioned temporal feature extraction module after the block energy image (such as the frequency domain features of the differential sequence) is input into the multi-stream fusion classification model. For example, road condition changes, vehicle acceleration, etc. It should be noted that the multi-stream fusion classification model outputs three independent streams consisting of the above-mentioned spatial stream, spatiotemporal stream, and time stream. Each stream processes different input data based on the pre-trained multi-stream fusion classification model. Each stream outputs its own independently calculated classification probability value (that is, the category probability distribution generated by the softmax function for each stream), rather than directly outputting the final detection result.

[0048] Step S104: input the motion feature image and the original video data into the multi-stream fusion classification model, and jointly train the multi-stream fusion classification model through a transfer learning strategy to obtain a trained multi-stream fusion classification model.

[0049] When processing video streams, the existing technology cannot take into account both static appearance and dynamic behavior features in a single feature stream, resulting in a surge in the misjudgment rate in complex scenes. In the embodiment of the present application, a multi-stream fusion classification model including spatial, spatiotemporal, and temporal feature extraction modules is constructed through step S103. Furthermore, by parallel processing the original video data of the vehicle and the motion feature image containing spatiotemporal information, the complementary enhancement of spatial static features and motion dynamic features can be achieved, showing stronger environmental adaptability in complex road scenes.

[0050] As an embodiment of the present application, the pre-trained deep convolutional network can be a VGG-16 architecture, and the transfer learning strategy can be: freezing the weight parameters of the first four convolutional blocks of the network; unfreezing the fifth convolutional block and subsequent fully connected layers for fine-tuning training. The joint training of the multi-stream fusion classification model through the transfer learning strategy can be: inputting a single-frame image randomly sampled from the original video data, and extracting the spatial stream by the spatial feature extraction module based on the VGG-16 architecture; inputting the block motion history image, and extracting the spatiotemporal stream by the spatiotemporal feature extraction module based on the VGG-16 architecture; inputting the block energy image, and extracting the temporal stream by the temporal feature extraction module based on the VGG-16 architecture. In the above embodiment, for the motion history image, before input, each sub-block thereof can be normalized for illumination and the CLAHE algorithm can be used to enhance the details of the low-contrast area. The block energy image is obtained by calculating the differential sequence of adjacent frames of the original video data, and then performing Fourier transform on the differential sequence to extract the frequency domain features. It should be noted that, on the one hand, since the vehicle motion frequency is usually concentrated in the range of 0-10Hz (for example, the steering angular velocity is ≤10Hz), retaining this frequency band can improve the signal-to-noise ratio of the motion feature. If it is extended to 20Hz, the high-frequency noise of the road vibration introduced will reduce the feature separability. On the other hand, compared with the Chebyshev filter, the Butterworth filter has a passband flatness better than 0.5dB, avoiding motion trajectory distortion caused by phase distortion. Therefore, in the above embodiment, when Fourier transforming the differential sequence to extract frequency domain features, the low-frequency components (for example, the 0-10Hz components) can be retained, while the high-frequency noise can be subjected to Butterworth low-pass filtering.

[0051] In addition, when fine-tuning pre-trained deep convolutional networks, regularization can be inserted between fully connected layers. This includes setting a dropout layer after the first fully connected layer with a dropout rate of 0.5, using a ReLU activation function and L2 regularization constraints. It should be noted that the dropout rate of the dropout layer is not fixed after setting and can be dynamically adjusted according to the training stage. For example, a higher dropout rate (e.g., 0.7) can be used in the initial training stage; the lower limit can be linearly reduced by 0.05 to 0.3 every 10 epochs.

[0052] Step S105: performing weighted fusion based on the classification probability values ​​output by each independent stream in the trained multi-stream fusion classification model to generate a final detection result of the vehicle driving state.

[0053] In order to avoid the problem in the existing technology that a single feature stream dominates the decision-making process, making it impossible for the advantages of each feature stream to complement each other, thereby causing noise interference to be directly transmitted to the final output, this application performs weighted fusion based on the classification probability values ​​output by each independent stream in the trained multi-stream fusion classification model to generate the final detection result of the vehicle's driving status. This mutual verification mechanism through multi-dimensional decision-making significantly suppresses the systemic risks brought about by single-view feature misjudgment, and improves the robustness of vehicle driving status detection in complex environments (such as bad weather).

[0054] In the above embodiment, when weighted fusion is performed based on the classification probability values ​​output by each independent stream in the trained multi-stream fusion classification model, the weight of the weighted fusion can adopt the following dynamic weight allocation strategy, that is, real-time monitoring of the classification confidence of each stream; a first weight is assigned to the mainstream stream with a confidence higher than a threshold in each independent stream, and the remaining independent streams are evenly divided into a second weight, wherein the first weight is greater than the second weight, for example, a weight of 0.7 is assigned to the mainstream stream with a confidence higher than a threshold in the independent stream, and the remaining two independent streams are evenly divided into a weight of 0.3 (meaning that each independent stream is assigned a weight of 0.15). It should be noted that the reason why the above-mentioned independent streams are assigned different weights is, on the one hand, to take into account that fixed weight fusion (such as average weighting) will cause the spatiotemporal stream misjudgment weight to be too high when the lighting changes suddenly. Dynamic allocation allows the system to automatically reduce the spatiotemporal stream weight to below a threshold (such as 0.3) in rainy and foggy weather, avoiding the decision-making dominated by erroneous features; on the other hand, through Monte Carlo simulation, it is found that when the mainstream weight exceeds 0.65, it can suppress the noise stream while retaining the multi-stream verification capability, reducing the false alarm rate by another 4.2% compared with the 0.6 weight scheme. If the first weight and the second weight are allocated in a ratio of 0.8 / 0.2, the risk of system crash when both streams fail at the same time is increased by 3 times. The weight allocation method of the above embodiment can greatly enhance the environmental adaptability and fault tolerance capabilities. For example, in a strong backlight scene, the dynamic reduction of the spatiotemporal stream weight reduces the vehicle positioning error to within 1.5 pixels; when a feature stream hardware fails (such as a spatiotemporal stream sensor is damaged), the 0.3 weight of the remaining stream can still maintain a basic detection accuracy of more than 80%.

[0055] Furthermore, in order to avoid the rigidity of weight distribution and deal with feature ambiguity scenarios, for example, if the confidence of the spatial stream and the spatiotemporal stream are 0.81 and 0.80 respectively (difference 0.01), the weights are still allocated according to the ratio of 0.7 / 0.3, which may amplify the errors caused by small differences. Or, in rainy and foggy weather, the confidence of the spatial feature (vehicle outline) and the temporal feature (motion trajectory) may decay at the same time and the difference is extremely small. Switching to average weighting can integrate multi-dimensional features and improve the detection stability in harsh environments. In the embodiment of the present application, when the confidence difference of the mainstream with a confidence higher than the threshold is less than the preset threshold, its weight is switched to the average of the second weight; the feature recalibration mechanism is started for the streams in each independent stream whose confidence is lower than the preset confidence threshold for three consecutive times. The above embodiment shows that when the mainstream confidence is close, the above dynamic weight distribution strategy can suppress the decision-making dominance of a single independent stream and optimize hardware utilization. The feature recalibration mechanism can ensure the long-term stable contribution of each independent stream, thereby significantly improving the robustness of vehicle status detection in complex scenarios.

[0056] From the above attached Figure 1 It can be seen from the example vehicle driving state detection method that, on the one hand, by modeling the motion features of each frame group of the acquired original video data, the problem of spatiotemporal information separation in single-frame analysis is broken through, and the short-term trajectory features and long-term trend features of vehicle motion are effectively integrated, which significantly improves the accuracy of steering state recognition and speed estimation; on the other hand, a multi-stream fusion classification model is constructed, including a spatial feature extraction module based on a pre-trained deep convolutional network, a spatiotemporal feature extraction module and a temporal feature extraction module, and the motion feature image and the original video data are input into the multi-stream fusion classification model, which realizes the complementary enhancement of spatial static features and motion dynamic features, and shows stronger environmental adaptability in complex road scenes; thirdly, the video frame group segmentation strategy is combined with the transfer learning strategy to jointly train the multi-stream fusion classification model, which can also significantly reduce the redundant computational amount of video processing.

[0057] Please see the attached Figure 2 , is a vehicle driving state detection device provided in an embodiment of the present application, which may include an acquisition module 201, a first generation module 202, a construction module 203, a training module 204, and a second generation module 205, as detailed below:

[0058] An acquisition module 201 is configured to acquire original video data of a vehicle traveling via a video acquisition device;

[0059] The first generating module 202 is configured to perform motion feature modeling on each frame group of the original video data to generate a motion feature image containing spatiotemporal information;

[0060] A construction module 203 is used to construct a multi-stream fusion classification model, wherein the multi-stream fusion classification model includes a spatial feature extraction module, a spatiotemporal feature extraction module, and a temporal feature extraction module based on a pre-trained deep convolutional network;

[0061] A training module 204 is configured to input the motion feature image and the original video data into a multi-stream fusion classification model, and jointly train the multi-stream fusion classification model using a transfer learning strategy to obtain a trained multi-stream fusion classification model;

[0062] The second generation module 205 is configured to perform weighted fusion based on the classification probability values ​​output by each independent stream in the trained multi-stream fusion classification model to generate a final detection result of the vehicle driving state.

[0063] From the above attached Figure 2 It can be seen from the example vehicle driving state detection device that, on the one hand, by modeling the motion features of each frame group of the acquired original video data, the problem of spatiotemporal information separation in single-frame analysis is broken through, and the short-term trajectory features and long-term trend features of vehicle motion are effectively integrated, which significantly improves the accuracy of steering state recognition and speed estimation; on the other hand, a multi-stream fusion classification model is constructed, including a spatial feature extraction module, a spatiotemporal feature extraction module and a temporal feature extraction module based on a pre-trained deep convolutional network, and the motion feature image and the original video data are input into the multi-stream fusion classification model, which realizes the complementary enhancement of spatial static features and motion dynamic features, and shows stronger environmental adaptability in complex road scenes; thirdly, the video frame group segmentation strategy is combined with the transfer learning strategy to jointly train the multi-stream fusion classification model, which can also significantly reduce the redundant computational amount of video processing.

[0064] Figure 3 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 3 As shown, the electronic device 3 of this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and capable of running on the processor 30, such as a program of the vehicle driving state detection method. When the processor 30 executes the computer program 32, the steps in the above-mentioned vehicle driving state detection method embodiment are implemented, such as Figure 1 Alternatively, when the processor 30 executes the computer program 32, the functions of the modules / units in the above-mentioned device embodiments are realized, for example Figure 2 The functions of the acquisition module 201, the first generation module 202, the construction module 203, the training module 204 and the second generation module 205 are shown.

[0065] Exemplarily, a computer program 32 for a vehicle driving state detection method primarily includes: acquiring raw video data of a vehicle driving through a video acquisition device; performing motion feature modeling on each frame group of the raw video data to generate a motion feature image containing spatiotemporal information; constructing a multi-stream fusion classification model, wherein the multi-stream fusion classification model includes a spatial feature extraction module, a spatiotemporal feature extraction module, and a temporal feature extraction module based on a pre-trained deep convolutional network; inputting the motion feature image and the raw video data into the multi-stream fusion classification model, jointly training the multi-stream fusion classification model through a transfer learning strategy to obtain a trained multi-stream fusion classification model; and performing weighted fusion based on the classification probability values ​​output by each independent stream in the trained multi-stream fusion classification model to generate a final detection result of the vehicle driving state. The computer program 32 can be divided into one or more modules / units, one or more of which are stored in the memory 31 and executed by the processor 30 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which are used to describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 can be divided into the functions of an acquisition module 201, a first generation module 202, a construction module 203, a training module 204 and a second generation module 205 (modules in the virtual device), and the specific functions of each module are as follows: the acquisition module 201 is used to obtain the original video data of the vehicle driving through the video acquisition device; the first generation module 202 is used to perform motion feature modeling on each frame group of the original video data to generate a motion feature image containing spatiotemporal information; the construction module 203 is used to construct a multi-stream fusion classification model, wherein the multi-stream fusion classification model includes a spatial feature extraction module, a spatiotemporal feature extraction module and a temporal feature extraction module based on a pre-trained deep convolutional network; the training module 204 is used to input the motion feature image and the original video data into the multi-stream fusion classification model, and jointly train the multi-stream fusion classification model through a transfer learning strategy to obtain a trained multi-stream fusion classification model; the second generation module 205 is used to perform weighted fusion based on the classification probability values ​​output by each independent stream in the trained multi-stream fusion classification model to generate the final detection result of the vehicle driving status.

[0066] The electronic device 3 may include but is not limited to a processor 30 and a memory 31. Those skilled in the art will appreciate that Figure 3 It is only an example of electronic device 3 and does not constitute a limitation of electronic device 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0067] The processor 30 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0068] The memory 31 can be an internal storage unit of the electronic device 3, such as a hard drive or memory of the electronic device 3. The memory 31 can also be an external storage device of the electronic device 3, such as a plug-in hard drive, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the electronic device 3. Furthermore, the memory 31 can include both an internal storage unit of the electronic device 3 and an external storage device. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 can also be used to temporarily store data that has been output or is about to be output.

[0069] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0070] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0071] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0072] In the embodiments provided in this application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0073] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0074] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0075] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program of the vehicle driving state detection method can be stored in a storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments, that is, obtaining the original video data of the vehicle driving through the video acquisition device; performing motion feature modeling on each frame group of the original video data to generate a motion feature image containing spatiotemporal information; constructing a multi-stream fusion classification model, wherein the multi-stream fusion classification model includes a spatial feature extraction module, a spatiotemporal feature extraction module and a temporal feature extraction module based on a pre-trained deep convolutional network; inputting the motion feature image and the original video data into the multi-stream fusion classification model, and jointly training the multi-stream fusion classification model through a transfer learning strategy to obtain a trained multi-stream fusion classification model; performing weighted fusion based on the classification probability values ​​output by each independent stream in the trained multi-stream fusion classification model to generate the final detection result of the vehicle driving state. Among them, computer programs include computer program code, which may be in source code form, object code form, executable files, or some intermediate form. Storage media may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, mobile hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunications signals, and software distribution media. It should be noted that the content contained in storage media may be appropriately increased or decreased based on the requirements of legislation and patent practice within a jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, storage media do not include electric carrier signals and telecommunications signals.

[0076] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application. The specific implementation methods described above further explain the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation method of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included in the protection scope of the present invention.

Claims

1. A vehicle driving state detection method, characterized in that: The method comprises: Obtaining original video data of vehicle driving through video acquisition equipment; Performing motion feature modeling on each frame group of the original video data to generate a motion feature image containing spatiotemporal information; Constructing a multi-stream fusion classification model, wherein the multi-stream fusion classification model includes a spatial feature extraction module, a spatiotemporal feature extraction module, and a temporal feature extraction module based on a pre-trained deep convolutional network; Inputting the motion feature image and the original video data into the multi-stream fusion classification model, and jointly training the multi-stream fusion classification model through a transfer learning strategy to obtain a trained multi-stream fusion classification model; The classification probability values ​​output by each independent stream in the trained multi-stream fusion classification model are weightedly fused to generate a final detection result of the vehicle driving state.

2. The vehicle driving state detection method according to claim 1, wherein: The performing motion feature modeling on each frame group of the original video data to generate a motion feature image containing spatiotemporal information includes: Splitting the original video data into N consecutive frame groups; calculating a motion history image for each of the N consecutive frame groups, wherein pixel brightness in the motion history image is positively correlated with a time when the motion occurs; The motion history image is divided into a plurality of sub-regions, and a gradient histogram feature is generated based on the gradient magnitude and direction distribution of each sub-region.

3. The vehicle driving state detection method according to claim 2, wherein: The method for determining N when dividing the original video data into N consecutive frame groups includes: Test the classification accuracy under different N values ​​through cross validation; The N value that takes into account both time resolution and computational efficiency is selected based on the inflection point of the accuracy curve.

4. The vehicle driving state detection method according to claim 1, wherein: The pre-trained deep convolutional network is a VGG-16 architecture, and the transfer learning strategy includes: Freeze the weight parameters of the first four convolutional blocks of the network; Unfreeze the fifth convolutional block and subsequent fully connected layers for fine-tuning training.

5. The vehicle driving state detection method according to claim 2, wherein: Generating a gradient histogram feature based on the gradient magnitude and direction distribution of each sub-region includes: Applying Sobel operators in the horizontal and vertical directions to each of the multiple sub-regions to calculate the gradient magnitude and direction of each pixel; The gradient direction is discretized into four intervals of 0°, 45°, 90°, and 135°, and the sum of the gradient amplitudes in each interval is counted.

6. The vehicle driving state detection method according to claim 1, wherein: The weight of the weighted fusion adopts the following dynamic weight allocation strategy: Real-time monitoring of the confidence level of each flow classification; A first weight is allocated to the mainstream stream with a confidence level higher than a threshold among the independent streams, and a second weight is allocated equally to the remaining independent streams, where the first weight is greater than the second weight.

7. The vehicle driving state detection method according to claim 6, characterized in that: The dynamic weight allocation strategy further includes: When the confidence difference of the mainstream is less than a preset threshold, switching its weight to the average of the second weights; A feature recalibration mechanism is initiated for a flow in each of the independent flows whose confidence is lower than a preset confidence threshold for three consecutive times.

8. A vehicle driving state detection device, characterized in that: The device comprises: An acquisition module is used to acquire original video data of the vehicle driving through a video acquisition device; A first generating module is used to perform motion feature modeling on each frame group of the original video data to generate a motion feature image containing spatiotemporal information; A construction module is used to construct a multi-stream fusion classification model, wherein the multi-stream fusion classification model includes a spatial feature extraction module, a spatiotemporal feature extraction module, and a temporal feature extraction module based on a pre-trained deep convolutional network; a training module, configured to input the motion feature image and the original video data into the multi-stream fusion classification model, and jointly train the multi-stream fusion classification model through a transfer learning strategy to obtain a trained multi-stream fusion classification model; The second generation module is used to perform weighted fusion based on the classification probability values ​​output by each independent stream in the trained multi-stream fusion classification model to generate a final detection result of the vehicle driving state.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.