Accident Detection Method, Device, Electronic Device and Storage Medium

By combining global and local characteristics for accident detection, the problem of low detection accuracy and reliability in the prior art is solved, and more accurate and reliable traffic accident detection is achieved.

CN114677618BActive Publication Date: 2025-06-10ANHUI IFLYTEK INTELLIGENT SYST +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210194552.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-01
Publication Date
2025-06-10
Estimated Expiration
2042-03-01

AI Technical Summary

Technical Problem

In the prior art, accident detection is carried out based on the relevant information of the target in the video, and the detection accuracy and reliability are low.

Method used

By determining the image frame sequence of the video to be detected, three-dimensional feature extraction is performed based on the global extraction network to obtain global features; based on the local extraction network, local features are determined using the detection target and target position in the image frame sequence; then, based on the fusion classification network, accident detection is performed in combination with global and local features.

Benefits of technology

It improves the accuracy and reliability of accident detection, and can accurately complete accident detection when the target changes drastically or the scene changes are not obvious, ensuring timely monitoring of traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114677618B_ABST
    Figure CN114677618B_ABST
Patent Text Reader

Abstract

The present invention provides an accident detection method, apparatus, electronic device and storage medium. The method includes: determining an image frame sequence of a video to be detected; performing three-dimensional feature extraction on the image frame sequence based on a global extraction network to obtain global features of the video to be detected; determining local features of the video to be detected based on a local extraction network by applying detection targets and target positions of each frame image in the image frame sequence; and determining an accident detection result of the video to be detected based on a fusion classification network by applying the global features and the local features. The method, apparatus, electronic device and storage medium provided by the present invention perform accident detection by combining the global features and the local features of the video to be detected. Whether it is the situation where target detection fails or the target is lost due to drastic changes in the target, or the situation where the scene change is not obvious, the accident detection can be accurately and reliably completed, so as to ensure that traffic accidents can be monitored in a timely manner and facilitate the timeliness of accident investigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to an accident detection method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development and application of video surveillance technology and the large-scale installation of road surveillance cameras, conditions are provided for monitoring road traffic conditions and promptly handling traffic accidents.

[0003] The traditional method for investigating traffic accidents is manual monitoring, which requires personnel to be on duty throughout the day and watch surveillance videos. This not only consumes a large amount of manpower but is also affected by uncontrollable factors such as the human eye's resolution ability and fatigue level, resulting in low reliability.

[0004] With the wide application of deep learning in computer vision tasks, video-based traffic accident detection methods have emerged. Currently, most such solutions obtain information on the appearance and motion of targets from videos, extract features therefrom, and perform accident detection based on these features. However, directly extracting features from the appearance and motion information of targets will result in the loss of some information, affecting the detection accuracy. Moreover, the extraction of features and the detection of accidents overly rely on the accuracy of object detection and tracking, which will also affect the reliability of accident detection. Summary of the Invention

[0005] The present invention provides an accident detection method, apparatus, electronic device, and storage medium to solve the problems of low detection accuracy and reliability in accident detection based on relevant information of targets in videos in the prior art.

[0006] The present invention provides an accident detection method, including:

[0007] Determine the image frame sequence of the video to be detected;

[0008] Based on a global extraction network, perform three-dimensional feature extraction on the image frame sequence to obtain the global features of the video to be detected;

[0009] Based on a local extraction network, use the detection targets and target positions of each frame image in the image frame sequence to determine the local features of the video to be detected;

[0010] Based on a fusion classification network, use the global features and the local features to determine the accident detection result of the video to be detected.

[0011] According to the accident detection method provided by the present invention, the step of performing three-dimensional feature extraction on the image frame sequence based on a global extraction network to obtain the global features of the video to be detected includes:

[0012] Based on the multi-layer 3D convolutional network in the global extraction network, perform multi-layer 3D convolution on the image frame sequence to obtain a first convolutional feature and a second convolutional feature, where the first convolutional feature is obtained by convolution before the second convolutional feature;

[0013] Based on the attention network in the global extraction network, apply the first convolutional feature to determine the attention weight of the second convolutional feature, and apply the attention weight to weight the second convolutional feature to obtain the global feature.

[0014] According to an accident detection method provided by the present invention, the first convolutional feature and the second convolutional feature are respectively the convolutional features output by the penultimate layer and the last layer of the multi-layer 3D convolution;

[0015] The applying the first convolutional feature to determine the attention weight of the second convolutional feature includes:

[0016] Perform single-layer 3D convolution on the first convolutional feature to obtain a third convolutional feature with the same dimension as the second convolutional feature;

[0017] Based on the third convolutional feature, determine the attention weight.

[0018] According to an accident detection method provided by the present invention, the determining the local feature of the video to be detected based on the local extraction network by applying the detection target and target position of each frame image in the image frame sequence includes:

[0019] Based on the target detection network in the local extraction network, determine the target feature and target position of the detection target in each frame image of the image frame sequence;

[0020] Based on the spatio-temporal extraction network in the local extraction network, perform spatio-temporal information extraction on the target feature map of each frame image to obtain the local feature, where the target feature map is determined based on the target feature and target position of the detection target in the corresponding image.

[0021] According to an accident detection method provided by the present invention, the performing spatio-temporal information extraction on the target feature map of each frame image based on the spatio-temporal extraction network in the local extraction network to obtain the local feature includes:

[0022] Based on the graph convolutional network in the spatio-temporal extraction network, perform spatial information extraction on the target feature map of each frame image to obtain the target spatial relationship of each frame image;

[0023] Based on the temporal extraction network in the spatio-temporal extraction network, perform temporal feature extraction on the target spatial relationship of each frame image to obtain the local feature of the video to be detected.

[0024] An accident detection method provided by the present invention, based on the fusion classification network, applies the global feature and the local feature to determine the accident detection result of the video to be detected, including:

[0025] Based on the fusion classification network, fuse the global feature and the local feature, and perform context extraction based on the fused feature to obtain a context feature, and apply the context feature to perform accident classification to determine the accident detection result.

[0026] An accident detection method provided by the present invention, the global extraction network, the local extraction network, and the fusion classification network are determined based on the following steps:

[0027] Construct an initial detection network based on an initial global extraction network, an initial local extraction network, and an initial fusion classification network;

[0028] Based on the first sample video with accident labels, train the initial detection network, and determine the global extraction network, the local extraction network, and the fusion classification network based on the trained initial detection network.

[0029] An accident detection method provided by the present invention, the initial global extraction network is obtained by jointly training with a global classification network based on a second sample video with accident labels, and the global classification network is used to perform accident detection based on the global feature.

[0030] The present invention also provides an accident detection device, including:

[0031] A sequence determination unit, configured to determine the image frame sequence of the video to be detected;

[0032] A global extraction unit, configured to perform three-dimensional feature extraction on the image frame sequence based on the global extraction network to obtain the global feature of the video to be detected;

[0033] A local extraction unit, configured to determine the local feature of the video to be detected based on the local extraction network by applying the detection target and the target position of each frame image in the image frame sequence;

[0034] A fusion classification unit, configured to determine the accident detection result of the video to be detected based on the fusion classification network by applying the global feature and the local feature.

[0035] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the steps of any one of the above accident detection methods are implemented.

[0036] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned accident detection methods are implemented.

[0037] The accident detection method, device, electronic device and storage medium provided by the present invention perform accident detection by combining the global features and local features of the video to be detected. Whether it is the situation where the target detection fails or the target is lost due to drastic changes in the target, or the situation where the scene change is not obvious, the accident detection can be accurately and reliably completed, so as to ensure that traffic accidents can be monitored in time and facilitate the timeliness of accident investigation. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly describe the drawings required for the implementation of the embodiments or the prior art descriptions. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 is a schematic flowchart of the accident detection method provided by the present invention;

[0040] Figure 2 is a schematic flowchart of step 120 in the accident detection method provided by the present invention;

[0041] Figure 3 is a schematic structural diagram of the global extraction network provided by the present invention;

[0042] Figure 4 is a schematic flowchart of step 130 in the accident detection method provided by the present invention;

[0043] Figure 5 is a schematic structural diagram of the local extraction network provided by the present invention;

[0044] Figure 6 is a schematic structural diagram of the first-stage training provided by the present invention;

[0045] Figure 7 is a schematic flowchart of the accident detection method provided by the present invention;

[0046] Figure 8 is a schematic structural diagram of the accident detection device provided by the present invention;

[0047] Figure 9 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0049] With the wide application of deep learning in computer vision tasks, video-based traffic accident detection methods have emerged.

[0050] Currently, most of such solutions obtain information on the appearance and motion of targets from videos, extract features therefrom, and perform accident detection based on these features. The features extracted from the appearance information of the target are used to determine whether there are obvious changes in vehicles and pedestrians compared to the normal state, such as whether there is vehicle rollover or pedestrian fall; the features extracted from the running information of the target are used to determine whether there are situations such as intersection of vehicle and pedestrian trajectories, sudden changes in speed and angular velocity, etc.

[0051] Directly extracting features from the appearance and running information of the target, overly focusing on local information, will result in loss of some information, ignoring the understanding of the entire traffic scene, and thus cannot well identify non-accident scenes such as traffic jams, affecting the accuracy of accident detection. In addition, the method of simultaneously capturing the fusion of appearance and motion in features requires modeling the information association in the target time series through object detection and object tracking networks, only focusing on the local information of the target, overly relying on the accuracy of object detection and tracking, and ignoring the global information of the video, which will also affect the reliability of accident detection.

[0052] In view of the above problems, an embodiment of the present invention provides an accident detection method, Figure 1 which is a schematic flowchart of the accident detection method provided by the present invention, as Figure 1 shown, the method includes:

[0053] Step 110, determining an image frame sequence of the video to be detected.

[0054] Specifically, the video to be recognized is the video for which accident detection needs to be performed. Here, the video to be recognized can be a pre-shot and stored video or a real-time acquired video stream. The embodiment of the present invention does not make specific limitations thereto. The image frame sequence is obtained by sampling the video to be recognized. The image frame sequence contains multiple frames of images. Each frame of image is derived from the video to be recognized, and the multiple frames of images are arranged in the time order in the video to be recognized, thereby forming an image frame sequence.

[0055] It should be noted that when collecting the video to be recognized, it is usually uniformly and sequentially collected based on the total number of frames of the video to be recognized, and the time intervals between each obtained frame image are equal. Alternatively, it can also be directly combining the frame images in the video to be recognized in sequence to form an image frame sequence. For example, every 16 frames can be used as a group of image frame sequences.

[0056] Step 120: Based on the global extraction network, perform three-dimensional feature extraction on the image frame sequence to obtain the global features of the video to be detected.

[0057] Here, the global extraction network is a pre-trained neural network for extracting global features from an image frame sequence. For example, the image frame sequence can be input into the global extraction network, and the global extraction network performs three-dimensional feature extraction on the input image frame sequence and takes the extracted three-dimensional features as the global features of the video to be detected.

[0058] Among them, performing three-dimensional feature extraction on the image frame sequence means extracting features from the entire image frame sequence in three dimensions, namely the two dimensions of the image frame itself and the dimension corresponding to the time sequence between the frame images in the image frame sequence. Here, the three-dimensional feature extraction can be implemented by 3D-CNN (Convolutional Neural Networks). Performing three-dimensional feature extraction on the image frame sequence through the global extraction network can ensure that the obtained global features can cover the information of each frame image and the time sequence between each frame image in the image frame sequence, and reflect the information of the video to be detected in the traffic scene as a whole.

[0059] Step 130: Based on the local extraction network, apply the detection targets and target positions of each frame image in the image frame sequence to determine the local features of the video to be detected.

[0060] Here, the local extraction network is a pre-trained neural network for extracting local target features from an image frame sequence. For example, the image frame sequence can be input into the local extraction network, and the local extraction network performs target detection on each frame image in the image frame sequence to determine the detection targets such as vehicles and pedestrians included in each frame image and their target positions. Based on this, the local features of the video to be detected are determined based on the positional relationships and positional changes of the detection targets in each frame image; alternatively, it can also perform target detection on each frame image in the image frame sequence first, and then input the detection targets and their target positions included in each frame image into the local extraction network for local feature extraction.

[0061] The local features extracted therefrom can reflect the information of each detection target in the video to be detected in terms of space and time, and represent the information of the video to be detected on detection targets such as vehicles and pedestrians from the level focusing on local targets.

[0062] It should be noted that step 120 and step 130 can be executed synchronously, or step 120 can be executed before or after step 130. The embodiments of the present invention do not make specific limitations on this.

[0063] Step 140: Based on the fusion classification network, apply the global feature and the local feature to determine the accident detection result of the video to be detected.

[0064] Here, the fusion classification network is a neural network that has been pre-trained for accident detection by combining global features and local features. For example, the global feature and the local feature obtained in step 120 and step 130 can be input into the fusion classification network respectively. The fusion classification network fuses the global feature and the local feature and then conducts accident classification to obtain the accident detection result. Or the global feature and the local feature can also be input into the fusion classification network, and the fusion classification network conducts accident classification on the global feature and the local feature respectively to obtain the classification result based on the global feature and the classification result based on the local feature, and integrates them based on this to obtain the accident detection result. The accident detection result here is used to reflect whether a traffic accident has occurred in the video to be detected. If a traffic accident has occurred, the accident detection result can further include the type or severity of the traffic accident, etc., or can further include the start frame and end frame of the accident.

[0065] In this process, for the accident detection of the video to be detected, features at two levels of global features and local features are combined. The application of global features can fully mine the features of the traffic scene, but it may be difficult to capture and identify small changes in the traffic scene. The application of local features can specifically give the features of the detection targets directly related to traffic accidents, but it overly depends on the progress of target detection and tracking. Missed detection or target loss will cause the accident detection to fail. The combination of the two can exactly make up for the deficiencies of a single type of feature in accident detection, and help improve the reliability and accuracy of accident detection.

[0066] The method provided by the embodiments of the present invention conducts accident detection by jointly using the global feature and the local feature of the video to be detected. Whether it is for the situation where target detection fails or the target is lost due to drastic changes in the target, or for the situation where the scene change is not obvious, it can accurately and reliably complete the accident detection, so as to ensure that traffic accidents can be monitored in time and facilitate the timeliness of accident investigation.

[0067] Based on the previous embodiment,Figure 2 It is a schematic flowchart of step 120 in the accident detection method provided by the present invention. As Figure 2 shown, step 120 includes:

[0068] Step 121: Based on the multi-layer 3D convolutional network in the global extraction network, perform multi-layer 3D convolution on the image frame sequence to obtain a first convolutional feature and a second convolutional feature, where the first convolutional feature is obtained by convolution before the second convolutional feature;

[0069] Step 122: Based on the attention network in the global extraction network, apply the first convolutional feature to determine the attention weight of the second convolutional feature, and apply the attention weight to weight the second convolutional feature to obtain the global feature.

[0070] Specifically, the global extraction network may include a multi-layer 3D convolutional network and an attention network. The multi-layer 3D convolutional network can be embodied as a plurality of cascaded 3D convolutional layers, and the output of the previous 3D convolutional layer is the input of the subsequent 3D convolutional layer.

[0071] In step 121, through the multi-layer 3D convolutional network, the image frame sequence can be used as a whole to perform layer-by-layer 3D convolutional feature extraction, so as to obtain a first convolutional feature and a second convolutional feature. Since each cascaded 3D convolutional layer in the multi-layer 3D convolutional network performs convolution layer by layer, there is also a difference in the order of the convolutional features output by the 3D convolutional layer. The first convolutional feature is obtained by convolution before the second convolutional feature, that is, the multi-layer 3D convolutional network first performs 3D convolution on the image frame sequence to obtain the first convolutional feature, and then performs 3D convolution on the first convolutional feature to obtain the second convolutional feature. For example, the multi-layer 3D convolutional network may include 5 cascaded 3D convolutional layers. The first convolutional feature may be obtained by the third layer convolution, and the second convolutional feature may be obtained by the fifth layer convolution.

[0072] Considering that the convolutional features obtained only by applying the multi-layer 3D convolutional network are relatively rough, which may contain a lot of useless background information and may cause certain interference to accident detection. In step 122 of the embodiment of the present invention, the attention network in the global extraction network is applied to improve the attention to the information related to traffic accidents in the overall information of the image frame sequence by determining the attention weight and then weighting, so as to obtain the global feature that can highlight the information related to traffic accidents while reflecting the global scene of the image frame sequence.

[0073] The method provided by the embodiment of the present invention introduces an attention mechanism in the process of global feature extraction, so that the global feature can highlight the information related to traffic accidents while reflecting the global scene of the image frame sequence, which helps to filter out the useless information carried in the features obtained by direct convolution and improve the reliability of accident detection.

[0074] Based on any of the above embodiments, the first convolutional feature and the second convolutional feature are respectively the convolutional features output by the penultimate layer and the last layer of the multi-layer three-dimensional convolution.

[0075] For example, the multi-layer three-dimensional convolution network may include 5 cascaded three-dimensional convolution layers. The first convolutional feature may be obtained by the fourth-layer convolution, and the second convolutional feature may be obtained by the fifth-layer convolution. Another example is that the multi-layer three-dimensional convolution network may include 6 cascaded three-dimensional convolution layers. The first convolutional feature may be obtained by the fifth-layer convolution, and the second convolutional feature may be obtained by the sixth-layer convolution.

[0076] In step 122, applying the first convolutional feature to determine the attention weight of the second convolutional feature includes:

[0077] Performing a single-layer three-dimensional convolution on the first convolutional feature to obtain a third convolutional feature with the same dimension as the second convolutional feature;

[0078] Determining the attention weight based on the third convolutional feature.

[0079] Specifically, since the first convolutional feature is obtained by convolution prior to the second convolutional feature, the feature dimension of the first convolutional feature is greater than that of the second convolutional feature. Therefore, it is necessary to perform an additional single-layer three-dimensional convolution on the first convolutional feature to regularize the feature dimension of the first convolutional feature into the same feature dimension as the second convolutional feature, thereby obtaining the third convolutional feature.

[0080] On this basis, the attention weight can be obtained by performing a 1*1 convolution on the third convolutional feature or performing a self-attention transformation on the third convolutional feature.

[0081] For example, Figure 3 is a schematic structural diagram of the global extraction network provided by the present invention. As Figure 3 shown, the image frame sequence includes images from time T to time T+t. The image frame sequence obtains the first convolutional feature and the second convolutional feature through a multi-layer three-dimensional convolution network (3D-CNN). Figure 3 The first convolutional feature in is output by the fourth three-dimensional convolution layer in the multi-layer three-dimensional convolution network, denoted as f4, and the second convolutional feature is output by the fifth three-dimensional convolution layer in the multi-layer three-dimensional convolution network, denoted as f5. The first convolutional feature f4 undergoes a spatio-temporal attention transformation to obtain the attention weight W. On this basis, after the second convolutional feature f5 is weighted by the attention weight W, it is added to the original second convolutional feature f5 as the final global feature.

[0082] Among them, the spatiotemporal attention conversion is to perform a three-dimensional convolution on the first convolution feature f4 to obtain a feature of the same dimension as the second convolution feature f5, and then reduce the dimension through a 1*1 convolution to obtain the attention weight W.

[0083] In the above process, the global feature can be recorded as f5', which is specifically expressed as the following formula:

[0084] f5'=f5+W*f5

[0085] Where * represents the bitwise product operation.

[0086] Based on any of the above embodiments, Figure 4 is a flow chart of step 130 in the accident detection method provided by the present invention, such as Figure 4 Said step 130 comprises:

[0087] Step 131, determining target features and target positions of detected targets in each frame image of the image frame sequence based on the target detection network in the local extraction network;

[0088] Step 132, based on the spatiotemporal extraction network in the local extraction network, extract spatiotemporal information from the target feature map of each frame image to obtain the local features, wherein the target feature map is determined based on the target features and target position of the detected target in the corresponding image.

[0089] Specifically, the local extraction network can include a target detection network and a spatiotemporal extraction network.

[0090] Wherein, the target detection network is used to realize the target detection and positioning of the input image. The target detection network here can be realized by a single-stage target detection method, or by a two-stage target detection method such as faster-rcn, which is not specifically limited in the embodiment of the present invention. In step 131, the image frame sequence can be input into the target detection network, and the target detection network detects and locates the targets such as vehicles and pedestrians contained in each frame image in the image frame sequence, so as to obtain the target features and target positions of the detected targets in each frame image. Wherein, for any detected target in any frame image, the target position of the detected target is the coordinate of the minimum bounding box of the target in the image, and the target feature of the detected target can be the corresponding feature in the image feature of the target area delineated based on the target position. The image feature here can be the feature of the image extracted during the target detection process. For example, when applying faster-rcn for target detection, the image feature can be the low-dimensional feature extracted by the fully connected layer (FC layer) in faster-rcn.

[0091] Considering that the occurrence of a traffic accident may be reflected in only one object in the image, and the object involved in the traffic accident usually has other objects with which it interacts. Therefore, after detecting the detection objects included in each frame of the image, an object feature map of each frame of the image can be constructed based on the object features and object positions of the detection objects. Here, in the object feature map, the object features of each detection object are used as nodes, the distances between the detection objects are calculated based on the respective object positions of the detection objects, and the weights of the connection edges between the nodes corresponding to the detection objects are determined based on the distances between the detection objects. The obtained object feature map G can be expressed as:

[0092] G = (V, E)

[0093] Wherein, V represents the object features of the detection objects extracted from the image, and E represents the distances between the detection objects.

[0094] The spatio-temporal extraction network can be used to extract features from the input object feature map in two dimensions of temporal information and spatial information, so as to obtain local features that can not only reflect the spatial relationships of the detection objects in the video to be detected, but also reflect the change information of the detection objects in each frame of the video to be detected in terms of time sequence.

[0095] Based on any of the above embodiments, step 132 includes:

[0096] Based on the graph convolutional network in the spatio-temporal extraction network, extract spatial information from the object feature map of each frame of the image to obtain the object spatial relationships of each frame of the image;

[0097] Based on the temporal extraction network in the spatio-temporal extraction network, extract temporal features from the object spatial relationships of each frame of the image to obtain the local features of the video to be detected.

[0098] Specifically, the spatio-temporal extraction network may include a graph convolutional network and a temporal extraction network. The graph convolutional network is used to realize the aggregation of spatial information of detection objects, and the temporal extraction network is used to realize the aggregation of temporal information of detection objects.

[0099] Among them, the graph convolutional network can be used to extract features from the input graph. Therefore, the object feature map of each frame of the image can be input into the graph convolutional network, and the graph convolutional network extracts the object features of each node reflected in the object feature map and the distances of the connection edges between the nodes, so as to aggregate the spatial information between the detection objects in the video to be detected in the traffic scene and obtain the object spatial relationships of each frame of the image.

[0100] The temporal extraction network can be used to aggregate the temporal relationships between the features of each input. Specifically, the target spatial relationships of the input images frame by frame can be input into the temporal extraction network. The temporal extraction network can memorize the features of the target spatial relationships extracted at the previous moment and apply them to the feature extraction of the target spatial relationships input at the current moment. The output of the temporal extraction network based on the target spatial relationships of the last frame image, that is, the local features containing the temporal associations of the target spatial relationships of all the images in the image frame sequence.

[0101] Here, the temporal extraction network can be a Long Short-Term Memory (LSTM), a Recurrent Neural Network (RNN), etc. The embodiments of the present invention do not make specific limitations thereto.

[0102] Based on any of the above embodiments, Figure 5 is a schematic structural diagram of the local extraction network provided by the present invention. As Figure 5 shown, the local extraction network includes a target detection network, a graph convolutional network, and a recurrent neural network, where the recurrent neural network plays the role of temporal extraction. After the target detection of each frame image in the image frame sequence is completed through the target detection network, the target feature maps of each frame image can be constructed based on the target features and target positions of the detected targets included in each frame image. By performing feature extraction on the target feature maps through a Graph Convolutional Network (GCN), the target spatial relationships of each frame image can be obtained. On this basis, the target spatial relationships of the input images frame by frame are input into the recurrent neural network, so that the input of the hidden layer of the recurrent neural network includes not only the output of the input layer but also the output of the hidden layer at the previous moment, thereby obtaining the local features containing the temporal associations of the target spatial relationships of all the images in the image frame sequence.

[0103] Based on any of the above embodiments, step 140 includes:

[0104] Based on the fusion classification network, fuse the global features and the local features, and perform context extraction based on the fused features to obtain context features, and apply the context features to accident classification to determine the accident detection result.

[0105] Specifically, after obtaining the global features and local features respectively, they can be fused through a fusion classification network, and classification for accident detection can be performed based on the fused features. In this process, the fusion of global features and local features can be achieved by means of concatenation or weighted summation, etc. Considering that traffic accidents have highly contextual characteristics, after obtaining the fused features, the fused features can be used to construct temporal information, that is, context extraction, through RNN or LSTM, etc., so as to obtain context features that can reflect the temporal sequence of the video to be detected. After obtaining the context features, classification can be performed based on the context features, so as to obtain the accident detection result of whether there is a traffic accident in the video to be analyzed, that is, the accident detection is completed.

[0106] Among them, the extraction of context features can be achieved through RNN, LSTM, etc. Considering that LSTM adds filtering of past states on the basis of RNN, so that it can select which states have more influence on the current, solves the problem of gradient disappearance in long-term modeling of RNN, and is more suitable for constructing long-term dependencies. As an option, LSTM can be set in the fusion classification network to achieve context feature extraction.

[0107] Based on any of the above embodiments, the global extraction network, the local extraction network, and the fusion classification network are determined based on the following steps:

[0108] Construct an initial detection network based on the initial global extraction network, the initial local extraction network, and the initial fusion classification network;

[0109] Based on the first sample video with accident labels, train the initial detection network, and determine the global extraction network, the local extraction network, and the fusion classification network based on the trained initial detection network.

[0110] Specifically, the initial global extraction network, the initial local extraction network, and the initial fusion classification network respectively correspond to the initialization networks of the global extraction network, the local extraction network, and the fusion classification network. Connect the outputs of the initial global extraction network and the initial local extraction network with the input of the initial fusion classification network, that is, an initial detection network is formed. Here, the network parameters of the initial global extraction network, the initial local extraction network, and the initial fusion classification network themselves can be initialized or pre-trained. The embodiments of the present invention do not make specific limitations on this.

[0111] After determining the initial detection network, the first sample video that has been pre-collected and labeled with accident labels indicating whether an accident has occurred can be applied to the training of the initial detection network, so as to realize the supervised training of the initial detection network. The trained initial detection network includes the global extraction network, the local extraction network, and the fusion classification network applied in accident detection.

[0112] Further, for the accident labels of the first sample videos, the normal first sample videos may not be labeled. For the first sample videos with accidents, the starting frame when the accident occurs and the ending frame after the accident occurrence can be labeled as accident labels. The position of the ending frame referred to here can be until all vehicles stop or the video ends.

[0113] Based on any of the above embodiments, during the training process of the overall initial detection network, taking the image frame sequence of the first sample video as the input of the initial detection network, the predicted detection results output by the initial detection network for the first sample video can be obtained. The predicted detection results may include the probability that the initial detection network predicts that the first sample video contains an accident. After obtaining the predicted detection results, the parameters of the initial detection network can be updated iteratively based on the following loss function Loss1:

[0114]

[0115] In the formula, N1 is the number of the first sample videos, y 1i is the accident label of the i-th first sample video, y 1i takes 0 to represent normal and 1 to represent the existence of an accident, p(y 1i ) is the predicted detection result of the i-th first sample video, and the value of p(y 1i ) ranges from 0 to 1.

[0116] Based on any of the above embodiments, the initial global extraction network is trained jointly with the global classification network based on the second sample videos carrying accident labels, and the global classification network is used for accident detection based on global features.

[0117] Specifically, in order to improve the training efficiency and training effect of the initial detection network, before constructing the initial detection network based on the initial global extraction network, the initial global extraction network can be pre-trained.

[0118] Here, when training the initial global extraction network, it needs to be implemented jointly with the initial classification network. The initial classification network is connected in series after the initial global extraction network. After the initial global extraction network extracts the global features of the input second sample video, the initial classification network can perform accident classification based on the global features extracted by the initial global extraction network and output the prediction results based on the global features. Here, the prediction results based on the global features may include the probability that the initial global extraction network and the global classification network jointly predict that the second sample video contains an accident. The loss function Loss2 for the joint training of the two can be expressed as the following formula:

[0119]

[0120] Wherein, N2 is the number of the second sample videos, and y 2i is the accident label of the i-th second sample video, and y 2i takes 0 to indicate normal and takes 1 to indicate the existence of an accident. p(y 2i ) is the prediction result based on the global features of the i-th second sample video, and the value of p(y 2i ) is between 0 and 1.

[0121] It should be noted that the second sample videos here and the first sample videos in the above embodiments may be the same batch of sample videos or different batches of sample videos. The embodiments of the present invention do not make specific limitations on this.

[0122] Based on any of the above embodiments, the global extraction network, the local extraction network, and the fusion classification network can be obtained through two-stage training:

[0123] In the first stage, an initial global extraction network is constructed, and the initial global extraction network and the global classification network are jointly trained based on the second sample videos carrying accident labels. Figure 6 is the structural schematic diagram of the first-stage training provided by the present invention. As Figure 6 shown, the output of the initial global extraction network is the input of the global classification network. The global classification network here can be a feed-forward network (FFN) or other types of classification networks.

[0124] In the second stage, the trained initial global extraction network is jointly constructed with the initial local extraction network and the initial fusion classification network to form an initial detection network, and the initial detection network is trained based on the first sample videos carrying accident labels, so as to obtain the trained global extraction network, local extraction network, and fusion classification network.

[0125] Based on any of the above embodiments, Figure 7 is the flow schematic diagram of the accident detection method provided by the present invention. As Figure 7 shown, in an actual scenario, a monitoring camera installed locally can be used to obtain a monitoring video as the video to be analyzed, and the image frame sequence of the video to be analyzed is transmitted to an accident detection network including a global extraction network, a local extraction network, and a fusion classification network.

[0126] Among them, the multi-layer three-dimensional convolutional network in the global extraction network performs multi-layer three-dimensional convolution on the image frame sequence to obtain the first convolutional feature and the second convolutional feature. The attention network then applies the first convolutional feature to determine the attention weight of the second convolutional feature, and applies the attention weight to weight the second convolutional feature to obtain the global feature.

[0127] The target detection network in the local extraction network determines the target features and target positions of the detected targets in each frame of the image frame sequence, and thereby determines the target feature map of each frame; the spatiotemporal extraction network in the local extraction network extracts spatiotemporal information from the target feature map of each frame to obtain local features.

[0128] Subsequently, the fusion classification network fuses the global features and local features output by the global extraction network and the local extraction network, respectively, and performs context extraction based on the fused features to obtain context features, and applies the context features to classify accidents, determine and output accident detection results. The accident detection result here can be the probability of an accident. If the probability is greater than a preset threshold, such as 0.7, 0.8, etc., it is judged that an accident has occurred and reported to the relevant departments. Alternatively, the accident detection result here can also be whether an accident has occurred. If an accident has occurred, it is directly reported to the relevant departments.

[0129] The method provided by the embodiment of the present invention divides the image frame sequence into two dimensions, global and local, and performs feature extraction on the image frame sequence. In the global extraction network, the network pays as much attention to the relevant information of the vehicle and its surroundings as possible by combining multi-layer three-dimensional convolution with spatiotemporal attention. In the local extraction network, the image frame sequence first passes through the target detection network to extract and aggregate the target information of the current frame, and then constructs a target feature graph. Each node in the graph represents the target feature contained in each frame. Through the graph convolution operation, the spatial relationship between local targets is aggregated, and on this basis, the spatiotemporal relationship of the target is obtained, that is, the local feature is obtained. Subsequently, the classification network fuses the features extracted from the two dimensions, and by extracting context features, the temporal features that aggregate the global context and local target relationship are obtained, that is, the context features. Finally, the probability of an accident, that is, the accident detection result, is output based on the context features. The network proposed in the embodiment of the present invention has better versatility. Compared with the 3D convolution traffic accident detection method in the related technology, it is difficult to identify traffic accidents with small scene changes, and the conventional traffic accident detection method of detecting and tracking targets is too dependent on the accuracy of target detection and tracking. If the target is not detected, the occurrence of a traffic accident cannot be determined at all. The speed of the target appearance in some accident scenes changes dramatically, which can easily lead to the inability to track the target. The method proposed in the embodiment of the present invention can better detect traffic accidents.

[0130] In addition, the application of fragmented image frame sequences, such as every 16 frames or every 20 frames as an image frame sequence for accident detection, can fully explore the characteristics of the accident scene and detect the occurrence of traffic accidents in real time. Compared with directly outputting the accident probability at the video level, it has stronger real-time and versatility.

[0131] Based on any of the above embodiments, Figure 8It is a schematic structural diagram of the accident detection device provided by the present invention. As Figure 8 shown, the device includes:

[0132] A sequence determination unit 810, configured to determine an image frame sequence of a video to be detected;

[0133] A global extraction unit 820, configured to perform three-dimensional feature extraction on the image frame sequence based on a global extraction network to obtain a global feature of the video to be detected;

[0134] A local extraction unit 830, configured to determine a local feature of the video to be detected based on a local extraction network by applying detection targets and target positions of each frame image in the image frame sequence;

[0135] A fusion classification unit 840, configured to determine an accident detection result of the video to be detected based on a fusion classification network by applying the global feature and the local feature.

[0136] The device provided by the embodiment of the present invention performs accident detection by combining the global feature and the local feature of the video to be detected. Whether it is for the situation where the target detection fails or the target is lost due to a drastic change in the target, or for the situation where the scene change is not obvious, it can accurately and reliably complete the accident detection, so as to ensure that traffic accidents can be monitored in time and facilitate the timeliness of accident investigation.

[0137] Based on any of the above embodiments, the global extraction unit is used for:

[0138] Based on a multi-layer three-dimensional convolutional network in the global extraction network, perform multi-layer three-dimensional convolution on the image frame sequence to obtain a first convolutional feature and a second convolutional feature, where the first convolutional feature is obtained by convolution before the second convolutional feature;

[0139] Based on an attention network in the global extraction network, apply the first convolutional feature to determine an attention weight of the second convolutional feature, and apply the attention weight to weight the second convolutional feature to obtain the global feature.

[0140] Based on any of the above embodiments, the first convolutional feature and the second convolutional feature are respectively convolutional features output by the penultimate layer and the last layer of the multi-layer three-dimensional convolution;

[0141] Specifically, the global extraction unit is used for:

[0142] Perform single-layer three-dimensional convolution on the first convolutional feature to obtain a third convolutional feature with the same dimension as the second convolutional feature;

[0143] Determine the attention weight based on the third convolutional feature.

[0144] Based on any of the above embodiments, the local extraction unit is configured to:

[0145] Based on the object detection network in the local extraction network, determine the object features and object positions of the detection objects in each frame image of the image frame sequence;

[0146] Based on the spatio-temporal extraction network in the local extraction network, perform spatio-temporal information extraction on the object feature maps of each frame image to obtain the local features, where the object feature maps are determined based on the object features and object positions of the detection objects in the corresponding images.

[0147] Based on any of the above embodiments, the local extraction unit is specifically configured to:

[0148] Based on the graph convolutional network in the spatio-temporal extraction network, perform spatial information extraction on the object feature maps of each frame image to obtain the object spatial relationships of each frame image;

[0149] Based on the temporal extraction network in the spatio-temporal extraction network, perform temporal feature extraction on the object spatial relationships of each frame image to obtain the local features of the video to be detected.

[0150] Based on any of the above embodiments, the fusion classification unit is configured to:

[0151] Based on the fusion classification network, fuse the global features and the local features, perform context extraction based on the fused features to obtain context features, and use the context features to perform accident classification to determine the accident detection result.

[0152] Based on any of the above embodiments, the apparatus further includes a training unit, configured to:

[0153] Construct an initial detection network based on an initial global extraction network, an initial local extraction network, and an initial fusion classification network;

[0154] Train the initial detection network based on a first sample video carrying accident labels, and determine the global extraction network, the local extraction network, and the fusion classification network based on the trained initial detection network.

[0155] Based on any of the above embodiments, the initial global extraction network is trained by jointly training with a global classification network based on a second sample video carrying accident labels, and the global classification network is used to perform accident detection based on global features.

[0156] Figure 9 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 9As shown in the figure, the electronic device may include: a processor 910, a communications interface 920, a memory 930, and a communication bus 940. Among them, the processor 910, the communications interface 920, and the memory 930 communicate with each other through the communication bus 940. The processor 910 may call the logical instructions in the memory 930 to execute an accident detection method, and the method includes:

[0157] Determine the image frame sequence of the video to be detected;

[0158] Based on the global extraction network, perform three-dimensional feature extraction on the image frame sequence to obtain the global features of the video to be detected;

[0159] Based on the local extraction network, apply the detection targets and target positions of each frame image in the image frame sequence to determine the local features of the video to be detected;

[0160] Based on the fusion classification network, apply the global features and the local features to determine the accident detection result of the video to be detected.

[0161] In addition, when the logical instructions in the above-mentioned memory 930 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0162] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the accident detection method provided by the above-mentioned various methods. The method includes:

[0163] Determine the image frame sequence of the video to be detected;

[0164] Based on the global extraction network, perform three-dimensional feature extraction on the image frame sequence to obtain the global features of the video to be detected;

[0165] Based on the local extraction network, the detection targets and target positions of each frame image in the image frame sequence are used to determine the local features of the video to be detected.

[0166] Based on the fusion classification network, the global features and the local features are used to determine the accident detection result of the video to be detected.

[0167] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the accident detection method provided above. The method includes:

[0168] Determine the image frame sequence of the video to be detected;

[0169] Based on the global extraction network, perform three-dimensional feature extraction on the image frame sequence to obtain the global features of the video to be detected;

[0170] Based on the local extraction network, the detection targets and target positions of each frame image in the image frame sequence are used to determine the local features of the video to be detected;

[0171] Based on the fusion classification network, the global features and the local features are used to determine the accident detection result of the video to be detected.

[0172] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0173] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An accident detection method, characterized in that, comprising: Determine the image frame sequence of the video to be detected; Based on the global extraction network, perform three-dimensional feature extraction on the image frame sequence to obtain the global features of the video to be detected, and the global features reflect the information of the video to be detected in the traffic scene as a whole; Based on the local extraction network, apply the detection targets and target positions of each frame image in the image frame sequence to determine the local features of the video to be detected. The local features reflect the information of the detection targets in the video to be detected in terms of space and time, and the local features reflect the information of the video to be detected in terms of detection targets from the level of local targets; Based on the fusion classification network, apply the global features and the local features to determine the accident detection result of the video to be detected; The step of determining the local features of the video to be detected by applying the detection targets and target positions of each frame image in the image frame sequence based on the local extraction network includes: Based on the target detection network in the local extraction network, determine the target features and target positions of the detection targets in each frame image of the image frame sequence; Based on the graph convolutional network in the spatio-temporal extraction network in the local extraction network, perform spatial information extraction on the target feature maps of each frame image to obtain the target spatial relationships of each frame image. The target feature maps are determined based on the target features and target positions of the detection targets in the corresponding images; Based on the temporal extraction network in the spatio-temporal extraction network, perform temporal feature extraction on the target spatial relationships of each frame image to obtain the local features of the video to be detected.

2. The accident detection method according to claim 1, characterized in that, The step of performing three-dimensional feature extraction on the image frame sequence based on the global extraction network to obtain the global features of the video to be detected includes: Based on the multi-layer three-dimensional convolutional network in the global extraction network, perform multi-layer three-dimensional convolution on the image frame sequence to obtain a first convolutional feature and a second convolutional feature. The first convolutional feature is obtained by convolution before the second convolutional feature; Based on the attention network in the global extraction network, apply the first convolutional feature to determine the attention weights of the second convolutional feature, and apply the attention weights to weight the second convolutional feature to obtain the global features.

3. The accident detection method according to claim 2, characterized in that, The first convolutional feature and the second convolutional feature are respectively the convolutional features output by the penultimate layer and the last layer of the multi-layer three-dimensional convolution; The step of applying the first convolutional feature to determine the attention weights of the second convolutional feature includes: Perform single-layer three-dimensional convolution on the first convolutional feature to obtain a third convolutional feature with the same dimension as the second convolutional feature; Determine the attention weights based on the third convolutional feature.

4. The accident detection method according to claim 1, characterized in that, The step of determining the accident detection result of the video to be detected by applying the global features and the local features based on the fusion classification network includes: Based on the fusion classification network, fuse the global feature and the local feature, and perform context extraction based on the fused feature to obtain a context feature, and apply the context feature to perform accident classification to determine the accident detection result.

5. The accident detection method according to any one of claims 1 to 4, wherein, the global extraction network, the local extraction network and the fusion classification network are determined based on the following steps: Construct an initial detection network based on an initial global extraction network, an initial local extraction network and an initial fusion classification network; Based on the first sample video with accident labels, train the initial detection network, and determine the global extraction network, the local extraction network and the fusion classification network based on the trained initial detection network.

6. The accident detection method according to claim 5, wherein, the initial global extraction network is obtained by jointly training with a global classification network based on a second sample video with accident labels, and the global classification network is used to perform accident detection based on global features.

7. An accident detection device, wherein, comprising: a sequence determination unit for determining an image frame sequence of a video to be detected; a global extraction unit for performing three-dimensional feature extraction on the image frame sequence based on a global extraction network to obtain the global feature of the video to be detected, and the global feature reflects the information of the video to be detected in a traffic scene as a whole; a local extraction unit for determining the local feature of the video to be detected based on a local extraction network by applying the detection target and the target position of each frame image in the image frame sequence, and the local feature reflects the information of the detection target in the video to be detected in terms of space and time, and the local feature reflects the information of the video to be detected on the detection target from the level of local targets; a fusion classification unit for determining the accident detection result of the video to be detected based on a fusion classification network by applying the global feature and the local feature; the local extraction unit is specifically used for: Determine the target feature and the target position of the detection target in each frame image of the image frame sequence based on the target detection network in the local extraction network; Based on the graph convolutional network in the spatio-temporal extraction network in the local extraction network, perform spatial information extraction on the target feature map of each frame image to obtain the target spatial relationship of each frame image, and the target feature map is determined based on the target feature and the target position of the detection target in the corresponding image; Based on the temporal extraction network in the spatio-temporal extraction network, perform temporal feature extraction on the target spatial relationship of each frame image to obtain the local feature of the video to be detected.

8. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, the steps of the accident detection method according to any one of claims 1 to 6 are implemented.

9. A non-transitory computer-readable storage medium, on which a computer program is stored, wherein, When the computer program is executed by a processor, the steps of the accident detection method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Traffic police command gesture recognition method based on skeleton joint point sequence

    CN110837778A