A method, device, equipment and readable storage medium for detecting the speed of a target object
By extracting and aligning the frame image sequences in the video data, combining with the speed recognition detection model, directly regressing the recognition speed of the target object, the problem of low speed detection accuracy in the prior art is solved, and more efficient and accurate speed detection is achieved.
Patent Information
- Application Number
- CN202111012847.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-08-31
AI Technical Summary
The prior art cannot directly obtain information about the target object in three-dimensional space, resulting in a long pipeline for speed detection and low accuracy.
By collecting video data, the frame image sequence is extracted, and the feature extraction model and behavior recognition model are used to align the feature maps. After the fusion, the velocity recognition detection model is input to directly regress the recognition speed of the target object.
The operation process of speed detection is optimized, the accuracy and continuity of speed detection is improved, and the post-processing process is streamlined, improving the performance of target object speed detection.
Smart Images

Figure CN113887284B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vision technology, and in particular, to a method, device, equipment and readable storage medium for detecting the speed of an object. Background Art
[0002] With the sharp increase in the number of automobiles, traffic accidents during vehicle driving are becoming increasingly frequent. In many traffic accidents, speed problems have become one of the main causes. At the same time, with the development of intelligent transportation and autonomous driving technologies, it is also necessary to be able to detect the speed of adjacent vehicles other than the vehicle driven by the driver through more precise means, so as to analyze the speed and driving conditions of the vehicle driven by oneself, and to reasonably and precisely plan the vehicle operation path in the subsequent stage. Therefore, it is very necessary to timely and accurately detect the running speed of other vehicles for the safety of the vehicle during its operation.
[0003] With the development of computer vision technology, computer vision technology has received great attention and is widely used in fields such as autonomous driving, area monitoring, image navigation guidance, urban security, terrain matching, etc. Stereo vision technology is an important part of computer vision, with advantages such as high measurement accuracy, low usage cost, and non-contact measurement. Using stereo vision technology, the detection and tracking of an object and the positioning and speed estimation of the object can be realized. For example, in an application based on video structuring, after the object passes through the tracking algorithm, a unique identifier and its corresponding motion trajectory will be obtained. Using these two data, some subsequent work can be done: speed detection (traffic application scenarios), counting (traffic application scenarios, security application scenarios), and behavior detection (traffic application scenarios, security application scenarios). Currently, the detection of an object such as a vehicle through a video monitoring system can only obtain the two-dimensional information of the object on the camera projection plane and the size of the object on the projection plane. The detected information is actually the two-dimensional information of the object on the camera projection plane, rather than the object itself in three-dimensional space. Since the actual size of the vehicle cannot be determined solely by the two-dimensional information in the camera image, real-time speed detection cannot be achieved. The commonly used existing technical means is to first obtain the information of the object in the two-dimensional space, then convert the obtained two-dimensional information into the information of the object in the three-dimensional space, and combine with a filtering algorithm to improve the problems of inaccurate tracking and easy loss in the original tracking algorithm. The overall algorithm flow (pipeline) of the above process is relatively long, resulting in low overall accuracy of speed estimation and difficult parameter adjustment.
[0004] In order to be able to timely detect the running speed of an object, a new method for detecting the speed of an object needs to be designed. Summary of the Invention
[0005] The present invention provides a method, apparatus, device and readable storage medium for detecting the speed of a target object, so as to solve the defect in the prior art that the information of the target object in the three-dimensional space cannot be directly obtained, resulting in a long pipeline and low accuracy, and to optimize the operation process of the entire pipeline, improve the speed detection accuracy and continuity, and streamline the post-processing process of speed detection.
[0006] The present invention provides a method for detecting the speed of a target object, including the following steps:
[0007] Collect video data to be detected, and extract a frame image sequence containing the target object to be detected from the video data to be recognized;
[0008] Input the feature map corresponding to the aligned first frame image sequence, the first frame image sequence, and the second frame image sequence into a speed recognition detection model, and obtain the recognition speed of the target object to be detected corresponding to the second frame image sequence output by the speed recognition detection model;
[0009] Wherein, the speed recognition detection model is trained based on a sample frame image sequence extracted from sample video data. The first frame image sequence and the second frame image sequence both belong to the frame image sequence, and the first frame image sequence is adjacent to and before the second frame image sequence. The feature map is obtained by inputting the first frame image sequence into a feature extraction model, and the feature map is aligned based on a first offset amount, and the first offset amount is obtained by inputting the feature map into a behavior recognition model.
[0010] According to the method for detecting the speed of a target object provided by the present invention, the step of inputting the feature map corresponding to the aligned first frame image sequence, the first frame image sequence, and the second frame image sequence into a speed recognition detection model, and obtaining the recognition speed of the target object to be detected corresponding to the second frame image sequence output by the speed recognition detection model specifically includes the following steps:
[0011] Extract the first frame image sequence and the second frame image sequence from the frame image sequence;
[0012] Input the first frame image sequence into the feature extraction model, and obtain the feature maps corresponding to each image in the first frame image sequence output by the feature extraction model, and use one of the feature maps as a reference feature map; wherein, the feature extraction model is trained based on a sample frame image sequence extracted from sample video data;
[0013] Based on the reference feature map, input the feature map corresponding to the first frame image sequence into the behavior recognition model to obtain the first offset output by the behavior recognition model; wherein, the behavior recognition model is trained based on sample feature maps and reference sample feature maps;
[0014] Based on the first offset, align the feature map corresponding to the first frame image sequence to the reference feature map;
[0015] After fusing the aligned feature map, the first frame image sequence, and the second frame image sequence, input them into the speed recognition and detection model to obtain the recognition speed output by the speed recognition and detection model.
[0016] According to the object speed detection method provided by the present invention, the feature extraction model is trained through the following steps:
[0017] Collect the sample video data and extract the third frame image sequence from the sample video data;
[0018] Use the third frame image sequence as the input data after training, and perform training in a deep learning manner to obtain the feature extraction model for generating the feature map of the frame image sequence to be predicted.
[0019] According to the object speed detection method provided by the present invention, the speed recognition and detection model is trained through the following steps:
[0020] Extract the sample frame image sequence from the sample video data, and extract the third frame image sequence and the fourth frame image sequence from the sample frame image sequence; wherein, both the third frame image sequence and the fourth frame image sequence belong to the sample frame image sequence, and the third frame image sequence is adjacent to the fourth frame image sequence and is located before the fourth frame image sequence;
[0021] Input the third frame image sequence into the feature extraction model to obtain the sample feature maps respectively corresponding to each image in the third frame image sequence output by the feature extraction model, and use one of the sample feature maps as the reference sample feature map;
[0022] Based on the reference sample feature map, input the sample feature map corresponding to the third frame image sequence into the behavior recognition model to obtain the second offset output by the behavior recognition model;
[0023] Based on the second offset, align the sample feature map corresponding to the third frame image sequence to the reference sample feature map;
[0024] Fuse the aligned sample feature map, the third frame image sequence, and the fourth frame image sequence, and use the fused result as the input data for training. Then, perform training using deep learning to obtain the speed recognition detection model for generating the frame image sequence to be predicted.
[0025] According to the object speed detection method provided by the present invention, the feature extraction model and the speed recognition detection model are the same physical model.
[0026] According to the object speed detection method provided by the present invention, the behavior recognition model is trained through the following steps:
[0027] Using the reference sample feature map as the alignment reference, use the sample feature map corresponding to the third frame image sequence as the input data for training, and perform training using deep learning to obtain the behavior recognition model for generating the offset of the feature map to be aligned.
[0028] The present invention also provides an object speed detection method, including:
[0029] A first acquisition module, configured to acquire video data to be detected, and extract a frame image sequence including the object to be detected from the video data to be recognized;
[0030] A speed detection module, configured to input the feature map corresponding to the aligned first frame image sequence, the first frame image sequence, and the second frame image sequence into the speed recognition detection model, and obtain the recognition speed of the object to be detected corresponding to the second frame image sequence output by the speed recognition detection model;
[0031] Wherein, the speed recognition detection model is trained based on a sample frame image sequence extracted from sample video data. The first frame image sequence and the second frame image sequence both belong to the frame image sequence, and the first frame image sequence is adjacent to the second frame image sequence and is located before the second frame image sequence. The feature map is obtained by inputting the first frame image sequence into the feature extraction model, and the feature map is aligned based on a first offset, and the first offset is obtained by inputting the feature map into the behavior recognition model.
[0032] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any one of the above object speed detection methods are implemented.
[0033] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned target object speed detection methods are implemented.
[0034] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above-mentioned target object speed detection methods are implemented.
[0035] The target object speed detection method, device, equipment and readable storage medium provided by the present invention use the continuous multi-frame image data included in the collected original video data as the second frame image sequence to be predicted, and use the continuous multi-frame image data included in the original video data that is adjacent and before the second frame image sequence as the first frame image sequence; input the first frame image sequence into the feature extraction model trained by this method, and then extract the corresponding feature map. After aligning the feature map, it is fused with the first frame image sequence and the second frame image sequence to be predicted and then input into the speed recognition and detection model together, directly regressing to the recognition speed of the target object to be detected. Based on the end-to-end deep learning method, the operation process of the entire pipeline is optimized, the speed detection accuracy and continuity are improved, and the post-processing process of speed detection is streamlined, making the recognition speed detection of the target object achieve better performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 is a schematic flowchart of the target object speed detection method provided by the present invention;
[0038] Figure 2 is a specific schematic flowchart of step S200 in the target object speed detection method provided by the present invention;
[0039] Figure 3 is a schematic structural diagram of the target object speed detection device provided by the present invention;
[0040] Figure 4 is a specific schematic structural diagram of the speed detection module in the target object speed detection device provided by the present invention;
[0041] Figure 5 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0043] The following will be combined with Figure 1 Describe the method for detecting the speed of the target object of the present invention. Taking the application in the speed detection of a vehicle as an example, the method includes the following steps:
[0044] S100. Collect the video data to be detected, and extract the frame image sequence containing the target object to be detected from the video data to be recognized. In this embodiment, the target object is a vehicle, more specifically, other vehicles other than the autonomous vehicle using this method. It can be understood that the video data to be detected collected through an in-vehicle camera and other means contains multiple frame image data. Each frame of image data may or may not contain a vehicle. And when the image data contains a vehicle, all the vehicles contained therein are used as the target objects to be detected, and the speed of each vehicle used as the target object is detected separately.
[0045] S200. Input the feature map corresponding to the aligned first frame image sequence, the first frame image sequence, and the second frame image sequence into the speed recognition detection model to obtain the recognition speed of the target object to be detected corresponding to the second frame image sequence output by the speed recognition detection model, that is, the recognition speed of the vehicle to be detected corresponding to the time frame of the second frame image sequence. Therefore, this recognition speed is the real-time speed of other vehicles, specifically, the relative moving speed of other vehicles in three-dimensional space relative to the autonomous vehicle using this method. This recognition speed can be used as a parameter for adjusting the autonomous driving strategy of the autonomous vehicle using this method later, and for serving the control and prediction (Prediction) in other aspects of autonomous driving. Specifically, through some post-processing, combined with information such as the position, distance, and speed of the vehicle, to determine which vehicles are key targets and which are non-key targets, and finally output the parameters of the key targets, and control and predict the autonomous driving process through the parameters of the above key targets.
[0046] The speed recognition detection model is trained based on a sequence of sample frame images extracted from sample video data; both the first frame image sequence and the second frame image sequence belong to the frame image sequence, and the first frame image sequence is adjacent to and before the second frame image sequence, that is, the first frame image sequence is a sequence of continuous multiple frames (at least two frames) of image data and is part of the frame image sequence, and the second frame image sequence is also a sequence of continuous multiple frames (at least two frames) of image data and is also part of the frame image sequence; the feature map is obtained by inputting the first frame image sequence into the feature extraction model, and the feature map is aligned based on the first offset, and the first offset is obtained by inputting the feature map into the behavior recognition model.
[0047] In this method, both the speed recognition detection model and the feature extraction model are Convolutional Neural Networks (CNN) models. Essentially, a CNN is a Multilayer Perceptron (MLP). A CNN uses local connections and shared weights. On the one hand, it reduces the number of weights, making the network easy to optimize. On the other hand, it reduces the risk of overfitting. A CNN is a type of neural network. Its weight-sharing network structure makes it more similar to a biological neural network, reducing the complexity of the network model and the number of weights. This advantage is more obvious when the input of the network is a multi-dimensional image, allowing the image to be directly used as the input data of the CNN model, avoiding the complex feature extraction and data reconstruction processes in traditional recognition algorithms. Currently, CNN models have many advantages in two-dimensional image processing. For example, the network can automatically extract image features including color, texture, shape, and the topological structure of the image; in dealing with two-dimensional image problems, especially in applications with good robustness and computational efficiency for recognizing displacements, scalings, and other forms of distortion invariance.
[0048] The CNN itself can adopt different combinations of neurons and learning rules. The CNN has some advantages that traditional technologies do not have: good fault tolerance, parallel processing ability, and self-learning ability. It can handle problems in complex environmental information, unclear background knowledge, and unclear inference rules. It allows for large defects and distortions in samples, has a fast operation speed, good adaptability, and high resolution. The CNN model integrates the feature extraction function into the MLP through structural reorganization and weight reduction, omitting the complex image feature extraction process before recognition. The CNN model consists of an input layer, an output layer, and multiple hidden layers. The hidden layers can be divided into a convolution layer, a pooling layer, a RELU layer, and a fully connected layer. Among them, the convolution layer is the core of the CNN model. The parameters of the convolution layer consist of a set of learnable filters or kernels, which have a small receptive field and extend to the entire depth of the input volume. During feedforward, each filter convolves the input data, calculates the dot product between the filter and the input data, and generates a two-dimensional activation map of the filter. The input data is generally a two-dimensional vector, but it may have a height. That is, the convolution layer is used to convolve the input layer to extract higher-level features.
[0049] In this method, the feature map is the feature output (extracted) after filtering a certain layer in the convolution layer of the feature extraction model. In this embodiment, specifically, it is the shared feature map output by the last layer of the convolution layer of the feature extraction model.
[0050] In this method, the speed recognition detection model and the feature extraction model can be the same physical model.
[0051] The CNN model adopts the deep learning method. During the training process of the deep learning method, a prediction result is obtained from the input end to the output end. There is only one neural network in the middle. Comparing with the real result, an error will be obtained. This error will be passed through each layer in the CNN model, and each layer will make adjustments according to this error until the CNN model converges or reaches the expected effect before ending. This is the end-to-end deep learning method, which can save the data annotation done before each independent learning task execution, improve the accuracy of the output result, and save the learning cost.
[0052] The object speed detection method of the present invention takes the continuous multi-frame image data contained in the collected original video data as the second frame image sequence to be predicted, and takes the continuous multi-frame image data contained in the same original video data that is adjacent and before the second frame image sequence as the first frame image sequence; inputs the first frame image sequence into the feature extraction model trained by this method, and then extracts the corresponding feature maps. After aligning the feature maps, they are fused with the first frame image sequence and the second frame image sequence to be predicted and then input into the speed recognition detection model together, directly regressing to the recognition speed of the object to be detected. Based on the end-to-end deep learning method, it optimizes the operation process of the entire pipeline, improves the speed detection accuracy and continuity, and simplifies the post-processing process of speed detection, making the recognition speed detection of the object achieve better performance.
[0053] In traditional speed detection methods, the motion speed of an object is estimated by combining object detection and object tracking. The recognition speed obtained by the present invention is the relative moving speed of the object in three-dimensional space relative to the autonomous driving vehicle using this method. Combining it with the moving speed of the autonomous driving vehicle itself using this method, such as the vehicle speed (obtained through in-vehicle sensors), the true speed of the object in three-dimensional space can be obtained, which is significantly different from the motion speed of the object obtained in traditional methods, which is the moving speed of the two-dimensional center point of the object in the two-dimensional image, i.e., the pixel speed.
[0054] The following combines Figure 2 to describe the object speed detection method of the present invention. Step S200 specifically includes the following steps:
[0055] S210. Extract the first frame image sequence and the second frame image sequence from the frame image sequence.
[0056] S220. Input the first frame image sequence into the feature extraction model to obtain the feature maps corresponding to each image in the first frame image sequence output by the feature extraction model, and use one of the feature maps as the reference feature map.
[0057] In this method, the feature extraction model is trained based on the sample frame image sequence extracted from the sample video data. In this embodiment, the feature extraction model adopts the supervised learning method in the deep learning process, that is, the sample feature map is used as the label for supervised learning.
[0058] S230. Based on the reference feature map, input the feature maps corresponding to the first frame image sequence into the behavior recognition model to obtain the first offset output by the behavior recognition model.
[0059] S240. Based on the first offset, align the feature maps corresponding to the first frame image sequence to the reference feature map.
[0060] In this method, the behavior recognition model is trained based on the sample feature map and the reference sample feature map. The role of the behavior recognition model is to learn an offset map to align the feature maps of multiple consecutive frames. Specifically, the input data of the behavior recognition model is the feature maps of multiple consecutive frames, that is, the feature maps corresponding to each image in the first frame image sequence, and the output data is the offset required to move other feature maps to the reference feature map, that is, the first offset. Through the first offset, the feature maps corresponding to the first frame image sequence can be aligned to the reference feature map.
[0061] S250. After fusing the aligned feature maps, the first frame image sequence, and the second frame image sequence, input them into the speed recognition and detection model to obtain the recognition speed directly regressed by the speed recognition and detection model. In this embodiment, the speed recognition and detection model adopts the supervised learning method during the deep learning process, that is, using the recognition speed corresponding to the fused data as a label for supervised learning.
[0062] The feature extraction model in this method is trained through the following steps:
[0063] A100. Collect sample video data and extract the third frame image sequence from the sample video data;
[0064] A200. Use the third frame image sequence as the input data after training, and adopt the deep learning method for training to obtain a feature extraction model for generating the feature maps of the frame image sequence to be predicted.
[0065] The speed recognition and detection model in this method is trained through the following steps:
[0066] A300. Extract the sample frame image sequence from the sample video data, and extract the third frame image sequence and the fourth frame image sequence from the sample frame image sequence. Among them, the third frame image sequence and the fourth frame image sequence both belong to the sample frame image sequence, and the third frame image sequence is adjacent to the fourth frame image sequence and is before the fourth frame image sequence.
[0067] A400. Input the third frame image sequence into the feature extraction model to obtain the sample feature maps corresponding to each image in the third frame image sequence output by the feature extraction model, and use one of the sample feature maps as the reference sample feature map.
[0068] A500. Based on the reference sample feature map, input the sample feature maps corresponding to the third frame image sequence into the behavior recognition model to obtain the second offset output by the behavior recognition model;
[0069] A600. Based on the second offset, align the sample feature maps corresponding to the third frame image sequence to the reference sample feature map;
[0070] A700. After fusing the aligned sample feature map, the third-frame image sequence, and the fourth-frame image sequence, use them as the input data for training, and adopt deep learning to train to obtain a speed recognition detection model for generating the recognition speed of the frame image sequence to be predicted. In step A700, during the deep learning process, the recognition speed corresponding to the fused data can be used as a label for supervised learning.
[0071] In this method, the behavior recognition model is trained through the following steps:
[0072] A800. Using the reference sample feature map as the alignment reference, use the sample feature map corresponding to the third-frame image sequence as the input data for training, and adopt deep learning to train to obtain a behavior recognition model for generating the offset of the feature map to be aligned.
[0073] The object speed detection device provided by the present invention will be described below. The object speed detection device described below can be mutually corresponding and referred to the object speed detection method described above.
[0074] Next, in combination with Figure 3 Describe the object speed detection device of the present invention. Taking the application in the speed detection of vehicles as an example, the device includes:
[0075] The first acquisition module 100 is used to acquire the video data to be detected and extract the frame image sequence containing the object to be detected from the video data to be recognized. In this embodiment, the object is a vehicle, more specifically, other vehicles other than the autonomous vehicle using this method. It can be understood that the video data to be detected collected by means of an on-vehicle camera and the like contains multiple frame image data. Each frame image data may or may not contain a vehicle. And when the image data contains a vehicle, all the vehicles contained therein are used as the objects to be detected, and the speed of each vehicle as an object is detected separately.
[0076] The speed detection module 200 is used to input the feature map corresponding to the aligned first frame image sequence, the first frame image sequence, and the second frame image sequence into the speed recognition detection model, and obtain the recognition speed of the target object to be detected in the second frame image sequence output by the speed recognition detection model, that is, the recognition speed of the vehicle to be detected in the time frame corresponding to the second frame image sequence. Therefore, this recognition speed is the real-time speed of the vehicle and can be used as a parameter for adjusting the vehicle's autonomous driving strategy later, and also serves for the control and prediction in other aspects of autonomous driving (Prediction). Specifically, through some post-processing, combined with information such as the vehicle's position, distance, and speed, to determine which vehicles are key targets and which are non-key targets, and finally output the parameters of the key targets, and control and predict the autonomous driving process through the parameters of the above key targets.
[0077] The speed recognition detection model is trained based on the sample frame image sequences extracted from the sample video data; both the first frame image sequence and the second frame image sequence belong to the frame image sequences, and the first frame image sequence is adjacent to the second frame image sequence and is before the second frame image sequence, that is, the first frame image sequence is continuous multi-frame (at least two frames) of image data and is a part of the frame image sequences, and the second frame image sequence is also continuous multi-frame (at least two frames) of image data and is also a part of the frame image sequences; the feature map is output by inputting the first frame image sequence into the feature extraction model, and the feature map is aligned based on the first offset, and the first offset is output by inputting the feature map into the behavior recognition model.
[0078] In this device, both the speed recognition detection model and the feature extraction model are CNN models. CNN is essentially an MLP. CNN adopts the way of local connection and shared weights. On the one hand, it reduces the number of weights, making the network easy to optimize. On the other hand, it reduces the risk of overfitting. CNN is a kind of neural network. Its weight-sharing network structure makes it more similar to the biological neural network, reducing the complexity of the network model and the number of weights. This advantage is more obvious when the input of the network is multi-dimensional images, enabling the images to be directly used as the input data of the CNN model, avoiding the complex feature extraction and data reconstruction processes in traditional recognition algorithms. Currently, the CNN model has many advantages in two-dimensional image processing, such as the network can automatically extract image features including color, texture, shape, and the topological structure of the image; in dealing with two-dimensional image problems, especially in applications with good robustness and computing efficiency for recognizing displacement, scaling, and other forms of distortion invariance.
[0079] The CNN itself can adopt different combinations of neurons and learning rules. The CNN has some advantages that traditional technologies do not have: good fault tolerance, parallel processing ability, and self-learning ability. It can handle problems in complex environmental information, unclear background knowledge, and unclear inference rules, allowing large defects and distortions in samples, with fast running speed, good adaptability, and high resolution. The CNN model integrates the feature extraction function into the MLP through structural reorganization and weight reduction, omitting the complex image feature extraction process before recognition. The CNN model consists of an input layer, an output layer, and multiple hidden layers. The hidden layers can be divided into a convolutional layer, a pooling layer, a RELU layer, and a fully connected layer. Among them, the convolutional layer is the core of the CNN model. The parameters of the convolutional layer consist of a set of learnable filters or kernels, which have a small receptive field and extend to the entire depth of the input volume. During forward propagation, each filter convolves the input data, calculates the dot product between the filter and the input data, and generates a two-dimensional activation map of the filter. The input data is generally a two-dimensional vector, but may have a height. That is, the convolutional layer is used to convolve the input layer to extract higher-level features.
[0080] In this device, the feature map is the feature output (extracted) after filtering a certain layer in the convolutional layer of the feature extraction model. In this embodiment, specifically, it is the shared feature map output by the last layer of the convolutional layer of the feature extraction model.
[0081] In this device, the speed recognition detection model and the feature extraction model can be the same physical model.
[0082] The CNN model adopts the deep learning method. During the training process of the deep learning method, a prediction result is obtained from the input end to the output end. There is only one neural network in the middle. Comparing with the real result will get an error. This error will be passed through each layer in the CNN model, and each layer will make adjustments according to this error until the CNN model converges or reaches the expected effect. This is an end-to-end learning method, which can save the data annotation done before each independent learning task execution, improve the accuracy of the output result and save the learning cost.
[0083] The target object speed detection device of the present invention uses the continuous multi-frame image data contained in the collected original video data as the second frame image sequence to be predicted, and the continuous multi-frame image data contained in the original video data that is adjacent and before the second frame image sequence as the first frame image sequence; inputs the first frame image sequence into the feature extraction model trained by this device, and then extracts the corresponding feature maps. After aligning the feature maps, they are fused with the first frame image sequence and the second frame image sequence to be predicted and then input into the speed recognition detection model together, directly regressing to the recognition speed of the target object to be detected. Based on the end-to-end deep learning method, it optimizes the operation process of the entire pipeline, improves the speed detection accuracy and continuity while streamlining the post-processing process of speed detection, making the recognition speed detection of the target object achieve better performance.
[0084] The following combines Figure 4 to describe the target object speed detection device of the present invention. The speed detection module 200 specifically includes:
[0085] An extraction unit 210, configured to extract a first frame image sequence and a second frame image sequence from the frame image sequence.
[0086] A first input unit 220, configured to input the first frame image sequence into the feature extraction model, obtain the feature maps corresponding to each image in the first frame image sequence output by the feature extraction model, and use one of the feature maps as the reference feature map.
[0087] In this device, the feature extraction model is trained based on the sample frame image sequence extracted from the sample video data. In this embodiment, the feature extraction model adopts the supervised learning method during the deep learning process, that is, using the sample feature map as the label for supervised learning.
[0088] A second input unit 230, configured to input the feature maps corresponding to the first frame image sequence into the behavior recognition model based on the reference feature map, and obtain the first offset output by the behavior recognition model.
[0089] An alignment unit 240, configured to align the feature maps corresponding to the first frame image sequence to the reference feature map based on the first offset.
[0090] In this device, the behavior recognition model is trained based on the sample feature map and the reference sample feature map. The role of the behavior recognition model is to learn an offset map to align the feature maps of multiple consecutive frames. Specifically, the input data of the behavior recognition model is the feature maps of multiple consecutive frames, that is, the feature maps corresponding to each image in the first frame image sequence, and the output data is the offset amount required to move other feature maps to the reference feature map, that is, the first offset amount. Through the first offset amount, the feature maps corresponding to the first frame image sequence can be aligned to the reference feature map.
[0091] The third input unit 250 is configured to fuse the aligned feature maps, the first frame image sequence, and the second frame image sequence and input them into the speed recognition detection model to obtain the recognition speed directly regressed by the speed recognition detection model. In this embodiment, the speed recognition detection model adopts the supervised learning method during the deep learning process, that is, the recognition speed corresponding to the fused data is used as a label for supervised learning.
[0092] The feature extraction model in this device is trained through the following modules:
[0093] The second acquisition module A100 is configured to acquire sample video data and extract the third frame image sequence from the sample video data;
[0094] The first input module 200 is configured to use the third frame image sequence as the input data after training, and perform training in the deep learning manner to obtain a feature extraction model for generating the feature maps of the frame image sequence to be predicted.
[0095] The speed recognition detection model in this device is trained through the following modules:
[0096] The extraction module A300 is configured to extract the sample frame image sequence from the sample video data, and extract the third frame image sequence and the fourth frame image sequence from the sample frame image sequence. Among them, both the third frame image sequence and the fourth frame image sequence belong to the sample frame image sequence, and the third frame image sequence is adjacent to the fourth frame image sequence and is before the fourth frame image sequence.
[0097] The second input module A400 is configured to input the third frame image sequence into the feature extraction model to obtain the sample feature maps corresponding to each image in the third frame image sequence output by the feature extraction model, and use one of the sample feature maps as the reference sample feature map.
[0098] The third input module A500 is configured to input the sample feature maps corresponding to the third frame image sequence into the behavior recognition model based on the reference sample feature map to obtain the second offset amount output by the behavior recognition model;
[0099] Alignment module A600, configured to align the sample feature maps corresponding to the third frame image sequence to the reference sample feature map based on the second offset;
[0100] Fourth input module A700, configured to fuse the aligned sample feature maps, the third frame image sequence, and the fourth frame image sequence as input data for training, and perform training in a deep learning manner to obtain a speed recognition detection model for generating the recognition speed of the frame image sequence to be predicted. In the fourth input module A700, during the deep learning process, the recognition speed corresponding to the fused data can be used as a label for supervised learning.
[0101] The behavior recognition model in this device is trained through the following steps:
[0102] Fifth input module A800, configured to use the reference sample feature map as the alignment reference, and use the sample feature maps corresponding to the third frame image sequence as input data for training, and perform training in a deep learning manner to obtain a behavior recognition model for generating the offset of the feature map to be aligned.
[0103] Figure 5 An example of the physical structure diagram of an electronic device is shown as Figure 5 shown. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete communication with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the target object speed detection method, and this method includes the following steps:
[0104] S100, collect the video data to be detected, and extract the frame image sequence containing the target object to be detected from the video data to be recognized;
[0105] S200, input the feature maps corresponding to the aligned first frame image sequence, the first frame image sequence, and the second frame image sequence into the speed recognition detection model, and obtain the recognition speed of the target object to be detected corresponding to the second frame image sequence output by the speed recognition detection model;
[0106] Among them, the speed recognition and detection model is trained based on a sample frame image sequence extracted from sample video data. The first frame image sequence and the second frame image sequence both belong to the frame image sequence, and the first frame image sequence is adjacent to and before the second frame image sequence. The feature map is obtained by inputting the first frame image sequence into a feature extraction model, and the feature map is aligned based on a first offset amount, and the first offset amount is obtained by inputting the feature map into a behavior recognition model.
[0107] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0108] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the target object speed detection method provided by the above-mentioned various methods. The method includes the following steps:
[0109] S100. Collect video data to be detected, and extract a frame image sequence including the target object to be detected from the video data to be recognized;
[0110] S200. Input the feature map corresponding to the aligned first frame image sequence, the first frame image sequence, and the second frame image sequence into the speed recognition and detection model to obtain the recognition speed of the target object to be detected corresponding to the second frame image sequence output by the speed recognition and detection model;
[0111] Among them, the speed recognition and detection model is trained based on a sample frame image sequence extracted from sample video data. The first frame image sequence and the second frame image sequence both belong to the frame image sequence, and the first frame image sequence is adjacent to the second frame image sequence and is located before the second frame image sequence. The feature map is obtained by inputting the first frame image sequence into a feature extraction model, and the feature map is aligned based on a first offset amount, and the first offset amount is obtained by inputting the feature map into a behavior recognition model.
[0112] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the target object speed detection method provided by the above-mentioned various methods. The method includes the following steps:
[0113] S100. Collect video data to be detected, and extract a frame image sequence including the target object to be detected from the video data to be recognized;
[0114] S200. Input the feature map corresponding to the aligned first frame image sequence, the first frame image sequence, and the second frame image sequence into the speed recognition and detection model to obtain the recognition speed of the target object to be detected corresponding to the second frame image sequence output by the speed recognition and detection model;
[0115] Among them, the speed recognition and detection model is trained based on a sample frame image sequence extracted from sample video data. The first frame image sequence and the second frame image sequence both belong to the frame image sequence, and the first frame image sequence is adjacent to the second frame image sequence and is located before the second frame image sequence. The feature map is obtained by inputting the first frame image sequence into a feature extraction model, and the feature map is aligned based on a first offset amount, and the first offset amount is obtained by inputting the feature map into a behavior recognition model.
[0116] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting the speed of an object, characterized in that, it includes the following steps: Collect video data to be detected, and extract a frame image sequence containing the object to be detected from the video data to be detected; Input the feature map corresponding to the aligned first frame image sequence, the first frame image sequence, and the second frame image sequence into the speed recognition detection model, and obtain the recognition speed of the object to be detected output by the speed recognition detection model corresponding to the second frame image sequence; Among them, the speed recognition detection model is trained based on the sample frame image sequence extracted from the sample video data. The first frame image sequence and the second frame image sequence both belong to the frame image sequence, and the first frame image sequence is adjacent to the second frame image sequence and is located before the second frame image sequence. The feature map is output by inputting the first frame image sequence into the feature extraction model, and the feature map is aligned based on the first offset amount, and the first offset amount is output by inputting the feature map into the behavior recognition model; Among them, the speed recognition detection model is trained through the following steps: Extract the sample frame image sequence from the sample video data, and extract the third frame image sequence and the fourth frame image sequence from the sample frame image sequence; among them, the third frame image sequence and the fourth frame image sequence both belong to the sample frame image sequence, and the third frame image sequence is adjacent to the fourth frame image sequence and is located before the fourth frame image sequence; Input the third frame image sequence into the feature extraction model, and obtain the sample feature maps corresponding to each image in the third frame image sequence output by the feature extraction model, and use one of the sample feature maps as the reference sample feature map; Based on the reference sample feature map, input the sample feature map corresponding to the third frame image sequence into the behavior recognition model, and obtain the second offset amount output by the behavior recognition model; Based on the second offset amount, align the sample feature maps corresponding to the third frame image sequence to the reference sample feature map; Fuse the aligned sample feature maps, the third frame image sequence, and the fourth frame image sequence as the input data for training, and use deep learning to train to obtain the speed recognition detection model for generating the recognition speed of the frame image sequence to be predicted.
2. The method for detecting the speed of an object according to claim 1, characterized in that, The step of inputting the feature map corresponding to the aligned first frame image sequence, the first frame image sequence, and the second frame image sequence into the speed recognition detection model, and obtaining the recognition speed of the object to be detected output by the speed recognition detection model corresponding to the second frame image sequence specifically includes the following steps: Extract the first frame image sequence and the second frame image sequence from the frame image sequence; Input the first frame image sequence into the feature extraction model to obtain the feature maps corresponding to each image in the first frame image sequence output by the feature extraction model, and use one of the feature maps as the reference feature map; wherein, the feature extraction model is trained based on the sample frame image sequence extracted from the sample video data; Based on the reference feature map, input the feature maps corresponding to the first frame image sequence into the behavior recognition model to obtain the first offset output by the behavior recognition model; wherein, the behavior recognition model is trained based on the sample feature maps and the reference sample feature maps; Based on the first offset, align the feature maps corresponding to the first frame image sequence to the reference feature map; After fusing the aligned feature maps, the first frame image sequence, and the second frame image sequence, input them into the speed recognition and detection model to obtain the recognition speed output by the speed recognition and detection model.
3. The method for detecting the speed of an object according to claim 2, wherein, the feature extraction model is trained through the following steps: Collect the sample video data, and extract the third frame image sequence from the sample video data; Use the third frame image sequence as the input data after training, and perform training in a deep learning manner to obtain the feature extraction model for generating the feature maps of the frame image sequence to be predicted.
4. The method for detecting the speed of an object according to claim 1, wherein, the feature extraction model and the speed recognition and detection model are the same physical model.
5. The method for detecting the speed of an object according to claim 1, wherein, the behavior recognition model is trained through the following steps: Use the reference sample feature map as the alignment reference, use the sample feature maps corresponding to the third frame image sequence as the input data for training, and perform training in a deep learning manner to obtain the behavior recognition model for generating the offsets of the feature maps to be aligned.
6. An apparatus for detecting the speed of an object, wherein, comprising: A first acquisition module, configured to acquire the video data to be detected, and extract the frame image sequence including the object to be detected from the video data to be detected; A speed detection module, configured to input the feature maps corresponding to the aligned first frame image sequence, the first frame image sequence, and the second frame image sequence into the speed recognition and detection model to obtain the recognition speed of the object to be detected corresponding to the second frame image sequence output by the speed recognition and detection model; Among them, the speed recognition detection model is trained based on a sample frame image sequence extracted from sample video data. The first frame image sequence and the second frame image sequence both belong to the frame image sequence, and the first frame image sequence is adjacent to the second frame image sequence and is located before the second frame image sequence. The feature map is obtained by inputting the first frame image sequence into a feature extraction model, and the feature map is aligned based on a first offset. The first offset is obtained by inputting the feature map into a behavior recognition model; Among them, the speed recognition detection model is trained through the following steps: Extract the sample frame image sequence from the sample video data, and extract a third frame image sequence and a fourth frame image sequence from the sample frame image sequence. Among them, the third frame image sequence and the fourth frame image sequence both belong to the sample frame image sequence, and the third frame image sequence is adjacent to the fourth frame image sequence and is located before the fourth frame image sequence; Input the third frame image sequence into the feature extraction model to obtain sample feature maps respectively corresponding to each image in the third frame image sequence output by the feature extraction model, and use one of the sample feature maps as a reference sample feature map; Based on the reference sample feature map, input the sample feature maps corresponding to the third frame image sequence into the behavior recognition model to obtain a second offset output by the behavior recognition model; Based on the second offset, align the sample feature maps corresponding to the third frame image sequence to the reference sample feature map; Fuse the aligned sample feature maps, the third frame image sequence, and the fourth frame image sequence as input data for training, and use deep learning methods for training to obtain the speed recognition detection model for generating a frame image sequence to be predicted to recognize the speed.
7. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the program, the steps of the target object speed detection method according to any one of claims 1 to 5 are implemented.
8. A non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, the steps of the target object speed detection method according to any one of claims 1 to 5 are implemented.
9. A computer program product, including a computer program, characterized in that, when the computer program is executed by a processor, the steps of the target object speed detection method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Monitoring terminal and road traffic early warning method and system
CN108492567A
Vehicle speed determination method and vehicle
CN111986472A