Vehicle target detection method and system based on LSTM and DETR

By combining LSTM and DETR models, processing video streams and image sequences, generating vehicle object detection results, and optimizing them with SVMD adaptive algorithm, the problems of missed and missed detection in complex scenarios in the prior art are solved, and more efficient and accurate vehicle object detection is achieved.

CN120032294APending Publication Date: 2025-05-23JIANGSU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510109857.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing vehicle object detection technology is susceptible to background interference in complex scenarios, resulting in missed detection and missed detection, and when processing video sequences, the modeling ability of timing changes is insufficient.

Method used

Using a vehicle object detection method based on LSTM and DETR, by inputting video streams and image sequences into the DETR architecture, feature extraction, Transformer encoding, LSTM timing fusion and Transformer decoding are performed, preliminary detection results are generated, and SVMD adaptive algorithm is used for determination and optimization.

Benefits of technology

It improves the accuracy and robustness of vehicle target detection, reduces missed detection and missed detection, especially in the case of complex background or changes in light, it can effectively capture dynamic changes and global features, and improves the accurate identification ability of vehicle targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032294A_ABST
    Figure CN120032294A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle target detection method and system based on LSTM and DETR, and relates to the technical field of target detection, and the method comprises the steps: taking a video stream containing vehicle information and an image sequence as input data, and carrying out the preprocessing of the input data, and obtaining an image frame sequence; processing the image frame sequence by constructing a deep learning model combining LSTM and DETR to obtain a preliminary detection result of the vehicle; and determining the preliminary detection result by constructing an SVMD adaptive algorithm to obtain a target detection result, and optimizing and outputting the target detection result. The SVMD-LSTM + DETR-based vehicle target detection model provides an efficient and accurate vehicle target detection solution, is suitable for intelligent traffic and automatic driving systems, and provides support for real-time decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a vehicle target detection method and system based on LSTM (Long Short-Term Memory Network) and DETR (Detection Transformer). Background Art

[0002] In recent years, the DETR model and its variants have received extensive attention and application in the field of target detection. The DETR model and Transformer architecture have shown broad application prospects and strong performance advantages in the field of target detection. With the continuous advancement of technology, existing vehicle target detection technology is committed to solving multiple challenges, such as different lighting and weather conditions, complex backgrounds, and real-time issues.

[0003] However, traditional target detection models are easily affected by background interference in complex scenes, especially when the contrast between vehicle targets and background is low, and missed detection and false detection often occur; at the same time, existing technologies can effectively capture the features of static images, but when processing video sequences, the modeling ability of temporal changes is still insufficient, which may lead to missed detection or wrong judgment in dynamic scenes, especially when detecting moving targets. Therefore, it is of great significance to construct an efficient vehicle target detection method. Summary of the invention

[0004] In view of the shortcomings in the prior art, the present invention provides a vehicle target detection method and system based on LSTM and DETR.

[0005] The present invention achieves the above technical objectives through the following technical means.

[0006] A vehicle target detection method based on LSTM and DETR, comprising:

[0007] The video stream and image sequence containing vehicle information are used as input data, and the input data is preprocessed to obtain an image frame sequence;

[0008] Construct a DETR architecture and perform feature processing on the image frame sequence to generate preliminary detection results. Feature processing includes feature extraction, Transformer encoding, LSTM time series fusion, and Transformer decoding.

[0009] The SVMD adaptive algorithm is used to determine the preliminary detection results to obtain the target detection results, and the target detection results are optimized and output.

[0010] Furthermore, the image frame sequence is obtained specifically as follows:

[0011] The video stream and image sequence containing vehicle information are decoded into an image frame sequence, and each frame of the image is scaled to a preset size; N frames of the image are extracted to construct a time series, where N is a preset number.

[0012] Furthermore, the specific steps of feature extraction are:

[0013] The image frame sequence is input into the convolutional neural network, and convolution and pooling operations are performed on each frame of the image to obtain the feature map of the image frame sequence; at the same time, the spatial structure information of the feature map is retained to express the characteristics of the vehicle position and relative relationship in the image frame sequence.

[0014] Furthermore, the feature map includes edge information, corner features, texture features, color histograms, shape features and statistical features.

[0015] Furthermore, the spatial structure information includes positional relationships, local context information, vehicle distance and layout, and boundary information.

[0016] Furthermore, the specific steps of the Transformer encoding are:

[0017] The feature map is converted into a single sequence as a feature matrix through a matrix and input into the Transformer encoder; the Transformer encoder performs position encoding on the feature matrix and uses the self-attention mechanism to capture the relative relationship of each feature in the feature matrix to obtain the attention output feature.

[0018] Furthermore, the specific steps of the LSTM time series fusion are:

[0019] The attention output features are input into the long short-term memory network, and the attention output features are processed frame by frame through the forget gate, input gate and output gate; by integrating the spatial structure information of each frame with the temporal relationship between consecutive frames, the temporal enhancement features are obtained.

[0020] Furthermore, the specific steps of the Transformer decoding are:

[0021] The temporal enhanced features are input into the Transformer decoder, which decodes the input features through the self-attention mechanism and feedforward network to generate preliminary detection results; the preliminary detection results include the vehicle's location, classification information, and classification confidence, where the classification information includes vehicle type, status information, functional attributes, and color information.

[0022] Furthermore, the specific steps of obtaining, optimizing and outputting the target detection results are:

[0023] The preliminary detection results are input into the support vector machine model constructed by SVMD, and the information density of the vehicle's location and classification information is analyzed through the support vector machine model; the classification threshold is adaptively adjusted according to the information density; the preliminary detection results are judged using the classification threshold and classification confidence, and the preliminary detection results that exceed the classification threshold are regarded as valid targets; the valid targets are optimized to obtain the detection results, and the detection results are output.

[0024] A vehicle target detection system based on LSTM and DETR is used to implement a vehicle target detection method based on LSTM and DETR, including:

[0025] Data processing module: takes the video stream and image sequence containing vehicle information as input data, and pre-processes the input data to obtain an image frame sequence;

[0026] DETR module: connected with the data processing module, constructs the DETR architecture to perform feature processing on the image frame sequence to generate preliminary detection results;

[0027] SVMD judgment module: connected to the DETR module, uses the SVMD adaptive algorithm to determine the preliminary detection results to obtain the target detection results, and optimizes and outputs the target detection results.

[0028] The beneficial effects of the present invention are:

[0029] (1) The present invention can effectively improve the accuracy and robustness of vehicle target detection by combining the long short-term memory network with the DETR model. LSTM has outstanding advantages in processing time series data and can capture dynamic changes in image frame sequences, thereby improving the detection of moving targets. The DETR model uses global feature modeling and a self-attention mechanism to deeply analyze image features, especially in capturing long-distance dependencies. The organic combination of the two enables the system to consider both static features and dynamic time series information when identifying vehicle targets, thereby improving the ability to accurately identify vehicle targets, especially in the case of complex backgrounds or changing lighting, reducing missed detections and false detections.

[0030] (2) The present invention adopts the SVMD (Successive Variational Mode Decomposition) adaptive algorithm to dynamically adjust the classification threshold, thereby significantly improving the accuracy and robustness of target judgment. Traditional target detection methods usually rely on fixed thresholds for target judgment. This method may lead to unstable prediction results due to environmental changes in complex scenarios, and is prone to misjudgment or missed judgment. The SVMD algorithm can flexibly adjust the classification criteria based on real-time data by analyzing the information density of the vehicle's location and classification information. In the case of dense vehicles or complex backgrounds, SVMD can increase the threshold to ensure that only highly reliable detection results are retained, reducing the occurrence of false detections. When the target distribution is relatively sparse or the environment is relatively simple, the threshold can be lowered to ensure that as many targets as possible are detected. In this way, efficient and reliable vehicle target detection results can be provided in a variety of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 A flow chart of a vehicle target detection method based on LSTM and DETR according to the present invention;

[0032] Figure 2 This is a system block diagram of a vehicle target detection system based on LSTM and DETR described in the present invention. DETAILED DESCRIPTION

[0033] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments, but the protection scope of the present invention is not limited thereto.

[0034] Example 1

[0035] like Figure 1 As shown, this embodiment provides a vehicle target detection method based on LSTM and DETR, including:

[0036] S1, taking a video stream and an image sequence containing vehicle information as input data, and preprocessing the input data to obtain an image frame sequence;

[0037] S2. Build a DETR architecture and perform feature processing on the image frame sequence to generate preliminary detection results; the feature processing includes feature extraction, Transformer encoding, LSTM time series fusion and Transformer decoding;

[0038] S3. Use the SVMD adaptive algorithm to determine the preliminary detection results to obtain the target detection results, and optimize and output the target detection results.

[0039] As described in the above steps S1-S3, the present invention effectively improves the accuracy and robustness of vehicle target detection by combining the long short-term memory network with the DETR model. LSTM has outstanding advantages in processing time series data, and can capture dynamic changes in image frame sequences, thereby improving the detection of moving targets; the DETR model uses global feature modeling and self-attention mechanism to deeply analyze image features, especially in capturing long-distance dependencies; the organic combination of the two enables the system to consider both static features and dynamic time series information when identifying vehicle targets, so as to improve the ability to accurately identify vehicle targets, especially in the case of complex backgrounds or changes in lighting, reducing missed detections and false detections. At the same time, the SVMD adaptive algorithm is used to dynamically adjust the classification threshold. Traditional target detection methods usually rely on fixed thresholds for target judgment. This method will lead to unstable prediction results due to environmental changes in complex scenarios, and is prone to misjudgment or missed judgment. The SVMD algorithm can flexibly adjust the classification criteria based on real-time data by analyzing the location of the vehicle and the information density of the classification information. In the case of dense vehicles or complex backgrounds, SVMD can increase the threshold to ensure that only highly reliable detection results are retained, reducing the occurrence of false detections. When the target distribution is relatively sparse or the environment is relatively simple, the threshold can be lowered to ensure that as many targets as possible are detected. In this way, efficient and reliable vehicle target detection results can be provided in a variety of application scenarios.

[0040] In step S1, the input data is preprocessed to obtain an image frame sequence, including:

[0041] S11, decoding the video stream and image sequence containing vehicle information into an image frame sequence, and scaling each frame of the image to a preset size;

[0042] S12, extracting N frames of images to construct a time series, where N is a preset number.

[0043] As described in the above steps S11-S12, a standard video decoder (such as FFmpeg or OpenCV) is used to process the video stream and convert it into a series of static image frame sequences. Each frame of the decoded image is scaled and uniformly adjusted to a preset size, wherein the preset size can be set to 640×480 or 512×512 to ensure the calculation consistency of subsequent processing steps; at the same time, the scaling process uses an interpolation algorithm, such as bilinear interpolation or bicubic interpolation, to maintain image quality, and extracts N consecutive frames from the decoded image frame sequence to construct a time series, wherein N is a preset number, which is set based on the accuracy requirement. If the accuracy requirement is high, the preset number is set to be high, and if the accuracy requirement is low, the preset number is set to be low, usually 4 to 10 frames, and the extracted N frames of images are combined into a multidimensional array in chronological order for subsequent input to the processing module.

[0044] The feature extraction in step S2 includes:

[0045] Input the image frame sequence into the convolutional neural network, perform convolution and pooling operations on each frame of the image, and obtain a feature map of the image frame sequence; wherein the feature map includes edge information, corner point features, texture features, color histograms, shape features, and statistical features;

[0046] At the same time, the spatial structure information of the feature map is retained to express the characteristics of the vehicle position and relative relationship in the image frame sequence; the spatial structure information includes position relationship, local context information, vehicle distance and layout, and boundary information;

[0047] As described above, the preprocessed image frame sequence is used as input and passed into a convolutional neural network, such as VGG, ResNet or EfficientNet architecture. Multiple convolution layers are applied to each frame of the image. These layers use convolution filters of different sizes (such as 3×3, 5×5) to extract multi-level features. Each layer can set a different number of convolution channels to capture more subtle features. After the convolution operation, a pooling layer (such as maximum pooling or average pooling) is used to reduce the size of the feature map, reduce the computational complexity and extract important features, and retain the most significant edge and shape information. The construction of the feature map involves multiple feature extraction modules, each of which focuses on capturing a specific feature type, such as extracting edge features through an edge detection algorithm, identifying corners in an image using Harris corner detection or FAST algorithm, applying texture descriptors (such as LBP, Gabor filters) to extract subtle texture features, calculating the color histogram of the image to obtain color distribution features, extracting the shape features of the object through a contour detection algorithm, and calculating basic statistical features (such as mean, variance, etc.) to describe the distribution of image brightness or color. When the feature map is generated, the spatial structure information of the image is retained using indexes or other data structures (such as Tensor) to facilitate the subsequent position and relative relationship analysis of the target. The feature map of each layer is associated with the corresponding spatial structure information to maintain the spatial coordinate relationship of the feature points, where the spatial structure information includes position relationship, local context information, vehicle distance and layout, and boundary information. In actual cases: by calculating the distance between feature points and their relative positions, the relationship between different vehicles is analyzed; by examining the pixel distribution of the local neighborhood of the feature, the group behavior of vehicles in a specific context is identified to obtain local context information; by estimating the distance between each target, their layout pattern is determined to distinguish whether they belong to the same trajectory or interfere with each other; through the segmentation algorithm, the boundary information of a specific target is obtained to enhance the recognition accuracy. Based on the above processing, information containing multiple features can be extracted from the image frame sequence, and the spatial structure information can be effectively retained, thereby providing a strengthening basis for subsequent feature encoding and time series fusion.

[0048] The Transformer encoding in step S2 includes:

[0049] The feature map is converted into a single sequence as a feature matrix through a matrix and input into the Transformer encoder;

[0050] The Transformer encoder performs position encoding on the feature matrix and uses the self-attention mechanism to capture the relative relationship of each feature in the feature matrix to obtain the attention output feature;

[0051] As described above, the feature map processed by the convolutional neural network is flattened to convert the multi-dimensional tensor into a one-dimensional array to form a feature matrix. In the actual case, it is assumed that the size of the feature map is C×H×W, where C is the number of channels, H is the height, and W is the width; the feature matrix obtained after flattening will have dimensions N×(C·H·W), where N is the number of image frames; during the flattening process, the batch dimension is also added to form a four-dimensional input tensor so that the Transformer encoder can process multi-frame information. According to the input requirements of the Transformer, the dimension of the feature matrix is ​​adjusted to be the same as the (Z, T, D) dimension, where Z is the batch size used to interpret the number of image frames in the input Transformer, T is the number of frame steps, and D is the dimension of the flattened feature for each frame step; the flattened feature matrix is ​​input to the Transformer encoder, and position encoding is added to each feature in the feature matrix to compensate for the lack of position encoding in the Transformer model. To solve the problem of lack of position information, the position encoding can use sine, cosine function or learned position encoding, and the dimension of each feature vector is consistent with D, where the feature vector is a high-dimensional numerical array extracted from the feature map. The generated position encoding is added to the feature matrix to obtain an enhanced feature matrix containing position information; then the enhanced feature matrix is ​​calculated through the Transformer encoder through the self-attention mechanism. In actual cases, each feature is linearly transformed to generate query (Q), key (K), and value (V) vectors, and the dot product of Q and K is calculated to generate an attention score, which is normalized by the softmax function to obtain the attention weight; the obtained attention weight is used to perform weighted summation on the value vector (V) to output a new weighted feature representation for the input feature vector as the attention output feature, reflecting the fusion of the relative relationship between the input features.

[0052] The LSTM time series fusion in step S2 includes:

[0053] Input the attention output features into the long short-term memory network, and process the attention output features frame by frame through the forget gate, input gate, and output gate;

[0054] By integrating the spatial structure information of each frame and the temporal relationship between consecutive frames, the temporal enhancement feature is obtained;

[0055] As mentioned above, the dimension of the attention output feature is adjusted to a format suitable for LSTM processing. Assume that the dimension of the attention output feature is (S, M, P), where S is the batch size used to interpret the dimension of the attention output feature, M is the number of time steps (number of image frames), and P is the feature dimension. LSTM needs to convert the input data into a three-dimensional tensor, including batch dimension, time dimension, and feature dimension, and input the adjusted attention output feature into the LSTM unit for processing. LSTM has three main working gates, including forget gate, input gate, and output gate; among them, the forget gate is used to decide which part of the old information to retain. It reads the current input and the previous hidden state through the sigmoid function to generate a vector with a value between 0 and 1; corresponding to the degree of forgetting of each value, the input gate is used to decide whether to store the current frame information in the memory unit, and selects which information needs to be updated through the sigmoid function, and the tanh function generates new candidate values; the output gate is used to decide which information to output from the LSTM unit. The current input and the previous hidden state are calculated by the sigmoid function, and the output is generated in combination with the current memory unit state. In the LSTM unit, the current hidden state and memory unit state are updated through the calculation of the forget gate, input gate and output gate. Through the mechanism of these gates, the system can effectively maintain and transmit important timing information; then the spatial structure information after each frame processing is integrated with the output features of the current LSTM unit. The current LSTM unit output can be combined with the corresponding spatial information using weighted averaging, splicing or other fusion methods to form rich timing enhancement features. This feature sequence can more comprehensively reflect the dynamic information of the current frame and the contextual information of the previous frame.

[0056] The Transformer decoding in step S2 includes:

[0057] The temporal enhancement features are input into the Transformer decoder, which decodes the input features through the self-attention mechanism and feedforward network to generate preliminary detection results; the preliminary detection results include the vehicle's location, classification information, and classification confidence, where the classification information includes but is not limited to vehicle type, status information, functional attributes, and color information;

[0058] As mentioned above, the dimensions of the temporal enhancement features output from LSTM are adjusted to a format suitable for processing by the Transformer decoder and input to the Transformer decoder. In the Transformer decoder, the self-attention mechanism is first used to calculate the relationship weights for the features of each time step. Each feature vector generates query (Q), key (K), and value (V) vectors through linear transformation, where the feature vector represents a high-dimensional numerical array extracted from the input features. Then the dot product of Q and K, i.e., the attention score, is calculated, and the weight is obtained by softmax normalization. The V vector is weighted and summed using the attention weights. Next, the features obtained by weighted summation are input into the feedforward neural network for training to obtain preliminary detection results.

[0059] In step S3, the SVMD adaptive algorithm is used to determine the preliminary detection result to obtain the target detection result, and the target detection result is optimized and output, including:

[0060] S31, inputting the preliminary detection results into the support vector machine model constructed by SVMD, and analyzing the location of the vehicle and the information density of the classification information through the support vector machine model;

[0061] S32, adaptively adjusting the classification threshold according to information density;

[0062] S33, using the classification threshold and the classification confidence to judge the preliminary detection results, and taking the preliminary detection results exceeding the classification threshold as valid targets;

[0063] S34, optimizing the effective target to obtain the detection result, and outputting the detection result;

[0064] As described in the above steps S31-S34, the vehicle position, classification information and classification confidence in the preliminary detection results are extracted as feature vectors. This feature vector is usually in the format of Features = [xcenter, ycenter, width, height, class_type, state, confidence], where xcenter and ycenter are the coordinates of the center of the bounding box, width and height are the size of the bounding box, class_type represents the vehicle type, state represents the state information, and confidence is the classification confidence; the extracted feature vector is input into the support vector machine model constructed by SVMD to analyze the information density of all vehicle detection results. The SVMD model analyzes the density value of each feature vector and uses kernel density estimation or other statistical methods to calculate the information density of each detection result in the feature space, where the information density provides information about the distribution of the detection results in the region. According to the calculated information density, the classification threshold is dynamically adjusted. For example, in the case of dense targets, the threshold is increased to reduce false detections. Conversely, in the case of sparse targets, the threshold is lowered to ensure that more valid targets are detected. Finally, the adaptive classification threshold is used to judge the preliminary detection results, and the classification confidence of each vehicle is traversed. If the confidence of a detection result exceeds the current classification threshold, it is marked as a valid target. All targets judged to be valid are summarized, including their location, classification information and confidence. At the same time, algorithms such as non-maximum suppression are used for valid targets to reduce redundant detections, merge them into a more accurate bounding box, obtain the detection results, and output the detection results.

[0065] Example 2

[0066] like Figure 2 As shown, this embodiment provides a vehicle target detection system based on LSTM and DETR, including:

[0067] Data processing module: takes the video stream and image sequence containing vehicle information as input data, and pre-processes the input data to obtain an image frame sequence;

[0068] DETR module: connected with the data processing module, constructs the DETR architecture to perform feature processing on the image frame sequence to generate preliminary detection results, where feature processing includes feature extraction, Transformer encoding, LSTM time series fusion and Transformer decoding;

[0069] SVMD judgment module: connected to the DETR module, uses the SVMD adaptive algorithm to determine the preliminary detection results to obtain the target detection results, and optimizes and outputs the target detection results.

[0070] The embodiments are preferred implementations of the present invention, but the present invention is not limited to the above-mentioned implementations. Any obvious improvements, substitutions or modifications that can be made by those skilled in the art without departing from the essential content of the present invention belong to the protection scope of the present invention.

Claims

1. A vehicle target detection method based on LSTM and DETR, characterized by: The video stream and image sequence containing vehicle information are used as input data, and the input data is preprocessed to obtain an image frame sequence; Construct a DETR architecture and perform feature processing on the image frame sequence to generate preliminary detection results. Feature processing includes feature extraction, Transformer encoding, LSTM time series fusion, and Transformer decoding. The SVMD adaptive algorithm is used to determine the preliminary detection results to obtain the target detection results, and the target detection results are optimized and output.

2. The vehicle target detection method based on LSTM and DETR according to claim 1, characterized in that: The image frame sequence is obtained specifically as follows: The video stream and image sequence containing vehicle information are decoded into an image frame sequence, and each frame of the image is scaled to a preset size; N frames of the image are extracted to construct a time series, where N is a preset number.

3. The vehicle target detection method based on LSTM and DETR according to claim 1, characterized in that: The specific steps of feature extraction are: The image frame sequence is input into the convolutional neural network, and convolution and pooling operations are performed on each frame of the image to obtain the feature map of the image frame sequence; at the same time, the spatial structure information of the feature map is retained to express the characteristics of the vehicle position and relative relationship in the image frame sequence.

4. The vehicle target detection method based on LSTM and DETR according to claim 3, characterized in that: The feature map includes edge information, corner features, texture features, color histograms, shape features, and statistical features.

5. The vehicle target detection method based on LSTM and DETR according to claim 3, characterized in that: The spatial structure information includes positional relationships, local context information, vehicle distances and layout, and boundary information.

6. The vehicle target detection method based on LSTM and DETR according to claim 3, characterized in that: The specific steps of the Transformer encoding are: The feature map is converted into a single sequence as a feature matrix through a matrix and input into the Transformer encoder; the Transformer encoder performs position encoding on the feature matrix and uses the self-attention mechanism to capture the relative relationship of each feature in the feature matrix to obtain the attention output feature.

7. The vehicle target detection method based on LSTM and DETR according to claim 6, characterized in that: The specific steps of the LSTM time series fusion are: The attention output features are input into the long short-term memory network, and the attention output features are processed frame by frame through the forget gate, input gate and output gate; by integrating the spatial structure information of each frame with the temporal relationship between consecutive frames, the temporal enhancement features are obtained.

8. The vehicle target detection method based on LSTM and DETR according to claim 7, characterized in that: The specific steps of the Transformer decoding are: The temporal enhanced features are input into the Transformer decoder, which decodes the input features through the self-attention mechanism and feedforward network to generate preliminary detection results; the preliminary detection results include the vehicle's location, classification information, and classification confidence, where the classification information includes vehicle type, status information, functional attributes, and color information.

9. The vehicle target detection method based on LSTM and DETR according to claim 8, characterized in that: The specific steps of obtaining, optimizing and outputting the target detection results are: The preliminary detection results are input into the support vector machine model constructed by SVMD, and the information density of the vehicle's location and classification information is analyzed through the support vector machine model; the classification threshold is adaptively adjusted according to the information density; The classification threshold and classification confidence are used to judge the preliminary detection results, and the preliminary detection results that exceed the classification threshold are regarded as valid targets; the valid targets are optimized to obtain the detection results, and the detection results are output.

10. A vehicle target detection system based on LSTM and DETR, used to implement the vehicle target detection method based on LSTM and DETR according to any one of claims 1 to 9, characterized in that: include: Data processing module: takes the video stream and image sequence containing vehicle information as input data, and pre-processes the input data to obtain an image frame sequence; DETR module: connected with the data processing module, constructs the DETR architecture to perform feature processing on the image frame sequence to generate preliminary detection results; SVMD judgment module: connected to the DETR module, uses the SVMD adaptive algorithm to determine the preliminary detection results to obtain the target detection results, and optimizes and outputs the target detection results.

Citation Information

Cited By

  • Transform-based ultraviolet plume target detection method

    CN120726309A