High-speed target detection and tracking method using infrared event information
By using a dual-mode infrared camera to obtain infrared event information and image grayscale information in the satellite system, and using a deep convolutional neural network of space-time-optical and signal fusion for processing, the problem of effective detection of low frame rates in the satellite system is solved, and efficient detection and tracking of high-speed moving targets is achieved.
Patent Information
- Application Number
- CN202510273121.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-27
AI Technical Summary
The existing information processing resources of the satellite system are limited, and the effective detection frame rate is often lower than 30Hz@1080p resolution images, which seriously restricts the system's effective detection and tracking capabilities of high-speed moving targets.
The infrared event detector with a dual-mode infrared camera or a common optical path and the infrared image detector simultaneously obtain infrared event information and image grayscale information, and process the information data through a deep convolutional neural network of space-time-light and signal fusion, extract local position and feature information of a specific target, and achieve effective tracking of the target.
By increasing the frame rate of infrared event information and image grayscale information, the detection time resolution of space-based infrared systems is improved, the detection capability of high-speed weak targets is enhanced, and more accurate and fast target detection and tracking is achieved.
Smart Images

Figure CN120047676A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of infrared remote sensing, and relates to a high-speed target detection and tracking method. Background Art
[0002] Due to the extremely small size of the targets in space-based detection and the long detection distance, the energy intensity captured during detection is very weak, and the effective information is often submerged in the background information and the noise of the detection device itself. Therefore, when detecting these "dark and weak" targets, it is very easy to cause false alarms, missed detections and mistracking, and the stability and reliability of the detection system are difficult to guarantee. Traditional space-based detection methods use equal intervals to trigger the continuous acquisition of infrared images at the current moment, and analyze the acquired images in real time to achieve long-term uninterrupted detection and tracking of "dark and weak" targets. This method requires a large amount of redundant background information to be processed from massive visual data for a long time, and due to the limited information processing resources of the onboard system, the effective detection frame rate is often lower than 30Hz@1080p resolution image, which seriously restricts the system's effective detection and tracking capabilities for high-speed moving targets. Summary of the invention
[0003] The purpose of the present invention is to provide a high-speed target detection and tracking method using infrared event information, which solves the problem that the existing on-board system information processing resources are limited and the effective detection frame rate is often lower than 30Hz@1080p resolution image, which seriously restricts the system's effective detection and tracking capabilities for high-speed moving targets.
[0004] To achieve the above object, the technical solution of the present invention is:
[0005] A high-speed target detection and tracking method using infrared event information, based on infrared event information and image grayscale information simultaneously acquired by a dual-mode infrared camera or by an infrared event detector and an infrared image detector in a common optical path, uses a deep convolutional neural network with spatiotemporal-optical information fusion to process information data; comprising the following steps,
[0006] Step S100, simultaneously acquiring infrared event information and image grayscale information through a dual-mode infrared camera or an infrared event detector and an infrared image detector in a common optical path;
[0007] Step S200, preprocessing the image grayscale information and the infrared event information respectively and converting them into an infrared information cube structure;
[0008] Step S300, processing the infrared information cube data through a deep convolutional neural network of spatiotemporal-optical information fusion to extract the local position and feature information of a specific target;
[0009] Step S400, completing the association of the target detection frame and drawing the trajectory curve to achieve effective tracking of the target.
[0010] The image grayscale information in step S100 refers to the linear conversion of radiation energy into a digital quantity in the range of 0-2N, where N is the number of bits; the infrared event information refers to the pixel output when the radiation energy exceeds the set threshold within a certain time interval, with +1 indicating an increase, -1 indicating a decrease, and 0 indicating no change; wherein the output frame rate of the infrared event information is greater than the output frame rate of the image grayscale information, and the output frame rate of the infrared event information is 5-10 times the output frame rate of the image grayscale information.
[0011] In step S200, the preprocessing of the image grayscale information includes Gaussian filtering, mean filtering, or frequency domain filtering denoising; the preprocessing of the infrared event information includes Gaussian filtering, mean filtering, or frequency domain filtering denoising after multiple frame accumulation;
[0012] The infrared information cube is a three-dimensional matrix, in which the x and y axes represent the length and width of the image respectively, the z axis represents the time dimension, and each slice on the z axis represents a processed image frame and event frame image; the infrared information cube is sent to the deep convolutional neural network of spatiotemporal-optical information fusion in step three for inference calculation.
[0013] The deep convolutional neural network of spatiotemporal-optical information fusion in step S300 refers to a neural network constructed through pre-training to output the target center position, length, width, and confidence. The network simultaneously inputs N consecutive frames of image information, each frame containing image grayscale information and accumulated infrared event information;
[0014] Since the frame rate of infrared event information is greater than the image grayscale information, the input frame event information is cumulative event information. Based on the grayscale of each frame image, the surrounding M×N frame event graphs are selected for accumulation and transformed into a cumulative event graph. Here, M and N are manually set positive integers, N represents the number of input continuous frames, and M represents the input ratio of the number of events and grayscale images.
[0015] Specifically, in step S301, the network uses a convolutional neural network to extract features from the image information of each frame to obtain a spatial feature matrix; and at the same time uses a long short-term memory-convolutional network ConvLSTM to obtain a motion feature matrix;
[0016] Step S302, using the coupling coefficient W to adjust the weights of the spatial features and the motion features to obtain a coupling feature matrix;
[0017] Step S303, extract coupling features through a convolutional neural network, flatten the feature map and learn the relationship between the features and the bounding box coordinates through a fully connected layer, and then directly predict the coordinate values of the bounding box in the output layer.
[0018] The long short-term memory-convolutional network ConvLSTM in step S301 is a 2-layer 2-slice structure, and each layer and each slice contains a number of nodes; the input of each node includes not only information from the same frame, but also information from other adjacent frames; the node output of the last slice and the last layer is selected as the target motion feature, and finally forms a motion feature matrix; the inference formula of the node is as follows:
[0019]
[0020] Among them, * represents matrix multiplication, e represents dot product, σ represents sigmoid activation function, α represents hyperbolic tangent activation function, W represents weight coefficient, which will be continuously updated in iterative training, C represents memory cell, H represents hidden state, and the final motion feature matrix H is obtained by connecting the hidden state H of each time step. (S,L,T) Obtained:
[0021] H = concat(H (s,l,1) ,H (s,l,2) ,H (s,l,3) ,H (s,l,4) ,H (s,l,5) )
[0022] Next, the coupling coefficient W is used to adjust the weights of spatial features and motion features to obtain the coupling feature matrix. Finally, the coupling features are extracted through a convolutional neural network, the feature map is flattened and passed through a fully connected layer to learn the relationship between the features and the bounding box coordinates, and then the coordinate values of the bounding box are directly predicted in the output layer.
[0023] The specific network architecture is based on YOLOv3 and Darknet53, where Darknet-53 is used as a feature extraction network, which contains 53 convolutional layers, each of which is followed by a batch normalization layer and a leaky ReLU activation function, connected through residuals. In addition, the network detection head uses the same decoupled detection head as YOLOv5; the number of channels is uniformly adjusted to 256 through the module-Conv+BN+SiLU activation function with kernel_size=11, stride=1, padding=0, and 256 convolution kernels. In the two parallel branches, kernel_size=33, stride=1, padding=0, and 256 convolution kernels are used. The upper branch is connected to an 11 convolution to obtain information for the prediction of target category information; the lower branch has two 11 convolution layers in parallel, one for predicting target regression parameters and the other for predicting target position and size information.
[0024] In step S400, the center size, length and width of each target detection frame in each frame are stored in the KDTree data structure, and the Euclidean distance formula is used to calculate and quickly find the closest target frame between frames to achieve the association between target frames.
[0025] The specific steps of S400 are as follows:
[0026] Step S401, feature extraction: for each prediction box in each frame, extract its geometric center coordinates (x, y) as the main features for subsequent association;
[0027] Step S402, constructing the size information of the target frame between frames; constructing the size information on the center coordinate set of the prediction frame of the previous frame to quickly find the nearest neighbor frame; assuming that the center point set of the target frame of the previous frame is {p 1 ,p 2 ,…,p m}, each point p i is a two-dimensional vector;
[0028] Step S403, nearest neighbor search: for each target frame in the current frame, search for the nearest neighbor frame of the previous frame through KDTree; let the center point of a target frame in the current frame be qj, and the Euclidean distance calculation formula between it and the nearest neighbor frame of the previous frame is:
[0029]
[0030] Step S404, distance threshold determination; in order to prevent mismatching, a distance threshold δ is set in this project; if the distance between the nearest neighbor frame found and the current frame is less than δ, the two are considered to be matched with the same target; otherwise, the frame in the current frame is considered to be a new target, or the target in the previous frame disappears;
[0031] Step S405, track update: the matched target frames are associated to form a continuous track, unmatched frames are regarded as newly appeared targets, unmatched targets in the previous frame are regarded as disappeared, and the final track is drawn.
[0032] The advantages of the present invention are as follows: 1. By utilizing high-frequency infrared event streams and low-frequency infrared images, information filtering and image denoising are performed on images accumulated from multiple frames to complete preprocessing of event and image information; 2. By combining sparse infrared event information with traditional image grayscale information, the detection time resolution of the space-based infrared system is improved, the feature details in the target moving process are enhanced, the information set is made more continuous, and the probability of discovering high-speed weak and small targets is effectively improved; 3. By adopting cumulative denoising for high-frequency event information, the image denoising effect is improved, and by accumulating time dimension information, the short-term trajectory information of the target is extracted, thereby reducing the difficulty of the target recognition task; 4. By combining the static information contained in the infrared image with the dynamic information in the event stream, and utilizing the advantages of the two types of data, more accurate and faster detection and tracking of high-speed weak and small targets are achieved. After being deployed in an infrared detection system, the detection sensitivity and detection frame rate of high-speed infrared moving small targets can be improved, the energy consumption of the system can be reduced, and the system can be widely used in infrared detection systems with limited processing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is the overall flow chart of the present invention;
[0034] Figure 2 This is a demonstration of the actual effect of the present invention in the deep space background. The left side of the figure is the event accumulation diagram, and the right side is the original infrared image;
[0035] Figure 3 This is a demonstration of the actual effect of the present invention in a complex background. The left side of the figure is an event accumulation diagram, and the right side is an original infrared image;
[0036] Figure 4 The figure is a block diagram of the deep convolutional neural network reasoning for the spatiotemporal-optical signal fusion used in the present invention;
[0037] Figure 5 It is a curve diagram of the actual detection rate of the present invention under complex background;
[0038] Figure 6 This is a demonstration of the actual trajectory tracking effect of the present invention. The right side of the figure is a complex background, and the left side is a deep space background. DETAILED DESCRIPTION
[0039] The present invention is further described below in conjunction with the accompanying drawings, which are only used for exemplary description and cannot be understood as limiting the present invention.
[0040] In order to more concisely describe the present embodiment, some parts known to those skilled in the art but not related to the main content of the present invention are omitted in the drawings or descriptions. In addition, for the convenience of description, some parts in the drawings are omitted, enlarged or reduced, but they do not represent the size or entire structure of the actual product.
[0041] The present invention discloses a high-speed target detection and tracking method using infrared event information, a high-speed target detection and tracking method using infrared event information, based on infrared event information and image grayscale information simultaneously acquired by using a dual-mode infrared camera or by using an infrared event detector with a common optical path and an infrared image detector, and using a deep convolutional neural network of spatiotemporal-optical information fusion to process information data; Figure 1 As shown, the following steps are included:
[0042] Step S100, simultaneously acquiring infrared event information and image grayscale information through a dual-mode infrared camera or an infrared event detector and an infrared image detector in a common optical path;
[0043] The image grayscale information in step S100 refers to the linear conversion of radiation energy into a digital quantity in the range of 0-2N, where N is the number of bits; the infrared event information refers to the pixel output when the radiation energy exceeds the set threshold within a certain time interval, with +1 indicating an increase, -1 indicating a decrease, and 0 indicating no change; wherein the output frame rate of the infrared event information is greater than the output frame rate of the image grayscale information, and the output frame rate of the infrared event information is 5-10 times the output frame rate of the image grayscale information.
[0044] Generally, the output frequency of low-frequency image information is 2 to 20 Hz, and the output frequency of high-frequency event information is 2 to 50 Hz, or even higher. The acquired image output frame sequence is grouped into N, where N is an adjustable parameter, which is set to 5 in this method example; the acquired event output frame sequence is grouped into n×N, where n is an integer from 2 to 50. There are n event frame images between image output frames, which can be set to 5, 10, or 20 in this method example.
[0045] Step S200: The dual-mode infrared camera performs data preprocessing on the image grayscale information and the infrared event information respectively and converts them into an infrared information cube structure.
[0046] In step S200, the preprocessing of the image grayscale information includes Gaussian filtering, mean filtering, or frequency domain filtering denoising; the preprocessing of the infrared event information includes Gaussian filtering, mean filtering, or frequency domain filtering denoising after multiple frame accumulation;
[0047] The infrared information cube is a three-dimensional matrix, in which the x and y axes represent the length and width of the image respectively, the z axis represents the time dimension, and each slice on the z axis represents a processed image frame and event frame image; the infrared information cube is sent to the deep convolutional neural network of spatiotemporal-optical information fusion in step three for inference calculation.
[0048] Perform operations such as Gaussian filtering denoising and size adjustment on the infrared image information; perform m-frame accumulation on the infrared event information, where m is an adjustable parameter, m < n. In this example, m is 5, and n is adjustable between 5, 10, and 20. Then perform Gaussian filtering denoising on the accumulated image, and the effect can be seen in Figure 2 , Figure 3 ; Reconstruct the infrared image and the event frame accumulated image into an infrared information cube, that is, a three-dimensional matrix, where the x and y axes represent the length and width of the picture respectively, and the z axis represents the time dimension. In this example, x is 512, y is 512, and z is 300. Each section on the z axis represents a processed image frame or event frame picture.
[0049] Step S300, through a spatio-temporal - optical signal fusion deep convolutional neural network, as shown in Figure 4 , process the infrared information cube data to extract the local position and feature information of specific targets;
[0050] The spatio-temporal - optical signal fusion deep convolutional neural network in step S300 refers to the output of the target center position, length and width, and confidence through a pre-trained neural network. The network simultaneously inputs continuous N-frame image information, and each frame contains image gray information and cumulative infrared event information;
[0051] Since the frame frequency of the infrared event information is greater than the image gray information, the input frame event information is cumulative event information. Based on the gray value of each frame image, select the surrounding M×N frame event maps for accumulation and transform them into a cumulative event map. Here, M and N are positive integers set manually. N represents the number of consecutive frames input, and M represents the input ratio of the number of events to the number of gray images;
[0052] Specifically, in step S301, the network uses a convolutional neural network to extract the feature of each frame of image information to obtain a spatial feature matrix; at the same time, use a long short-term memory - convolutional network ConvLSTM to obtain a motion feature matrix; ConvLSTM is a deep learning model that combines the spatial feature extraction ability of a convolutional neural network (CNN) and the time series processing ability of a long short-term memory network (LSTM), and is suitable for processing data with spatial and time dimensions, such as video analysis and time series prediction;
[0053] Step S302, use the coupling coefficient W to adjust the weights of the spatial feature and the motion feature to obtain a coupled feature matrix;
[0054] Step S303, extract the coupled feature through a convolutional neural network, flatten the feature map and learn the relationship between the feature and the bounding box coordinates through a fully connected layer, and then directly predict the coordinate values of the bounding box at the output layer.
[0055] The long short-term memory-convolutional network ConvLSTM in step S301 is a 2-layer 2-slice structure, and each layer and each slice contains a number of nodes; the input of each node includes not only information from the same frame, but also information from other adjacent frames; the node output of the last slice and the last layer is selected as the target motion feature, and finally forms a motion feature matrix; the inference formula of the node is as follows:
[0056]
[0057] Among them, * represents matrix multiplication, e represents dot product, σ represents sigmoid activation function, α represents hyperbolic tangent activation function, W represents weight coefficient, which will be continuously updated in iterative training, C represents memory cell, H represents hidden state, and the final motion feature matrix H is obtained by connecting the hidden state H of each time step (S,L,T) Obtained:
[0058] H = concat(H (s,l,1) ,H (s,l,2) ,H (s,l,3) ,H (s,l,4) ,H (s,l,5) )
[0059] Next, the coupling coefficient W is used to adjust the weights of spatial features and motion features to obtain the coupling feature matrix. Finally, the coupling features are extracted through a convolutional neural network, the feature map is flattened and passed through a fully connected layer to learn the relationship between the features and the bounding box coordinates (center point coordinates, width, and height), and then the coordinate values of the bounding box are directly predicted in the output layer.
[0060] The specific network architecture is based on YOLOv3 and Darknet53, where Darknet-53 is used as a feature extraction network, which contains 53 convolutional layers, each of which is followed by a batch normalization layer and a leaky ReLU activation function, connected through residuals. In addition, the network detection head uses the same decoupled detection head as YOLOv5. Through the module with kernel_size=11, stride=1, padding=0, and 256 convolution kernels (Conv+BN+SiLU activation function), the number of channels is uniformly adjusted to 256. In the two parallel branches, kernel_size=33, stride=1, padding=0, and 256 convolution kernels are used. The upper branch is connected to an 11 convolution to obtain information for the prediction of target category information; the lower branch has two 11 convolution layers in parallel, one for predicting target regression parameters and the other for predicting target position and size information.
[0061] In complex background scenes, the target detection rate of this method can be stabilized at more than 90%, see Figure 5 .
[0062] Step S400, completing the association of the target detection frame and drawing the trajectory curve to achieve effective tracking of the target.
[0063] In step S400, the center size, length and width of each target detection frame in each frame are stored in the KDTree data structure, and the Euclidean distance formula is used to calculate and quickly find the closest target frame between frames to achieve the association between target frames.
[0064] The specific steps of S400 are as follows:
[0065] Step S401, feature extraction: for each prediction box in each frame, extract its geometric center coordinates (x, y) as the main features for subsequent association;
[0066] Step S402, constructing the size information of the target frame between frames; constructing the size information on the center coordinate set of the prediction frame of the previous frame to quickly find the nearest neighbor frame; assuming that the center point set of the target frame of the previous frame is {p 1 ,p 2 ,…,p m}, each point p i is a two-dimensional vector;
[0067] Step S403, nearest neighbor search: for each target frame in the current frame, search for the nearest neighbor frame of the previous frame through KDTree; let the center point of a target frame in the current frame be qj, and the Euclidean distance calculation formula between it and the nearest neighbor frame of the previous frame is:
[0068]
[0069] Step S404, distance threshold determination; in order to prevent mismatching, a distance threshold δ is set in this project; if the distance between the nearest neighbor frame found and the current frame is less than δ, the two are considered to be matched with the same target; otherwise, the frame in the current frame is considered to be a new target, or the target in the previous frame disappears;
[0070] Step S405, track update: associate the matched target frames to form a continuous track, treat the unmatched frames as newly appeared targets, and treat the unmatched targets of the previous frame as disappeared, and draw the final track. Figure 6 shown.
[0071] The above description is only a preferred embodiment of the present invention and is not intended to limit the scope of the present invention. That is, any equivalent changes and modifications made according to the content of the patent application scope of the present invention should be within the technical scope of the present invention.
Claims
1. A high-speed target detection and tracking method using infrared event information, characterized in that: Based on the infrared event information and image grayscale information obtained simultaneously by using a dual-mode infrared camera or using an infrared event detector and an infrared image detector in a common optical path, the information data is processed by using a deep convolutional neural network of spatiotemporal-optical information fusion; including the following steps, Step S100, simultaneously acquiring infrared event information and image grayscale information through a dual-mode infrared camera or an infrared event detector and an infrared image detector in a common optical path; Step S200, preprocessing the image grayscale information and the infrared event information respectively and converting them into an infrared information cube structure; Step S300, processing the infrared information cube data through a deep convolutional neural network of spatiotemporal-optical information fusion to extract the local position and feature information of a specific target; Step S400, completing the association of the target detection frame and drawing the trajectory curve to achieve effective tracking of the target.
2. The high-speed target detection and tracking method according to claim 1, characterized in that: The image grayscale information in step S100 refers to the linear conversion of radiation energy into a digital quantity in the range of 0-2N, where N is the number of bits; the infrared event information refers to the pixel output when the radiation energy exceeds the set threshold within a certain time interval, with +1 indicating an increase, -1 indicating a decrease, and 0 indicating no change; wherein the output frame rate of the infrared event information is greater than the output frame rate of the image grayscale information, and the output frame rate of the infrared event information is 5-10 times the output frame rate of the image grayscale information.
3. The high-speed target detection and tracking method according to claim 1, characterized in that: In step S200, the preprocessing of the image grayscale information includes Gaussian filtering, mean filtering, or frequency domain filtering denoising; the preprocessing of the infrared event information includes Gaussian filtering, mean filtering, or frequency domain filtering denoising after multiple frame accumulation; The infrared information cube is a three-dimensional matrix, in which the x and y axes represent the length and width of the image respectively, the z axis represents the time dimension, and each slice on the z axis represents a processed image frame and event frame image; the infrared information cube is sent to the deep convolutional neural network of spatiotemporal-optical information fusion in step three for inference calculation.
4. The high-speed target detection and tracking method according to claim 1, characterized in that: The deep convolutional neural network of spatiotemporal-optical information fusion in step S300 refers to a neural network constructed through pre-training to output the target center position, length, width, and confidence. The network simultaneously inputs N consecutive frames of image information, each frame containing image grayscale information and accumulated infrared event information; Since the frame rate of infrared event information is greater than the image grayscale information, the input frame event information is cumulative event information. Based on the grayscale of each frame image, the surrounding M×N frame event graphs are selected for accumulation and transformed into a cumulative event graph. Here, M and N are manually set positive integers, N represents the number of input continuous frames, and M represents the input ratio of the number of events and grayscale images. Specifically, in step S301, the network uses a convolutional neural network to extract features from the image information of each frame to obtain a spatial feature matrix; and at the same time uses a long short-term memory-convolutional network ConvLSTM to obtain a motion feature matrix; Step S302, using the coupling coefficient W to adjust the weights of the spatial features and the motion features to obtain a coupling feature matrix; Step S303, extract coupling features through a convolutional neural network, flatten the feature map and learn the relationship between the features and the bounding box coordinates through a fully connected layer, and then directly predict the coordinate values of the bounding box in the output layer.
5. The high-speed target detection and tracking method according to claim 4, characterized in that: The long short-term memory-convolutional network ConvLSTM in step S301 is a 2-layer 2-slice structure, and each layer and each slice contains a number of nodes; the input of each node includes not only information from the same frame, but also information from other adjacent frames; the node output of the last slice and the last layer is selected as the target motion feature, and finally a motion feature matrix is formed; The inference formula of the node is as follows: Among them, * represents matrix multiplication, e represents dot product, σ represents sigmoid activation function, α represents hyperbolic tangent activation function, W represents weight coefficient, which will be continuously updated in iterative training, C represents memory cell, H represents hidden state, and the final motion feature matrix H is obtained by connecting the hidden state H of each time step. (S,L,T) Obtained: H=concat(H (s,l,1) ,H (s,l,2) ,H (s,l,3) ,H (s,l,4) ,H (s,l,5) ) Next, the coupling coefficient W is used to adjust the weights of spatial features and motion features to obtain the coupling feature matrix. Finally, the coupling features are extracted through a convolutional neural network, the feature map is flattened and passed through a fully connected layer to learn the relationship between the features and the bounding box coordinates, and then the coordinate values of the bounding box are directly predicted in the output layer. The specific network architecture is based on YOLOv3 and Darknet53, where Darknet-53 is used as a feature extraction network, which contains 53 convolutional layers, each of which is followed by a batch normalization layer and a leaky ReLU activation function, connected through residual connections; in addition, the network detection head uses the same decoupled detection head as YOLOv5; the number of channels is uniformly adjusted to 256 through the module-Conv+BN+SiLU activation function with kernel_size=11, stride=1, padding=0, and 256 convolution kernels. In the two parallel branches, the convolution module with kernel_size=33, stride=1, padding=0, and 256 convolution kernels is used; the upper branch is connected to an 11 convolution to obtain information for the prediction of target category information; the lower branch has two 11 convolutional layers in parallel, one for predicting target regression parameters and the other for predicting target position and size information.
6. The high-speed target detection and tracking method according to claim 1, characterized in that: In step S400, the center size, length and width of each target detection frame in each frame are stored in the KDTree data structure, and the Euclidean distance formula is used to calculate and quickly find the closest target frame between frames to achieve the association between target frames.
7. The high-speed target detection and tracking method according to claim 6, characterized in that: The specific steps of step S400 are as follows: Step S401, feature extraction: for each prediction box in each frame, extract its geometric center coordinates (x, y) as the main features for subsequent association; Step S402, constructing the size information of the target frame between frames; constructing the size information on the center coordinate set of the prediction frame of the previous frame to quickly find the nearest neighbor frame; Assume that the target box center point set of the previous frame is {p1,p2,…,p m }, each point p i is a two-dimensional vector; Step S403, nearest neighbor search: for each target frame in the current frame, search for the nearest neighbor frame of the previous frame through KDTree; let the center point of a target frame in the current frame be qj, and the Euclidean distance calculation formula between it and the nearest neighbor frame of the previous frame is: Step S404, distance threshold determination; in order to prevent mismatching, a distance threshold δ is set in this project; if the distance between the nearest neighbor frame found and the current frame is less than δ, the two are considered to be matched with the same target; otherwise, the frame in the current frame is considered to be a new target, or the target in the previous frame disappears; Step S405, track update: the matched target frames are associated to form a continuous track, unmatched frames are regarded as newly appeared targets, unmatched targets in the previous frame are regarded as disappeared, and the final track is drawn.
Citation Information
Cited By
Intelligent event identification method and system based on high-speed camera
CN121121021A
Infrared high-speed target detection tracking method and system based on alternation of full-frame patrol and multi-window burst reading
CN122420626A