Dynamic video processing method for target tracking
Through dynamic video processing methods, including preprocessing, pulsed neural network feature extraction, YOLO object detection and Kalman filtering tracking, the tracking accuracy and real-time problems of traditional methods in target occlusion, lighting changes and background complex situations are solved, efficient and robust target tracking is achieved, and computational costs are reduced.
Patent Information
- Application Number
- CN202411816467.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional target tracking methods cannot effectively deal with the problems of target occlusion, lighting changes and complex backgrounds, resulting in poor tracking accuracy and real-time performance, inability to conduct real-time monitoring and intelligent analysis, and high network computing costs, making it impossible to achieve efficient target tracking in resource-constrained environments.
Dynamic video processing methods are adopted, including obtaining original video data and preprocessing, extracting feature information based on pulsed neural network, building an object detection model based on YOLO object detection algorithm, dynamic tracking is performed through Kalman filtering algorithm, and adjusting the tracking results using the result optimization algorithm.
It improves the accuracy and real-time nature of target tracking, can effectively deal with target occlusion, lighting changes and complex background problems, is highly robust, and realizes real-time monitoring and intelligent analysis of dynamic video through FPGA accelerator, reducing network computing costs.
Smart Images

Figure CN119942397A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video scene processing, and in particular to a dynamic video processing method for target tracking. Background Art
[0002] With the rapid development of video surveillance technology, dynamic video processing is increasingly used in security, traffic monitoring, intelligent manufacturing and other fields. By tracking the targets of interest in dynamic videos, real-time monitoring and positioning of video targets can be achieved.
[0003] However, traditional target tracking methods cannot effectively deal with problems such as target occlusion, lighting changes, and complex backgrounds, which may lead to reduced tracking accuracy or even target tracking failure. In addition, the tracking real-time performance of existing methods is poor, and dynamic video cannot be monitored and intelligently analyzed in real time. The network computing cost is high, and efficient target tracking cannot be achieved in a resource-constrained environment.
[0004] Therefore, it is necessary to provide a dynamic video processing method for target tracking to solve the above technical problems. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a dynamic video processing method for target tracking, which is used to solve the problems that traditional target tracking methods cannot effectively deal with target occlusion, lighting changes and complex backgrounds, the tracking accuracy and real-time performance are poor, and real-time monitoring and intelligent analysis of dynamic videos cannot be performed. In addition, the network computing cost is high, and efficient target tracking cannot be achieved in a resource-constrained environment.
[0006] The present invention provides a dynamic video processing method for target tracking, the processing method comprising:
[0007] Acquire original video data, and pre-process the original video data to generate corresponding dynamic video data;
[0008] Extracting feature information from the dynamic video data based on a spiking neural network to generate corresponding dynamic feature data;
[0009] Building a target detection model based on the YOLO target detection algorithm, dynamically detecting the dynamic feature data through the target detection model, and determining the tracking target in the dynamic video data;
[0010] Dynamically track the tracking target based on the Kalman filter algorithm, predict and update the state of the tracking target, and generate a corresponding target tracking result;
[0011] The target tracking result is evaluated based on a result optimization algorithm, and the target tracking result is adjusted and optimized according to the evaluation result.
[0012] Preferably, the obtaining of original video data and preprocessing the original video data to generate corresponding dynamic video data specifically includes:
[0013] Acquire the original video data, and adjust the resolution of the original video data to generate corresponding first video data;
[0014] Performing color correction on the first video data, that is, adjusting the brightness, contrast and saturation of the first video data, to generate corresponding second video data;
[0015] A dynamic visual sensor is used to capture brightness changes in the second video data to generate corresponding dynamic video data.
[0016] Preferably, the pulse neural network is mapped and deployed on an FPGA accelerator.
[0017] Preferably, the extracting feature information from the dynamic video data based on the pulse neural network to generate corresponding dynamic feature data specifically includes:
[0018] The corresponding pulse emission state is determined according to the membrane potential of the neuron in the spiking neural network, and the corresponding calculation formula is as follows:
[0019]
[0020] In the formula, S i (t) represents the pulse emission state of neuron i at time t, that is, S i (t) = 1 means that neuron i emits a pulse at time t, and S i (t) = 0 means that neuron i does not emit a pulse at time t; V i (t) represents the membrane potential of neuron i at time t; θ i Represents the preset membrane potential threshold of neuron i.
[0021] Preferably, the dynamic feature data includes pulse emission states of a plurality of neurons in the spiking neural network, and each neuron corresponds to a video region in the dynamic video data;
[0022] In the time window M, the pulse emission states of all neurons belonging to the video region j are summed to obtain the jth feature element of the dynamic feature data. The corresponding calculation formula is as follows:
[0023]
[0024] In the formula, F j represents the jth feature element of the dynamic feature data; M represents the time window; Si (t) represents the pulse emission state of neuron i at time t, that is, S i (t) = 1 means neuron i fires a pulse, S i (t) = 0 means that neuron i does not emit a pulse; ij represents the indicator function. When neuron i belongs to video region j, δ ij = 1, when neuron i does not belong to video region j, δ ij =0;
[0025] All calculated characteristic elements are summarized to generate the dynamic characteristic data F.
[0026] Preferably, the target detection model is constructed based on the YOLO target detection algorithm, and the dynamic feature data is dynamically detected by the target detection model to determine the tracking target in the dynamic video data, which specifically includes:
[0027] The target detection model is constructed based on the YOLO target detection algorithm, and the extracted dynamic feature data is input into the target detection model, and the tracking target in the dynamic video data is determined after dynamic detection, and the coordinates, width and height of the center point of the bounding box corresponding to the tracking target are obtained;
[0028] The calculation formula for the coordinates of the center point of the bounding box corresponding to the tracking target is as follows:
[0029]
[0030] Where x represents the horizontal coordinate of the center point of the bounding box corresponding to the tracking target; represents the transposition of the weight vector corresponding to the horizontal coordinate of the center point of the bounding box of the tracking target determined by the target detection model; F represents the dynamic feature data; b x represents the bias item corresponding to the horizontal coordinate of the center point of the bounding box of the tracked target determined by the target detection model; y represents the vertical coordinate of the center point of the bounding box corresponding to the tracked target; represents the transpose of the weight vector corresponding to the ordinate of the center point of the bounding box of the tracked target determined by the target detection model; b y Represents the bias term corresponding to the ordinate of the center point of the bounding box of the tracked target determined by the target detection model.
[0031] Preferably, the calculation formula for the width and height of the bounding box corresponding to the tracking target is as follows:
[0032]
[0033] Where w represents the width of the bounding box corresponding to the tracked target; represents the transpose of the weight vector corresponding to the width of the bounding box of the tracking target determined by the target detection model; F represents the dynamic feature data; b w represents the bias term corresponding to the width of the bounding box of the tracked target determined by the target detection model; h represents the height of the bounding box corresponding to the tracked target; represents the transpose of the weight vector corresponding to the bounding box height of the tracked target determined by the target detection model; b h It represents the bias term corresponding to the bounding box height of the tracked target determined by the target detection model; exp() represents an exponential function with the natural constant e as the base.
[0034] Preferably, the dynamically tracking the tracking target based on the Kalman filter algorithm, predicting and updating the state of the tracking target, and generating a corresponding target tracking result specifically include:
[0035] During the dynamic tracking process of the tracking target, the predicted estimated value of the state of the tracking target is:
[0036] ZT(k|k-1)=G(k)ZT(k-1|k-1)+u(k)
[0037] Wherein, ZT(k|k-1) represents the predicted estimated value of the state of the tracking target at time k; G(k) represents the state transfer matrix of the tracking target at time k; ZT(k-1|k-1) represents the predicted estimated value of the state of the tracking target at time k-1; u(k) represents the control amount of the tracking target at time k, that is, the driving input vector;
[0038] Update the state covariance of the tracking target to obtain a predicted estimated value of the state covariance of the tracking target:
[0039] ZTXFC(k|k-1)=G(k)ZTXFC(k-1|k-1)G(k) T +Q(k)
[0040] Where ZTXFC(k|k-1) represents the predicted estimated value of the state covariance of the tracking target at time k; G(k) represents the state transfer matrix of the tracking target at time k; G(k) T represents the transpose of the state transfer matrix of the tracking target at time k; ZTXFC(k-1|k-1) represents the predicted estimate of the state covariance of the tracking target at time k-1; Q(k) represents the covariance matrix of the process noise of the tracking target at time k.
[0041] Preferably, the Kalman gain update formula of the tracking target is:
[0042] ZY(k)=ZTXFC(k|k-1)H(k) T(H(k)ZTXFC(k|k-1)H(k) T +R(k)) -1
[0043] Where ZY(k) represents the Kalman gain of the tracking target at time k; ZTXFC(k|k-1) represents the predicted estimated value of the state covariance of the tracking target at time k; H(k) represents the measurement matrix of the tracking target at time k; H(k) T represents the transpose of the measurement matrix of the tracking target at time k; R(k) represents the covariance matrix of the measurement noise of the tracking target at time k; () -1 Represents the inverse of a matrix.
[0044] Preferably, the state of the tracking target is updated to obtain an estimated value of the state of the tracking target:
[0045] ZT(k|k)=ZT(k|k-1)+ZY(k)(z(k)-H(k)ZT(k|k-1))
[0046] Wherein, ZT(k|k) represents the estimated value of the state of the tracking target at time k; ZY(k) represents the Kalman gain of the tracking target at time k; z(k) represents the measured value of the tracking target at time k; H(k) represents the measurement matrix of the tracking target at time k; ZT(k|k-1) represents the predicted estimated value of the state of the tracking target at time k;
[0047] Update the state covariance of the tracking target, and obtain the estimated value of the state covariance of the tracking target as:
[0048] ZTXFC(k|k)=I-ZY(k)H(k)ZTXFC(k|k-1)
[0049] Wherein, ZTXFC(k|k) represents the estimated value of the state covariance of the tracking target at time k; I represents the unit matrix; ZY(k) represents the Kalman gain of the tracking target at time k; H(k) represents the measurement matrix of the tracking target at time k; ZTXFC(k|k-1) represents the predicted estimated value of the state covariance of the tracking target at time k;
[0050] Predict and update the state of the tracking target during the dynamic tracking process, and generate the corresponding target tracking result.
[0051] Compared with the related art, the dynamic video processing method for target tracking provided by the present invention has the following beneficial effects:
[0052] The present invention obtains original video data and pre-processes the original video data to generate corresponding dynamic video data; extracts feature information in dynamic video data based on a pulse neural network to generate corresponding dynamic feature data; constructs a target detection model based on a YOLO target detection algorithm, dynamically detects dynamic feature data through the target detection model, and determines the tracking target in the dynamic video data; dynamically tracks the tracking target based on a Kalman filter algorithm, predicts and updates the state of the tracking target, and generates a corresponding target tracking result; evaluates the target tracking result based on a result optimization algorithm, and adjusts and optimizes the target tracking result according to the evaluation result. The present invention can accurately detect dynamic targets by introducing a pulse neural network and a YOLO target detection algorithm, and uses a Kalman filter algorithm to perform real-time target tracking, thereby improving the accuracy and real-time performance of target tracking; the method of the present invention can effectively deal with the problems of target occlusion, illumination change, and complex background, and has strong robustness; by introducing an FPGA accelerator architecture, hardware acceleration of the pulse neural network algorithm can be achieved, thereby achieving real-time monitoring and intelligent analysis of dynamic video; at the same time, the network computing cost is reduced, and efficient target tracking can also be achieved in a resource-constrained environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 A flow chart of a dynamic video processing method for target tracking according to the present invention;
[0054] Figure 2 The present invention is a flow chart of generating dynamic video data through preprocessing of original video data. DETAILED DESCRIPTION
[0055] The present invention will be further described below in conjunction with the accompanying drawings and implementation modes.
[0056] Embodiment 1
[0057] like Figure 1 As shown, a dynamic video processing method for target tracking, the processing method comprises:
[0058] S1, obtaining original video data, and preprocessing the original video data to generate corresponding dynamic video data;
[0059] S2, extracting feature information from the dynamic video data based on a pulse neural network to generate corresponding dynamic feature data;
[0060] S3, building a target detection model based on the YOLO target detection algorithm, dynamically detecting the dynamic feature data through the target detection model, and determining the tracking target in the dynamic video data;
[0061] S4, dynamically tracking the tracking target based on the Kalman filter algorithm, predicting and updating the state of the tracking target, and generating a corresponding target tracking result;
[0062] S5, evaluating the target tracking result based on a result optimization algorithm, and adjusting and optimizing the target tracking result according to the evaluation result.
[0063] First, the raw video data can be captured, and then preprocessed to generate dynamic video data suitable for subsequent analysis. The preprocessing stage includes resolution adjustment, color correction, and brightness change capture, which can ensure the accuracy and consistency of the dynamic video data.
[0064] Using the spiking neural network (SNN) technology, we can extract feature information with significant discrimination from the preprocessed dynamic video data. SNN can simulate the biological neuron pulse emission mechanism and generate high-dimensional dynamic feature data.
[0065] Subsequently, a target detection model based on the YOLO target detection algorithm can be constructed to perform real-time dynamic detection of dynamic feature data and accurately identify the tracking target in dynamic video data. The target detection model can predict the bounding box corresponding to the tracking target through a single forward pass, thereby significantly improving the accuracy and real-time performance of dynamic monitoring.
[0066] In order to achieve continuous and smooth tracking of the target, the Kalman filter algorithm can be used to effectively fuse the current observation information with the previous state estimation. In this way, when facing complex scenarios such as fast movement or occlusion of the target, the state information such as the position and speed of the target can be accurately predicted and updated to generate accurate target tracking results.
[0067] Finally, the result optimization algorithm can be introduced to evaluate the target tracking results, including the evaluation of multiple dimensions such as tracking accuracy, continuity and stability. Based on the evaluation results, the target tracking results can be further optimized through parameter adjustment and other means to ensure the accuracy and reliability of the output tracking information.
[0068] In the specific implementation process, Figure 2 As shown, the obtaining of original video data and preprocessing of the original video data to generate corresponding dynamic video data specifically includes:
[0069] Acquire the original video data, and adjust the resolution of the original video data to generate corresponding first video data;
[0070] Performing color correction on the first video data, that is, adjusting the brightness, contrast and saturation of the first video data, to generate corresponding second video data;
[0071] A dynamic visual sensor is used to capture brightness changes in the second video data to generate corresponding dynamic video data.
[0072] It is understandable that after acquiring the original video data, the resolution can be adjusted first to ensure the clarity of the video data and the compatibility of the subsequent processing flow. Through the image processing algorithm, the resolution of the original video data can be adjusted to a preset high resolution standard to generate the corresponding first video data.
[0073] Next, the first video data can be color corrected. Color correction involves precise adjustment of the brightness, contrast, and saturation of the video data, thereby optimizing the color parameters of the first video data to eliminate color deviation, enhance color expression, and generate second video data with balanced colors and better visual effects.
[0074] After the color correction is completed, the second video data can be processed accordingly using a dynamic vision sensor. The dynamic vision sensor can capture and extract the brightness change information in the second video data in real time and convert it into a digital signal to generate the corresponding dynamic video data. The dynamic video data not only retains the basic information of the original video data, but also incorporates the dynamic characteristics of brightness changes, which facilitates the subsequent advanced video processing tasks of target tracking and behavior analysis.
[0075] The spiking neural network is mapped and deployed on an FPGA accelerator.
[0076] In practical applications, the spiking neural network algorithm can be mapped and deployed on FPGA, and a series of key strategies can be used to accelerate the operation of the spiking neural network on the FPGA accelerator, thereby optimizing data flow and effectively improving the storage and computing efficiency of the spiking neural network.
[0077] The extracting feature information from the dynamic video data based on the pulse neural network to generate corresponding dynamic feature data specifically includes:
[0078] The corresponding pulse emission state is determined according to the membrane potential of the neuron in the spiking neural network, and the corresponding calculation formula is as follows:
[0079]
[0080] In the formula, S i (t) represents the pulse emission state of neuron i at time t, that is, S i (t) = 1 means that neuron i emits a pulse at time t, and S i (t) = 0 means that neuron i does not emit a pulse at time t; V i (t) represents the membrane potential of neuron i at time t; θi Represents the preset membrane potential threshold of neuron i.
[0081] The dynamic feature data includes the pulse emission states of a plurality of neurons in the spiking neural network, and each neuron corresponds to a video region in the dynamic video data;
[0082] In the time window M, the pulse emission states of all neurons belonging to the video region j are summed to obtain the jth feature element of the dynamic feature data. The corresponding calculation formula is as follows:
[0083]
[0084] In the formula, F j represents the jth feature element of the dynamic feature data; M represents the time window; S i (t) represents the pulse emission state of neuron i at time t, that is, S i (t) = 1 means neuron i fires a pulse, S i (t) = 0 means that neuron i does not emit a pulse; ij represents the indicator function. When neuron i belongs to video region j, δ ij = 1, when neuron i does not belong to video region j, δ ij =0;
[0085] All calculated characteristic elements are summarized to generate the dynamic characteristic data F.
[0086] The target detection model is constructed based on the YOLO target detection algorithm, and the dynamic feature data is dynamically detected by the target detection model to determine the tracking target in the dynamic video data, specifically including:
[0087] The target detection model is constructed based on the YOLO target detection algorithm, and the extracted dynamic feature data is input into the target detection model, and the tracking target in the dynamic video data is determined after dynamic detection, and the coordinates, width and height of the center point of the bounding box corresponding to the tracking target are obtained;
[0088] The calculation formula for the coordinates of the center point of the bounding box corresponding to the tracking target is as follows:
[0089]
[0090] Where x represents the horizontal coordinate of the center point of the bounding box corresponding to the tracking target; represents the transposition of the weight vector corresponding to the horizontal coordinate of the center point of the bounding box of the tracking target determined by the target detection model; F represents the dynamic feature data; b xrepresents the bias item corresponding to the horizontal coordinate of the center point of the bounding box of the tracked target determined by the target detection model; y represents the vertical coordinate of the center point of the bounding box corresponding to the tracked target; represents the transpose of the weight vector corresponding to the ordinate of the center point of the bounding box of the tracked target determined by the target detection model; b y Represents the bias term corresponding to the ordinate of the center point of the bounding box of the tracked target determined by the target detection model.
[0091] The calculation formula for the width and height of the bounding box corresponding to the tracking target is as follows:
[0092]
[0093] Where w represents the width of the bounding box corresponding to the tracked target; represents the transpose of the weight vector corresponding to the width of the bounding box of the tracking target determined by the target detection model; F represents the dynamic feature data; b w represents the bias term corresponding to the width of the bounding box of the tracked target determined by the target detection model; h represents the height of the bounding box corresponding to the tracked target; represents the transpose of the weight vector corresponding to the bounding box height of the tracked target determined by the target detection model; b h It represents the bias term corresponding to the bounding box height of the tracked target determined by the target detection model; exp() represents an exponential function with the natural constant e as the base.
[0094] It should be noted that the pulse emission state can be determined based on the change of neuron membrane potential based on the pulse neural network. Specifically, the pulse emission state of each neuron is determined by its membrane potential and the preset membrane potential threshold. When the membrane potential reaches or exceeds the preset membrane potential threshold, the neuron emits a pulse; conversely, when the membrane potential is less than the preset membrane potential threshold, the neuron does not emit a pulse.
[0095] The dynamic feature data includes the pulse emission states of multiple neurons in the spiking neural network, and each neuron corresponds to a video region in the dynamic video data. In a specific time window, the pulse emission states of all neurons belonging to a specific video region can be summed to generate feature elements of the dynamic feature data, and all feature elements can be summarized to obtain the dynamic feature data.
[0096] Finally, the YOLO target detection algorithm can be used to build a target detection model to achieve dynamic detection of dynamic feature data. Specifically, the extracted dynamic feature data can be input into the target detection model, and the target detection model can be used to analyze and determine the tracking target in the dynamic video data, and calculate the coordinates, width and height of the center point of the bounding box corresponding to the tracking target.
[0097] The dynamically tracking the tracking target based on the Kalman filter algorithm, predicting and updating the state of the tracking target, and generating the corresponding target tracking result specifically include:
[0098] During the dynamic tracking process of the tracking target, the predicted estimated value of the state of the tracking target is:
[0099] ZT(k|k-1)=G(k)ZT(k-1|k-1)+u(k)
[0100] Wherein, ZT(k|k-1) represents the predicted estimated value of the state of the tracking target at time k; G(k) represents the state transfer matrix of the tracking target at time k; ZT(k-1|k-1) represents the predicted estimated value of the state of the tracking target at time k-1; u(k) represents the control amount of the tracking target at time k, that is, the driving input vector;
[0101] Update the state covariance of the tracking target to obtain a predicted estimated value of the state covariance of the tracking target:
[0102] ZTXFC(k|k-1)=G(k)ZTXFC(k-1|k-1)G(k) T +Q(k)
[0103] Where ZTXFC(k|k-1) represents the predicted estimated value of the state covariance of the tracking target at time k; G(k) represents the state transfer matrix of the tracking target at time k; G(k) T represents the transpose of the state transfer matrix of the tracking target at time k; ZTXFC(k-1|k-1) represents the predicted estimate of the state covariance of the tracking target at time k-1; Q(k) represents the covariance matrix of the process noise of the tracking target at time k.
[0104] The Kalman gain update formula of the tracking target is:
[0105] ZY(k)=ZTXFC(k|k-1)H(k) T (H(k)ZTXFC(k|k-1)H(k) T +R(k)) -1
[0106] Where ZY(k) represents the Kalman gain of the tracking target at time k; ZTXFC(k|k-1) represents the predicted estimated value of the state covariance of the tracking target at time k; H(k) represents the measurement matrix of the tracking target at time k; H(k) T represents the transpose of the measurement matrix of the tracking target at time k; R(k) represents the covariance matrix of the measurement noise of the tracking target at time k; () -1 Represents the inverse of a matrix.
[0107] Update the state of the tracking target, and obtain an estimated value of the state of the tracking target as follows:
[0108] ZT(k|k)=ZT(k|k-1)+ZY(k)(z(k)-H(k)ZT(k|k-1))
[0109] Wherein, ZT(k|k) represents the estimated value of the state of the tracking target at time k; ZY(k) represents the Kalman gain of the tracking target at time k; z(k) represents the measured value of the tracking target at time k; H(k) represents the measurement matrix of the tracking target at time k; ZT(k|k-1) represents the predicted estimated value of the state of the tracking target at time k;
[0110] Update the state covariance of the tracking target, and obtain the estimated value of the state covariance of the tracking target as:
[0111] ZTXFC(k|k)=I-ZY(k)H(k)ZTXFC(k|k-1)
[0112] Wherein, ZTXFC(k|k) represents the estimated value of the state covariance of the tracking target at time k; I represents the unit matrix; ZY(k) represents the Kalman gain of the tracking target at time k; H(k) represents the measurement matrix of the tracking target at time k; ZTXFC(k|k-1) represents the predicted estimated value of the state covariance of the tracking target at time k;
[0113] Predict and update the state of the tracking target during the dynamic tracking process, and generate the corresponding target tracking result.
[0114] Among them, the state prediction estimate at the previous moment and the state transfer matrix at the current moment can be used, combined with the control quantity at the current moment, to calculate the state prediction estimate at the current moment, so that the state information at the previous moment can be mapped to the current moment through the state transfer matrix, while considering the influence of the external control quantity to achieve dynamic prediction of the state.
[0115] Subsequently, the state covariance of the target can be updated to reflect the uncertainty of the state prediction. The state covariance prediction estimate at the current moment is calculated using the state transfer matrix and the state covariance prediction estimate at the previous moment, combined with the covariance matrix of the process noise.
[0116] Next, the Kalman gain can be calculated to determine the contribution of the measurement value in the state update. The Kalman gain can be calculated from the state covariance prediction estimate, the measurement matrix, and the covariance matrix of the measurement noise.
[0117] By using the Kalman gain and the current measurement value, the state of the tracking target can be updated to obtain the current state estimate, so that the predicted state can be combined with the measurement value to achieve accurate state estimation.
[0118] Finally, the state covariance of the tracked target can be updated to obtain the estimated state covariance at the current moment. The updated state covariance can be calculated through the unit matrix, Kalman gain, measurement matrix and state covariance prediction estimate, so that the state of the target can be tracked in real time and accurately, and the corresponding target tracking result can be generated.
[0119] Through the introduction of the above embodiments, the present invention uses a dynamic video processing method for target tracking, obtains original video data, and pre-processes the original video data to generate corresponding dynamic video data; extracts feature information in dynamic video data based on a pulse neural network to generate corresponding dynamic feature data; constructs a target detection model based on a YOLO target detection algorithm, dynamically detects dynamic feature data through the target detection model, and determines the tracking target in the dynamic video data; dynamically tracks the tracking target based on a Kalman filter algorithm, predicts and updates the state of the tracking target, and generates a corresponding target tracking result; evaluates the target tracking result based on a result optimization algorithm, and adjusts and optimizes the target tracking result according to the evaluation result. The present invention can accurately detect dynamic targets by introducing a pulse neural network and a YOLO target detection algorithm, and uses a Kalman filter algorithm to perform real-time target tracking, thereby improving the accuracy and real-time performance of target tracking; the method of the present invention can effectively deal with the problems of target occlusion, illumination changes, and complex backgrounds, and has strong robustness; by introducing an FPGA accelerator architecture, hardware acceleration of the pulse neural network algorithm can be achieved, thereby achieving real-time monitoring and intelligent analysis of dynamic videos; at the same time, the network computing cost is reduced, and efficient target tracking can also be achieved in a resource-constrained environment.
[0120] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0121] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0122] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
Claims
1. A dynamic video processing method for target tracking, characterized in that: The processing method comprises: Acquire original video data, and pre-process the original video data to generate corresponding dynamic video data; Extracting feature information from the dynamic video data based on a pulse neural network to generate corresponding dynamic feature data; Building a target detection model based on the YOLO target detection algorithm, dynamically detecting the dynamic feature data through the target detection model, and determining the tracking target in the dynamic video data; Dynamically track the tracking target based on the Kalman filter algorithm, predict and update the state of the tracking target, and generate a corresponding target tracking result; The target tracking result is evaluated based on a result optimization algorithm, and the target tracking result is adjusted and optimized according to the evaluation result.
2. A dynamic video processing method for target tracking according to claim 1, characterized in that: The obtaining of original video data and preprocessing the original video data to generate corresponding dynamic video data specifically includes: Acquire the original video data, and adjust the resolution of the original video data to generate corresponding first video data; Performing color correction on the first video data, that is, adjusting the brightness, contrast and saturation of the first video data, to generate corresponding second video data; A dynamic visual sensor is used to capture brightness changes in the second video data to generate corresponding dynamic video data.
3. The dynamic video processing method for target tracking according to claim 1, characterized in that: The spiking neural network is mapped and deployed on an FPGA accelerator.
4. The dynamic video processing method for target tracking according to claim 1, characterized in that: The extracting feature information from the dynamic video data based on the pulse neural network to generate corresponding dynamic feature data specifically includes: The corresponding pulse emission state is determined according to the membrane potential of the neuron in the spiking neural network, and the corresponding calculation formula is as follows: In the formula, S i (t) represents the pulse emission state of neuron i at time t, that is, S i (t) = 1 means that neuron i emits a pulse at time t, and S i (t) = 0 means that neuron i does not emit a pulse at time t; V i (t) represents the membrane potential of neuron i at time t; θ i Represents the preset membrane potential threshold of neuron i.
5. A dynamic video processing method for target tracking according to claim 4, characterized in that: The dynamic feature data includes the pulse emission states of a plurality of neurons in the spiking neural network, and each neuron corresponds to a video region in the dynamic video data; In the time window M, the pulse emission states of all neurons belonging to the video region j are summed to obtain the jth feature element of the dynamic feature data. The corresponding calculation formula is as follows: In the formula, F j Represents the jth feature element of dynamic feature data; M represents the time window; S i (t) represents the pulse emission state of neuron i at time t, that is, S i (t) = 1 means neuron i fires a pulse, S i (t) = 0 means that neuron i does not fire a pulse; δ ij represents the indicator function. When neuron i belongs to video region j, δ ij = 1, when neuron i does not belong to video region j, δ ij =0; All calculated characteristic elements are summarized to generate the dynamic characteristic data F.
6. The dynamic video processing method for target tracking according to claim 1, characterized in that: The target detection model is constructed based on the YOLO target detection algorithm, and the dynamic feature data is dynamically detected by the target detection model to determine the tracking target in the dynamic video data, specifically including: The target detection model is constructed based on the YOLO target detection algorithm, and the extracted dynamic feature data is input into the target detection model, and the tracking target in the dynamic video data is determined after dynamic detection, and the coordinates, width and height of the center point of the bounding box corresponding to the tracking target are obtained; The calculation formula for the coordinates of the center point of the bounding box corresponding to the tracking target is as follows: Where x represents the horizontal coordinate of the center point of the bounding box corresponding to the tracking target; represents the transposition of the weight vector corresponding to the horizontal coordinate of the center point of the bounding box of the tracking target determined by the target detection model; F represents the dynamic feature data; b x represents the bias item corresponding to the horizontal coordinate of the center point of the bounding box of the tracked target determined by the target detection model; y represents the vertical coordinate of the center point of the bounding box corresponding to the tracked target; represents the transpose of the weight vector corresponding to the ordinate of the center point of the bounding box of the tracked target determined by the target detection model; b y Represents the bias term corresponding to the ordinate of the center point of the bounding box of the tracked target determined by the target detection model.
7. A dynamic video processing method for target tracking according to claim 6, characterized in that: The calculation formula for the width and height of the bounding box corresponding to the tracking target is as follows: Where w represents the width of the bounding box corresponding to the tracked target; represents the transpose of the weight vector corresponding to the width of the bounding box of the tracked target determined by the target detection model; F represents the dynamic feature data; b w represents the bias term corresponding to the width of the bounding box of the tracked target determined by the target detection model; h represents the height of the bounding box corresponding to the tracked target; represents the transpose of the weight vector corresponding to the bounding box height of the tracked target determined by the target detection model; b h Represents the bias term corresponding to the bounding box height of the tracked target determined by the target detection model; exp() represents an exponential function with the natural constant e as the base.
8. The dynamic video processing method for target tracking according to claim 1, characterized in that: The dynamically tracking the tracking target based on the Kalman filter algorithm, predicting and updating the state of the tracking target, and generating the corresponding target tracking result specifically include: During the dynamic tracking process of the tracking target, the predicted estimated value of the state of the tracking target is: ZT(k|k-1)=G(k)ZT(k-1|k-1)+u(k) Wherein, ZT(k|k-1) represents the predicted estimated value of the state of the tracking target at time k; G(k) represents the state transfer matrix of the tracking target at time k; ZT(k-1|k-1) represents the predicted estimated value of the state of the tracking target at time k-1; u(k) represents the control amount of the tracking target at time k, that is, the driving input vector; Update the state covariance of the tracking target to obtain a predicted estimated value of the state covariance of the tracking target: ZTXFC(k|k-1)=G(k)ZTXFC(k-1|k-1)G(k) T +Q(k) Where ZTXFC(k|k-1) represents the predicted estimated value of the state covariance of the tracking target at time k; G(k) represents the state transfer matrix of the tracking target at time k; G(k) T represents the transpose of the state transfer matrix of the tracking target at time k; ZTXFC(k-1|k-1) represents the predicted estimate of the state covariance of the tracking target at time k-1; Q(k) represents the covariance matrix of the process noise of the tracking target at time k.
9. A dynamic video processing method for target tracking according to claim 8, characterized in that: The Kalman gain update formula of the tracking target is: ZY(k)=ZTXFC(k|k-1)H(k) T (H(k)ZTXFC(k|k-1)H(k) T +R(k)) -1 Where ZY(k) represents the Kalman gain of the tracking target at time k; ZTXFC(k|k-1) represents the predicted estimated value of the state covariance of the tracking target at time k; H(k) represents the measurement matrix of the tracking target at time k; H(k) T represents the transpose of the measurement matrix of the tracking target at time k; R(k) represents the covariance matrix of the measurement noise of the tracking target at time k; () -1 Represents the inverse of a matrix.
10. A dynamic video processing method for target tracking according to claim 9, characterized in that: Update the state of the tracking target, and obtain an estimated value of the state of the tracking target as follows: ZT(k|k)=ZT(k|k-1)+ZY(k)(z(k)-H(k)ZT(k|k-1)) Wherein, ZT(k|k) represents the estimated value of the state of the tracking target at time k; ZY(k) represents the Kalman gain of the tracking target at time k; z(k) represents the measured value of the tracking target at time k; H(k) represents the measurement matrix of the tracking target at time k; ZT(k|k-1) represents the predicted estimated value of the state of the tracking target at time k; Update the state covariance of the tracking target, and obtain the estimated value of the state covariance of the tracking target as: ZTXFC(k|k)=I-ZY(k)H(k)ZTXFC(k|k-1) Wherein, ZTXFC(k|k) represents the estimated value of the state covariance of the tracking target at time k; I represents the unit matrix; ZY(k) represents the Kalman gain of the tracking target at time k; H(k) represents the measurement matrix of the tracking target at time k; ZTXFC(k|k-1) represents the predicted estimated value of the state covariance of the tracking target at time k; Predict and update the state of the tracking target during the dynamic tracking process, and generate the corresponding target tracking result.
Citation Information
Patent Citations
Target detection processing method and device based on pulse signal, equipment and medium
CN114997235A
Target tracking method based on event data and spiking neural network
CN116822592A
High-frame-rate target tracking method based on deep pulse neural network
CN117876927A
Target tracking method based on Deepsort algorithm
CN118115538A
Multi-target object motion tracking method
CN118229728A