A pop-up target real-time identification tracking method and device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]本发明的主要目的是提供一种闪现目标实时识别跟踪方法及装置,旨在克服现有视觉预测模型与控制算法的滞后缺陷实现对高速机动且短暂闪现目标的毫秒级精准探测预测与稳定锁定伺服跟踪
[0026]采用上述技术方案具有以下优点:
Smart Images

Figure CN122205240B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision, sensor fusion and visual servo control technology, and in particular to a method and apparatus for real-time identification and tracking of flashing targets. Background Technology
[0002] Real-time recognition and tracking of fleeting targets primarily addresses the challenge of stable detection, recognition, and tracking of targets that appear briefly in the field of view, move at high speed, and then quickly disappear. Traditional visual perception mechanisms for handling fleeting targets largely rely on cameras that output images at a fixed frame rate. To capture high-speed targets, a high frame rate is typically required, which significantly increases the amount of data and puts pressure on transmission and processing; conversely, reducing the frame rate can easily miss the moment the target appears and produce motion blur. Because discrete sampling between frames loses continuous temporal information about the target's motion, motion detection based on frame difference is highly prone to failure when the target is moving at high speed.
[0003] Existing tracking algorithms generally employ motion prediction models based on the assumption of continuous smoothness. This assumption struggles to describe the sudden maneuvering behavior of non-cooperative targets, which mathematically manifest as discrete jumps in the state space, leading to significant discrepancies between predicted and actual positions. Relying solely on a single-modal sensor is severely inadequate in robustness under complex environments, and multimodal fusion methods often remain at the level of simple feature stitching, failing to fully utilize the complementary characteristics of different sensors in spatiotemporal resolution. The computational overhead of fusion is enormous, making it difficult to meet the real-time requirements of millisecond-level response. Existing servo control strategies fail to deeply couple sensor recognition confidence with platform dynamic characteristics. Particularly when facing the risk of targets potentially leaving the field of view or permanently disappearing, there is a lack of proactive avoidance mechanisms, resulting in extremely high tracking loss rates.
[0004] Therefore, overcoming the lag defects of existing visual prediction models and control algorithms to achieve millisecond-level accurate detection, prediction, and stable locking servo tracking of high-speed maneuvering targets that appear only briefly has become an urgent technical challenge. Summary of the Invention
[0005] The main objective of this invention is to provide a method and apparatus for real-time identification and tracking of flashing targets, aiming to overcome the lag defects of existing visual prediction models and control algorithms to achieve millisecond-level accurate detection, prediction, and stable locking servo tracking of high-speed maneuvering targets that flash briefly.
[0006] To achieve the above objectives, this invention proposes a real-time identification and tracking method for flashing targets, applied to a device comprising a visual tracking platform, a binocular collaborative sensing unit, and a recognition processing module. The binocular collaborative sensing unit includes an event camera and a visible light camera, and includes the following steps: The asynchronous event stream output by the event camera and the image output by the visible light camera are acquired. The asynchronous event stream is subjected to motion saliency-guided spatiotemporal voxelization to generate an event tensor. The event tensor is encoded and inferred through a hybrid stochastic differential equation consisting of a continuous diffusion process and discrete jump events driven by a Poisson process. The position probability distribution of the target's future position and the maneuver discrimination probability of the target performing a maneuver are output simultaneously. Based on the dynamic threat attention map and the location probability distribution, the computational resource allocation weight is calculated and obtained. Combined with the maneuver discrimination probability and historical state, a cognitive leap gate signal is generated. According to the cognitive leap gate signal, the spatial attention focus of the image output by the visible light camera is selected to extract appearance features by motion guidance, or a lightweight feature update is performed to calculate the target recognition confidence and occlusion probability. A nonlinear model prediction controller is constructed to control the visual tracking platform. The optimization objective function of the nonlinear model prediction controller includes a penalty term for the probability of the target leaving the field of view and a penalty term for the probability of the target permanently disappearing. The penalty term for the probability of the target leaving the field of view is determined by the spatial uncertainty of the position probability distribution and the field of view contraction effect caused by the platform motion. The penalty term for the probability of the target permanently disappearing is jointly generated by the product relationship between the decay of the recognition confidence and the occlusion probability. The nonlinear model predictive controller generates angle control commands to drive the visual tracking platform to track the target.
[0007] Preferably, when performing motion saliency-guided spatiotemporal voxelization on the asynchronous event stream, a contribution weight is assigned to each event in the asynchronous event stream. :
[0008] in, and This is the weighting adjustment coefficient; For the first The time interval between an event and its reference neighboring events; For image grayscale functions synchronized with the event stream, For the image at the event location Gradient magnitude at; To prevent the default positive number with a denominator of zero, the weighted event density of pixel positions within the sliding time window is calculated. When the weighted event density is greater than the adaptive threshold estimated by the mean and standard deviation of the background noise, the temporal resolution of the event tensor is reduced.
[0009] Preferably, in the hybrid stochastic differential equations composed of discrete jump events driven by the Poisson process, the differential equations describing instantaneous acceleration jumps include pulse maneuver event terms:
[0010] in, This is the acceleration damping coefficient; Here is the diffusion intensity matrix; For Wiener process; Total number of maneuver events; For the first The moment when the maneuver occurs; It is a one-dimensional Dirac impulse function; Let be the intensity of the maneuver mutation, where the intensity at the moment the maneuver occurs follows a non-homogeneous Poisson process, and the maneuver intensity is exponentially positively correlated with the square of the target's current velocity; when the maneuver discrimination probability exceeds a set threshold, the predicted covariance matrix of the position probability distribution is... Adaptively magnified:
[0011] in, Based on the prediction of the covariance matrix; This is the covariance amplification factor; The probability of the maneuver is defined as follows; This is the covariance increment matrix induced by the maneuver.
[0012] Preferably, the dynamic threat attention map The Gaussian probability distribution density of the target's predicted location and the local environmental entropy and historical trajectory memory items Common weighted generation:
[0013] in, For the Sigmoid function; Predict the mean value for the target's future location; To predict the covariance matrix; Background prior probability; , and For weight fusion.
[0014] Preferably, the historical trajectory memory item The reaction-diffusion-convection partial differential equation preserves the field for the evolutionary history trajectory of the tracked target on the image plane. The result of superposition is:
[0015] in, This is the memory decay time constant; The spatial diffusion coefficient; The target velocity vector; and These are the Laplacian operator and the spatial gradient operator, respectively. An energy injection term that is positively correlated with identification confidence and negatively correlated with target acceleration is used; the reappearance location of the target after occlusion is predicted based on the local maxima of the historical trajectory preservation field.
[0016] Preferably, the cognitive transition gating signal Calculate using the following formula:
[0017] in, For the Sigmoid function; Global average pooling for event feature maps; Global average pooling is used for the highest layer of the multi-scale appearance feature map; The maximum confidence level of the confidence heatmap identified in the previous frame; This is the cognitive transition gating signal from the previous frame; and These are learnable parameters.
[0018] Preferably, the penalty term for the probability of the target leaving the field of view in the optimization objective function is... and the penalty for the probability of the target permanently disappearing The calculation formulas are as follows:
[0019]
[0020] in, It is the standard normal cumulative distribution function; To predict the location; Center of the field of view; The effective radius of the field of view; This is the field-of-view contraction coefficient; This is the platform rotation speed command for the visual tracking platform. To predict the standard deviation of uncertainty; To identify confidence levels; This represents the occlusion probability.
[0021] This application also discloses a real-time identification and tracking device for flashing targets, including: A visual tracking platform, comprising a robotic arm with independent pitch and orientation degrees of freedom and a joint motor for driving the robotic arm to rotate; A binocular collaborative sensing unit is fixed to the end of the robotic arm. The binocular collaborative sensing unit includes an event camera and a visible light camera that are fixedly connected to each other. The recognition processing module is communicatively connected to the visual tracking platform and the binocular collaborative perception unit; the recognition processing module includes a processor and a memory storing a computer program, which, when executed, implements the real-time recognition and tracking method for flashing targets as described in any of the preceding claims.
[0022] Preferably, when generating angle control commands, the recognition processing module uses a pre-calibrated field-joint space mapping matrix. The motion vector of the target centroid on the pixel plane is converted into a basic angle increment, and an exponential decay factor related to the center distance is introduced:
[0023] in, and These are the angle increment commands for pitch and azimuth, respectively. and The pixel displacement of the target centroid in the horizontal and vertical directions; To predict the target position vector; The vector representing the center position of the image; This is the distance attenuation coefficient.
[0024] Preferably, when constructing the nonlinear model predictive controller, the identification and processing module models the dynamic response characteristics of the robotic arm as a nonlinear friction term containing LuGre terms. Second-order dynamic systems:
[0025] in, , For the internal state variables of the LuGre friction model, The first derivative of the internal state variable is... Angle command; From a practical perspective; This is the system's inherent frequency matrix; Here is the damping ratio matrix; , , These are nonlinear friction parameters.
[0026] The above technical solution has the following advantages: The method provided in this invention generates an event tensor by performing spatiotemporal voxelization on an asynchronous event stream guided by motion saliency. This process effectively suppresses complex background noise and significantly enhances the motion edge features of faintly appearing targets. The system uses a hybrid stochastic differential equation composed of discrete jump events driven by continuous diffusion and Poisson processes to infer the event tensor. This accurately predicts the target position probability distribution and maneuver discrimination probability. This processing greatly improves the predictive robustness against sudden maneuvers of non-cooperative targets. Subsequently, resource allocation weights are calculated based on the dynamic threat attention map and position probability distribution. The system combines the maneuver discrimination probability and historical state to generate a cognitive transition gating signal. This achieves spatial attention focusing and appearance feature extraction guided by motion in visible light images. This mechanism improves recognition accuracy while reducing computational redundancy. A nonlinear model predictive controller is then constructed, including penalties for the probability of the target leaving the field of view and the probability of the target permanently disappearing. This controller fully considers the spatial uncertainty of position prediction and the field of view contraction effect caused by platform motion. This allows the platform to accelerate and turn in advance before the target leaves the field of view. By combining the confidence decay and occlusion probability multiplication to generate a permanent disappearance probability penalty term, the system effectively distinguishes between temporary occlusion and permanent disappearance of the target. Finally, the system generates angle control commands to drive the vision tracking platform by solving a nonlinear model to predict the controller. This achieves extremely rapid response and stable tracking of high-speed flashing targets. Attached Figure Description
[0027] The present invention will now be described in detail with reference to specific embodiments and accompanying drawings, wherein: Figure 1 This is a schematic diagram of the real-time identification and tracking device for flashing targets provided in an embodiment of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. The illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention. Figure 1 As shown.
[0029] Example 1 This embodiment provides a real-time identification and tracking method for fleeting targets, applied to a real-time identification and tracking device for fleeting targets comprising a visual tracking platform, a binocular collaborative perception unit, and a recognition and processing module. In real-world applications, fleeting targets typically exhibit extremely short appearance times (e.g., less than 500ms), extremely high movement speeds (e.g., angular velocities exceeding 200° / s), and trajectories involving sudden maneuvers such as sharp turns, leaps, or evasive maneuvers. Furthermore, they are often accompanied by complex environmental conditions such as sudden changes in illumination, occlusion, and smoke interference. Traditional fixed-frame-rate visual perception mechanisms suffer from inherent latency and information redundancy, and traditional motion prediction models cannot effectively describe the sudden maneuvers of non-cooperative targets. To address these technical problems, this embodiment constructs a gaze-pointing servo tracking architecture that coordinates event vision with a robotic arm for visual servo capture, and proposes a three-level cascaded processing architecture inspired by a collaborative mechanism of peripheral perception and fine-grained center recognition in biological vision systems.
[0030] At the hardware architecture level, the visual tracking platform includes a robotic arm with independent pitch and azimuth degrees of freedom and joint motors that drive the robotic arm's rotation. The robotic arm is constructed using modular joint motor modules, with the two degrees of freedom completely decoupled, enabling independent control and avoiding tracking errors caused by mechanical coupling in traditional two-axis gimbals. A binocular collaborative perception unit is fixed to the end of the robotic arm, comprising an event camera and a visible light camera fixedly connected to each other. The event camera, acting as a dynamic visual sensor, simulates peripheral perception in biological vision systems, exhibiting high sensitivity to motion and rapid response; the visible light camera simulates fine central recognition in biological vision systems, possessing high resolution and responsible for extracting detailed features. The recognition processing module is communicatively connected to the visual tracking platform and the binocular collaborative perception unit. The internal memory of the recognition processing module stores a computer program, which, when executed by the processor, implements the real-time recognition and tracking method for flashing targets provided in this embodiment.
[0031] The real-time identification and tracking method for flashing targets provided in this embodiment specifically includes the following steps.
[0032] First, the recognition and processing module acquires the asynchronous event stream output by the event camera and the image output by the visible light camera, and performs motion saliency-guided spatiotemporal voxelization on the asynchronous event stream to generate an event tensor. Traditional fixed-time-window accumulation methods suffer from insufficient signal accumulation when the time window is too short, and motion blur when the time window is too long. This embodiment enhances the motion edge features of the flashing target by assigning contribution weights to each event in the asynchronous event stream. This contribution weight The calculation formula is:
[0033] in, and This is the weighting adjustment coefficient. For the first The time interval between an event and its reference neighboring events. For image grayscale functions synchronized with the event stream, For the image at the event location The gradient magnitude at that point. To prevent the use of default positive numbers with a denominator of zero.
[0034] To obtain the contribution weight of each event Then, the weighted event density of pixel positions within the sliding time window is calculated. Specifically, the mean background noise... variance of background noise The calculation formulas are as follows: The results are obtained through online estimation using an exponentially weighted moving average.
[0035] in, It is the exponential smoothing coefficient. The basic threshold used to distinguish between background events and salient events. This is the indicative function. For regions where the weighted event density is greater than the adaptive threshold, the recognition processing module determines that the region is a potential flashing target region, and then adaptively reduces the temporal bin resolution of the event tensor, specifically to 50. For the background region, the threshold is coarsened to 2ms. This adaptive threshold is estimated from the mean and standard deviation of the background noise. The final generated four-dimensional event tensor is... The normalized cumulative responses of positive and negative polarity events in this voxel are as follows:
[0036]
[0037] in, For the first The spatiotemporal region corresponding to an individual element The normalization constant is set. The above design can utilize the high spatiotemporal continuity of the event triggered by the edge of the flashing target movement to effectively suppress background noise under extremely low signal-to-noise ratio conditions and generate an event tensor that is highly sensitive to the flashing target.
[0038] Secondly, the recognition and processing module encodes the event tensor and infers its position using a hybrid stochastic differential equation consisting of a continuous diffusion process and discrete jump events driven by a Poisson process. Simultaneously, it outputs the probability distribution of the target's future position and the probability of the target maneuvering. Traditional prediction models based on uniform velocity or uniform acceleration are prone to failure when facing targets with sudden changes in direction at the millisecond level. This embodiment models the target's motion state as a hybrid state system, whose continuous evolution is governed by the formula... and The description, and the differential equation describing the instantaneous acceleration jump, contains impulsive maneuvering event terms:
[0039] in, This is the acceleration damping coefficient. This is the diffusion intensity matrix. This is the Wiener process. This represents the total number of maneuver events. For the first The moment when the maneuver occurs. It is a one-dimensional Dirac impulse function. Let be the intensity of the maneuver mutation. The intensity at the moment the maneuver occurs follows a non-homogeneous Poisson process, specifically expressed as:
[0040] in Based on basic mobility strength, The speed threshold scale indicates that the maneuver intensity is exponentially positively correlated with the square of the target's current speed. This reflects the physical fact that the higher the target speed, the more likely it is to undergo violent maneuvers, and the intensity of this maneuver abrupt change... It is also constrained by maximum maneuverability.
[0041] The recognition and processing module obtains the hidden state by performing 3D convolutional encoding on the event tensor. Then, two decoding branches are used to infer the continuous motion state and the discrete maneuver probability, respectively. The output of the continuous motion state is: Discrete maneuver probability output is When the inferred maneuver discrimination probability When the set threshold is exceeded, the model automatically switches to high-maneuver mode, and the predicted covariance matrix of the position probability distribution is... Adaptively magnified: ,in, Based on the prediction of the covariance matrix. This is the covariance amplification factor. The probability of motor discrimination. This is the covariance increment matrix induced by the maneuver. The final predicted output's position probability distribution not only indicates the most likely location of the target but also provides a quantified uncertainty confidence range, thus guiding the expansion of the subsequent search range when the target may perform a large maneuver.
[0042] Subsequently, the identification and processing module calculates and obtains computational resource allocation weights based on the dynamic threat attention map and position probability distribution. It then generates a cognitive transition gating signal by combining the maneuver discrimination probability and historical states. Based on this signal, it selects either motion-guided spatial attention focusing on the image output by the visible light camera to extract appearance features, or performs lightweight feature updates to calculate the target's recognition confidence and occlusion probability. This process simulates the collaborative mechanism of peripheral perception and fine-grained center recognition in biological vision systems. Specifically, the identification and processing module constructs a dynamic threat attention map on the image plane. It is composed of the Gaussian probability distribution density of the target prediction location and the local environmental entropy. and historical trajectory memory items Common weighted generation:
[0043] in, This is the Sigmoid function. The mean of the predicted future position of the target. To predict the covariance matrix. The background prior probability. , and For fusion weights. This is the local environment entropy. Local texture entropy, edge density entropy, and saliency entropy are all considered. The image plane is divided into an adaptive grid, with each grid... Allocated computing resources The specific weighting formula is proportional to the threat attention integral within the grid:
[0044] in, The total computing resource budget for the system. To ensure that each grid receives at least a minimal constant of non-zero resources.
[0045] To achieve cognitive transition between foveal and peripheral vision, cognitive transition gating signals are needed. Calculate using the following formula:
[0046] in, This is a global average pooling of the event feature map. This is the global average pooling of the highest layer of the multi-scale appearance feature map. The maximum confidence level of the confidence heatmap identified in the previous frame. This is the cognitive transition gating signal from the previous frame. and These are learnable parameters. When event features detect strong motion signals but appearance feature recognition confidence is low, the cognitive transition gating signal... When the value approaches 1, the system activates a deep cognitive mode, utilizing a motion-guided spatial attention map to spatially weight and modulate the channel dimensions of the image features output from the visible light camera. In this deep cognitive mode, the spatial attention map... The formula for generating it is:
[0047] Subsequently, it was upsampled and spatially weighted and focused on the appearance feature map to obtain... Simultaneously, based on motion characteristics, the RGB features are modulated in terms of channel dimensions to enhance motion-related feature channels, resulting in...
[0048] in This is the channel modulation matrix. When the target has been stably identified and the confidence level is high, the cognitive transition gating signal... When the value approaches 0, the system enters steady-state tracking mode, using only lightweight feature updates, the update formula of which is: .
[0049] The aforementioned dynamic threat attention map Historical trajectory memory items used Derived from the historical trajectory preservation field. The recognition and processing module evolves the historical trajectory preservation field for the tracked target on the image plane using reaction-diffusion-convection partial differential equations. The historical trajectory memory term is obtained by superposition. The partial differential equation is:
[0050] in, This is the memory decay time constant. is the spatial diffusion coefficient. This is the target velocity vector. and These are the Laplacian operator and the spatial gradient operator, respectively. This is an energy injection term that is positively correlated with recognition confidence and negatively correlated with target acceleration. A larger acceleration indicates more violent target maneuvering, and a weaker corresponding stabilization injection. Based on the local maxima of this historical trajectory-preserving field, the recognition processing module can effectively predict the reappearance location of the target after occlusion, solving the trajectory memory problem in cases of temporary target disappearance. The specific formula for calculating this reappearance location is as follows:
[0051] The predicted location and its confidence level will be fed back into the aforementioned threat attention map construction as an overlay of historical trajectory memory items.
[0052] Next, the recognition and processing module constructs a nonlinear model predictive controller (MMC) for controlling the visual tracking platform. Traditional control algorithms fail to address the specific risk of the target permanently disappearing, easily leading to meaningless blind tracking. In this embodiment, the optimization objective function of the MMC includes, in addition to the conventional error penalty and control quantity smoothing penalty, a penalty term for the probability of the target leaving the field of view and a penalty term for the probability of the target permanently disappearing. The penalty term for the probability of the target leaving the field of view... Determined by the spatial uncertainty of the position probability distribution and the field-of-view contraction effect caused by platform motion:
[0053] In this formula, It is the standard normal cumulative distribution function. To predict the location. It is the center of the field of view. Let be the effective radius of the field of view. is the field shrinkage coefficient. This is the platform rotation speed command for the vision tracking platform. The standard deviation of the predicted uncertainty is used as a penalty. This penalty ensures that when the predicted position is closer to the edge of the field of view, the uncertainty is greater, or the platform speed is higher, the controller will automatically accelerate and turn in advance, effectively preventing high-speed maneuvering targets from running out of the field of view.
[0054] Optimize the penalty term for the probability of the target permanently disappearing in the objective function. It is generated jointly by the product of the recognition confidence decay and the occlusion probability:
[0055] In this formula, To identify confidence levels. This is the occlusion probability predicted by the motion state and scene context. Based on this penalty, the system can accurately distinguish between temporary occlusion and permanent disappearance of a target. Once the probability of permanent disappearance of a target rises rapidly, the controller will automatically reduce the tracking weight, achieving extremely fast multi-target switching.
[0056] When constructing the aforementioned nonlinear model predictive controller, the recognition and processing module also models the dynamic response characteristics of the robotic arm as a nonlinear friction term containing LuGre terms. A second-order dynamic system is used to accurately describe the delay between command and actual movement of the joint motor:
[0057] in, , For the internal state variables of the LuGre friction model, The first derivative of the internal state variable is... This is an angle command. From a practical perspective. This is the system's inherent frequency matrix. Let be the damping ratio matrix. , , These are nonlinear friction parameters.
[0058] Finally, the recognition and processing module generates angle control commands by solving a nonlinear model to predict the controller, driving the vision tracking platform to track the target. During the conversion of the generated angle control commands, the recognition and processing module uses a pre-calibrated field-to-joint space mapping matrix. The motion vector of the target centroid on the pixel plane is converted into a basic angle increment, and an exponential decay factor related to the center distance is introduced:
[0059] in, and These are the angle increment commands for pitch and azimuth, respectively. and The pixel displacement of the target centroid in the horizontal and vertical directions. To predict the target position vector. This is the vector representing the center position of the image. This is the distance attenuation coefficient. The introduction of this exponential attenuation factor ensures that the control gain is automatically reduced as the target approaches the center of the field of view, minimizing the risk of overshoot oscillation in the robotic arm, while maintaining a high response level when the target deviates significantly. This direct conversion method based on dynamic prediction and kinematic mapping keeps the end-to-end delay of the visual perception of the mechanical response extremely low, thus achieving target acquisition. Furthermore, control commands are distributed to each joint motor module via a CAN bus-based distributed drive architecture, with a bus cycle configurable to 200 cycles. This ensures that the end-to-end latency is controlled within 1ms. Furthermore, the mapping matrix... During operation, online updates are performed using recursive least squares to adapt to mechanical wear, changes in connection gaps, and temperature drift. The entire platform's capture workflow strictly follows the four stages of biomimetic mapping: First, the alert stage, where the event camera continuously scans and waits for the target to appear; second, the locking stage, where the event signal triggers the visible light camera to complete accurate identification and centroid position calculation; next, the pointing stage, where centroid deviation directly drives the robotic arm to rotate and pull the target back to the center of the field of view; and finally, the holding stage, where the system enters fine-tuning tracking mode after the target stabilizes in the center, completing target capture.
[0060] Example 2 Building upon Example 1, this example further elaborates on the specific working process and variant schemes of the real-time identification and tracking device for flashing targets in multi-target rapid switching tracking scenarios and high-speed flying target monitoring scenarios. This example further enriches the internal network structure and multi-target collaborative scheduling strategy of the identification processing module.
[0061] The encoder architecture within the recognition and processing module includes a 3D convolutional encoder for processing event tensors and a visual transformer encoder for processing images output from the visible light camera. The spatiotemporal motion feature map is obtained by processing the event tensors using the 3D convolutional encoder. Visible light images are processed by a vision transformer encoder to obtain multi-scale appearance feature maps. .
[0062] When faced with a scenario where multiple targets appear alternately, the identification and processing module performs multi-target threat assessment and attention scheduling mechanisms. Each detected target... Assigned a dynamic threat level The calculation of this dynamic threat level comprehensively considers target type, movement status, and behavior pattern. The calculation formula is as follows:
[0063] in, The prior threat coefficient for the target type. This is the upper bound for velocity normalization. This is the upper bound for acceleration normalization. It is a symbolic function. The degree of proximity between the target and the protected target. This represents the current confidence level. to These are the weight parameters for each item.
[0064] The above degree of proximity The calculation formula is: in, For the goal Distance from the protected target. This is the distance attenuation scale.
[0065] The identification and processing module determines the dynamic threat level of all detected targets. Sort the targets, select the most threatening ones, and output their precise locations. The corresponding identification confidence scores are then used by the nonlinear model predictive controller. The location is calculated using the heatmap-weighted centroid method, with the specific formula as follows:
[0066]
[0067] in, To identify confidence level heatmaps. Set the sharpening factor. A value greater than 1 can effectively suppress the interference of low-confidence regions on centroid calculation, thereby improving the robustness of localization.
[0068] To adapt to mechanical wear, changes in connection gaps, and temperature drift in complex field environments, the recognition and processing module establishes a camera-to-gimbal joint parameter model that considers installation errors and temperature drift. This joint parameter model is specifically represented as follows:
[0069] in, For image plane pixel coordinates, For imaging scale factor, This is the temperature-dependent camera intrinsic parameter matrix. These represent the distance, azimuth, and elevation angles in the base coordinate system, respectively. and These are the rotation matrix and translation vector of the camera coordinate system relative to the base coordinate system, respectively. This is a Lie group exponent mapping used to describe small misalignment perturbations between the camera and the gimbal tip, where For perturbation Lie algebra parameters, These are the corresponding basis vectors. Perturbation parameters. The online update is achieved using the extended Kalman filter method, where the observation model is related to the perturbation parameters. The Jacobian matrix is denoted as:
[0070] in, It is a nonlinear observation function.
[0071] Regarding the constraints of the control law, the nonlinear model predictive controller, in addition to considering the target loss risk avoidance in Example 1, also comprehensively incorporates the physical limiting and thermal constraints of the joint motor. Specific constraints include speed constraints:
[0072] Acceleration constraints:
[0073] Power constraints:
[0074] And energy constraints:
[0075] The identification and processing module employs a fast distributed optimization algorithm based on the alternating direction multiplier method (ADMM) to solve the optimization problem in real time, with a solution cycle of less than 500. This generates smooth and physically achievable optimal control commands.
[0076] In high-speed target monitoring scenarios, incoming targets travel at high speeds. Traditional feedback controllers suffer from inherent feedback lag, resulting in low acquisition success rates. The system provided in this embodiment predicts the target's future position and, based on the aforementioned constraints and penalties, pre-turns, thereby improving acquisition success rates, reducing reacquisition delays, and decreasing energy consumption under simulation or testing conditions. Furthermore, in multi-target rapid switching tracking scenarios, traditional methods often require several seconds to confirm loss when a target briefly disappears, leading to continued blind tracking, wasting valuable processing time and generating numerous missed alarms. This embodiment dynamically determines the target's state based on the probability of permanent disappearance, rapidly switching to the next target within 30ms after a target is determined to be permanently disappeared. This demonstrates good engineering adaptability and robustness in complex environments with rapidly appearing targets.
[0077] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A real-time identification and tracking method for a flashing target, applied in a device comprising a visual tracking platform, a binocular collaborative sensing unit, and a recognition processing module, wherein the binocular collaborative sensing unit includes an event camera and a visible light camera, characterized in that, Includes the following steps: The asynchronous event stream output by the event camera and the image output by the visible light camera are acquired. The asynchronous event stream is subjected to motion saliency-guided spatiotemporal voxelization to generate an event tensor. The event tensor is encoded and inferred through a hybrid stochastic differential equation consisting of a continuous diffusion process and discrete jump events driven by a Poisson process. The position probability distribution of the target's future position and the maneuver discrimination probability of the target performing a maneuver are output simultaneously. Based on the dynamic threat attention map and the location probability distribution, the computational resource allocation weight is calculated and obtained. Combined with the maneuver discrimination probability and historical state, a cognitive leap gate signal is generated. According to the cognitive leap gate signal, the spatial attention focus of the image output by the visible light camera is selected to extract appearance features by motion guidance, or a lightweight feature update is performed to calculate the target recognition confidence and occlusion probability. A nonlinear model prediction controller is constructed to control the visual tracking platform. The optimization objective function of the nonlinear model prediction controller includes a penalty term for the probability of the target leaving the field of view and a penalty term for the probability of the target permanently disappearing. The penalty term for the probability of the target leaving the field of view is determined by the spatial uncertainty of the position probability distribution and the field of view contraction effect caused by the platform motion. The penalty term for the probability of the target permanently disappearing is jointly generated by the product relationship between the decay of the recognition confidence and the occlusion probability. The nonlinear model predictive controller generates angle control commands to drive the visual tracking platform to track the target.
2. The real-time identification and tracking method for flashing targets according to claim 1, characterized in that, When performing motion saliency-guided spatiotemporal voxelization on the asynchronous event stream, a contribution weight is assigned to each event in the asynchronous event stream. : in, and This is the weighting adjustment coefficient; For the first The time interval between an event and its reference neighboring events; For image grayscale functions synchronized with the event stream, For the image at the event location Gradient magnitude at; To prevent the default positive number with a denominator of zero, the weighted event density of pixel positions within the sliding time window is calculated. When the weighted event density is greater than the adaptive threshold estimated by the mean and standard deviation of the background noise, the temporal resolution of the event tensor is reduced.
3. The real-time identification and tracking method for flashing targets according to claim 1, characterized in that, In the hybrid stochastic differential equations composed of discrete jump events driven by Poisson processes, the differential equations describing instantaneous acceleration jumps include pulse maneuver event terms: in, This is the acceleration damping coefficient; Here is the diffusion intensity matrix; For Wiener process; Total number of maneuver events; For the first The moment when the maneuver occurs; It is a one-dimensional Dirac impulse function; Let be the intensity of the maneuver mutation, where the intensity at the moment the maneuver occurs follows a non-homogeneous Poisson process, and the maneuver intensity is exponentially positively correlated with the square of the target's current velocity; when the maneuver discrimination probability exceeds a set threshold, the predicted covariance matrix of the position probability distribution is... Adaptively magnified: in, Based on the prediction of the covariance matrix; This is the covariance amplification factor; The probability of the maneuver is defined as follows; This is the covariance increment matrix induced by the maneuver.
4. The real-time identification and tracking method for flashing targets according to claim 1, characterized in that, The dynamic threat attention map The Gaussian probability distribution density of the target's predicted location and the local environmental entropy and historical trajectory memory items Common weighted generation: in, For the Sigmoid function; Predict the mean value for the target's future location; To predict the covariance matrix; Background prior probability; , and For weight fusion.
5. The real-time identification and tracking method for a flashing target according to claim 4, characterized in that, The historical trajectory memory item The reaction-diffusion-convection partial differential equation preserves the field for the evolutionary history trajectory of the tracked target on the image plane. The result of superposition is: in, This is the memory decay time constant; The spatial diffusion coefficient; The target velocity vector; and These are the Laplacian operator and the spatial gradient operator, respectively. An energy injection term that is positively correlated with identification confidence and negatively correlated with target acceleration is used; the reappearance location of the target after occlusion is predicted based on the local maxima of the historical trajectory preservation field.
6. The real-time identification and tracking method for a flashing target according to claim 1, characterized in that, The cognitive transition gating signal Calculate using the following formula: in, For the Sigmoid function; Global average pooling for event feature maps; Global average pooling is used for the highest layer of the multi-scale appearance feature map; The maximum confidence level of the confidence heatmap identified in the previous frame; This is the cognitive transition gating signal from the previous frame; and These are learnable parameters.
7. The real-time identification and tracking method for a flashing target according to claim 1, characterized in that, The penalty term for the probability of the target leaving the field of view in the optimization objective function. and the penalty for the probability of the target permanently disappearing The calculation formulas are as follows: in, It is the standard normal cumulative distribution function; To predict the location; Center of the field of view; The effective radius of the field of view; This is the field-of-view contraction coefficient; This is the platform rotation speed command for the visual tracking platform. To predict the standard deviation of uncertainty; To identify confidence levels; This represents the occlusion probability.
8. A real-time identification and tracking device for a flashing target, characterized in that, include: A visual tracking platform, comprising a robotic arm with independent pitch and orientation degrees of freedom and a joint motor for driving the robotic arm to rotate; A binocular collaborative sensing unit is fixed to the end of the robotic arm. The binocular collaborative sensing unit includes an event camera and a visible light camera that are fixedly connected to each other. The identification processing module is communicatively connected to the visual tracking platform and the binocular collaborative perception unit; the identification processing module includes a processor and a memory storing a computer program, which, when executed, implements the real-time identification and tracking method for flashing targets as described in any one of claims 1 to 7.
9. The real-time identification and tracking device for flashing targets according to claim 8, characterized in that, When generating angle control commands, the recognition processing module uses a pre-calibrated field-joint space mapping matrix. The motion vector of the target centroid on the pixel plane is converted into a basic angle increment, and an exponential decay factor related to the center distance is introduced: in, and These are the angle increment commands for pitch and azimuth, respectively. and The pixel displacement of the target centroid in the horizontal and vertical directions; To predict the target position vector; The vector representing the center position of the image; This is the distance attenuation coefficient.
10. The real-time identification and tracking device for flashing targets according to claim 8, characterized in that, When constructing the nonlinear model predictive controller, the identification and processing module models the dynamic response characteristics of the robotic arm as a nonlinear friction term containing LuGre terms. Second-order dynamic systems: in, , For the internal state variables of the LuGre friction model, The first derivative of the internal state variable is... Angle command; From a practical perspective; This is the system's inherent frequency matrix; Here is the damping ratio matrix; , , These are nonlinear friction parameters.
Citation Information
Patent Citations
Multi-modal visual fusion complex scene small target detection tracking method and system
CN121438218A
Visual tracking of an object
US20160086344A1