A method and system for fast identification and tracking of low, slow and small targets

By constructing a large-scale video device array and a distributed computing architecture, the problem of difficulty in identifying and tracking small, slow targets in complex environments was solved, achieving fast and accurate target identification and tracking, reducing computational and network load, and improving the system's self-reflection ability and robustness.

CN121190518BActive Publication Date: 2026-03-24BEIJING HANGHUI DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing radar detection and photoelectric identification methods are difficult to effectively identify and track low-flying, slow-moving, and small targets, especially in complex environments with interference and obstruction. The accuracy and real-time performance of target detection are insufficient, and the unreasonable allocation of computing resources leads to a high false alarm rate, frequent switching of target identity, and a lack of self-reflection capabilities, making it impossible to meet the requirements for rapid identification and real-time tracking.

Method used

A large-scale video device array is constructed, employing a distributed computing architecture. Each video device independently and in parallel performs local processing and real-time detection, uploading only valid metadata. The monitoring server performs high-level collaborative computing to achieve multi-view data association and 3D positioning, and dynamically adjusts the status of the video devices to optimize the allocation of computing resources.

Benefits of technology

It enables rapid identification and tracking of low, slow, and small targets, reduces network transmission load, shortens response time, improves system accuracy and reliability, adapts to multi-target monitoring in swarm scenarios, and has self-sensing and self-diagnostic capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190518B_ABST
    Figure CN121190518B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of low slow small target fast identification tracking method and system;The method comprises: constructing the video device array for target area;Each video device has first state and second state, and independently and in parallel video stream acquisition and local processing are carried out;Based on current motion target information and historical motion target information, each motion target is carried out multi-view data association and three-dimensional positioning, and the associated video device set of each motion target is determined in real time;The advantage video device group of each motion target is determined in real time, and based on advantage video device group team motion target, fine identification and tracking are carried out.The present application realizes early warning, real-time identification, continuous tracking and threat assessment to unmanned aerial vehicle, model airplane, bird and other low slow small targets.The system can be widely applied to key scenes such as important place air defense, border patrol, major event security, airport clearance protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent recognition technology, and in particular relates to a method and system for rapid recognition and tracking of low-speed, small targets. Background Technology

[0002] Low-altitude, slow-moving, and small-sized flying targets refer to aircraft characterized by slow speed, small size, low altitude, and difficulty in detection by military and civilian radars. The identification and tracking of these targets is facing an increasingly urgent development need. This urgency stems from the severe security challenges posed by these targets; they are easy to acquire and control, and pose a potential danger. With the rapid development of the low-altitude economy, unmanned aerial vehicle (UAV) technology has been widely applied in various fields, such as logistics, agricultural monitoring, and security. A typical scenario is low-altitude, slow-moving, and small-sized UAVs. Their widespread use has brought significant challenges to traditional UAV detection and countermeasure systems. These UAVs are typically small in size, slow in speed, and have complex flight trajectories, often flying in complex environments such as cities and airports. This makes it difficult for existing detection technologies to effectively identify their presence, especially in the presence of significant interference and obstruction, where the accuracy and real-time performance of target detection often fail to meet requirements. Especially for swarms, the current focus is on developing high-resolution phased array radars and multiple-input multiple-receiver (MIRV) radars. The former uses agile beams to rapidly scan the airspace, attempting to resolve multiple points; the latter utilizes its virtual aperture characteristics to improve the ability to distinguish dense targets. However, radar still has inherent limitations: it is difficult to distinguish targets that are very close together, and it cannot perform identification or distinguish between drones and flocks of birds.

[0003] However, traditional radar detection methods struggle to detect targets with extremely small radar cross-sections and are difficult to classify precisely. For swarm radar, communication modes may employ anti-jamming techniques such as frequency hopping and networking, but these techniques become ineffective once the swarm performs its pre-programmed autonomous tasks. Single photoelectric identification methods are limited by field of view and environmental influences, making large-scale continuous monitoring difficult. Existing identification methods primarily rely on single sensors or simple multi-sensor patchwork, which has drawbacks. First, these methods have shallow data fusion levels, failing to fundamentally solve the correlation problem between observation data from different sensors in geometric space, leading to high false alarm rates and frequent target identification switching. Second, computational resource allocation is unreasonable, often employing a coarse-grained mode of full-load operation across all sensors, deeply processing a large amount of invalid or low-value information, resulting in huge computational power waste and making it difficult to support large-scale, long-term deployment applications. More importantly, existing systems lack the ability to perceive their own monitoring status, unable to determine in real time whether tracking quality has deteriorated or correlation is incorrect; their monitoring process is a black box that cannot self-examine, raising questions about its reliability. More importantly, the significant reduction in the cost of image sensors, processors, and communication hardware in recent years has laid a solid material foundation for a paradigm shift in low-speed, small-target recognition and tracking technology. This trend has directly given rise to two key possibilities: First, the democratization of hardware costs has enabled the deployment of large-scale, high-density camera arrays from expensive proof-of-concept to large-scale engineering applications, providing the physical prerequisite for achieving wide-area, multi-angle, and blind-spot-free collaborative perception. Second, the decline in computing hardware prices has made it possible to move computing power from a central server to the edge of each camera node, fundamentally overturning the traditional acquisition-transmission-centralized processing model. Architecture; Based on the aforementioned hardware changes, this invention improves upon traditional technologies, with its core being the distributed array setup and hierarchical pre-processing of computation. Traditional solutions, limited by cost, typically employ a small number of high-priced sensors and transmit raw data back to a central server for centralized processing via a network. This post-processing model inevitably introduces network latency and bandwidth bottlenecks, resulting in slow system response and difficulty in meeting the requirements for rapid identification and real-time tracking of small, slow targets. When faced with numerous, small, and coordinated bee swarms, traditional single-sensor systems are prone to missed detections and misjudgments due to narrow field of view, computational saturation, or feature confusion.

[0004] This invention fully leverages hardware cost advantages by deploying a large-scale camera array and completely circumvents the aforementioned problems by layering and prioritizing computational tasks: each video device node independently and in parallel completes lightweight computations such as local moving target detection and preliminary identification, thereby achieving partial computational pre-processing. Only effective metadata, rather than massive video streams, is uploaded, while maintaining complex computational functions for the advantageous video devices. Subsequently, higher-level collaborative scheduling computation is performed on the monitoring server, greatly reducing network transmission load and shortening the closed-loop response time from data acquisition to decision output. This enables the entire system to react to high-speed, maneuvering, slow, and small targets in near real-time with unprecedented efficiency, thus achieving dual optimization of performance and cost. In particular, it demonstrates unique technical advantages and excellent detection capabilities when dealing with swarm scenarios of slow, small targets. Through the collaborative layout of a large-scale distributed video device array, it achieves wide-area seamless coverage of the airspace at the detection level. Its multi-view, overlapping observation characteristics make it difficult for individual targets in the swarm to hide using blind spots. Through the almost simultaneous detection information of the array structure, the overall appearance of the swarm can be quickly perceived and confirmed, and its size and approximate spatial distribution can be preliminarily judged. Summary of the Invention

[0005] To address the aforementioned problems in the prior art, this invention proposes a method and system for fast identification and tracking of slow, small targets, the method comprising:

[0006] Step S1: Construct a video device array for the target area; the video device array contains multiple video devices; the field of view of each video device in the video device array may or may not overlap;

[0007] Step S2: Each video device has a first state and a second state, and performs video stream acquisition and local processing independently and in parallel; each video device is stationary in the first state and performs real-time detection of moving targets within its field of view; when it is one of the dominant video device groups for a moving target, it enters and maintains the second state to perform high-precision identification and tracking of the moving target.

[0008] Step S3: All video devices send the current moving target information to the monitoring server in real time; based on the current moving target information and historical moving target information, multi-view data association is performed on each moving target to determine the associated video device set for each moving target in real time; the associated video device set is the set of video devices that can effectively observe the moving target.

[0009] Step S4: Based on the associated video device set, determine the dominant video device group for each moving target in real time. When a change in the dominant video device group is detected, send a state switching message to the current dominant video device group and the new dominant video device group. The corresponding video device in the current dominant video device group closes the second state and maintains the first state for the changed moving target. The corresponding video device in the new dominant video device group opens the second state and maintains the first state for the changed moving target.

[0010] Furthermore, a group of superior video devices is used to perform high-precision identification and tracking of the moving target. In the first state, each video device detects moving targets within its detection range. Specifically, for a known global moving target, video devices are selected from all its associated video device sets to form a group of superior video devices. The group of superior video devices is used for high-precision identification and tracking of the moving target. The group of superior video devices includes a primary superior video device and secondary superior video devices. The primary superior video device is used for high-confidence identification and image forensics. The secondary superior video devices are used in conjunction with the primary superior video device to perform stereo vision triangulation and output high-precision three-dimensional coordinates in real time.

[0011] Furthermore, a unified second state is set for all moving targets in each video device, and entry and maintenance are managed. Under the unified setting, when the video device does not enter the second state due to any moving target, it enters the second state for the first time and maintains the second state upon receiving a state switching message. When maintaining the second state, if a state switching message is received and it is necessary to close the second state for a moving target, only the second state setting maintained by the video device due to that moving target needs to be closed, while the second state setting maintained due to other moving targets is retained. If the video device maintains the second state only due to one moving target, the second state can be closed directly.

[0012] Furthermore, in each video device, a second state is independently set for each moving target, and its entry and maintenance are managed. In the case of independent setting, the second state associated with the moving target is entered and maintained, or closed, based on the state switching message; when the second state is maintained, the superior video device group performs high-precision identification and tracking of the moving target.

[0013] Furthermore, the real-time determination of the advantageous video device group for each moving target specifically involves: calculating the advantage score of each video device in the associated video devices based on the size of the moving target, the centering degree, the image clarity, and the resolution of the video devices; selecting the two video devices with the highest advantage scores to form the advantageous video device group, wherein: the one with the highest advantage score is the primary advantageous video device, and the one with the second highest advantage score is the secondary advantageous video device.

[0014] Furthermore, the moving target information includes one or more of image coordinates, features, and timestamps; the detection server performs timestamp alignment on all received moving target information.

[0015] Furthermore, the monitoring server determines the sequence of associated video devices for each moving target based on the set of associated video devices for each moving target; and performs real-time anomaly monitoring based on changes in the sequence of associated video devices.

[0016] A platform for rapid identification and tracking of low-speed, small targets, the platform being used to implement the aforementioned method for rapid identification and tracking of low-speed, small targets.

[0017] A server for fast identification and tracking of slow, small targets includes a processor coupled to a memory. The memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the fast identification and tracking method for slow, small targets is implemented.

[0018] A fast identification and tracking system for low-speed, small targets, the system being used to implement the fast identification and tracking method for low-speed, small targets.

[0019] A computer-readable storage medium includes a program that, when run on a computer, causes the computer to perform the described method for fast identification and tracking of slow, small targets.

[0020] The beneficial effects of this invention include:

[0021] (1) By constructing a redundant full coverage range through a video device array, partial computation is pre-processed, only valid metadata is uploaded instead of massive video streams, and complex computation functions are maintained for the advantageous video devices. Each video device works in the first state and the triggerable second state, providing fine classification for each moving target, independently maintaining the associated video device set and its advantageous video device group for each moving target, thereby providing differentiated customized recognition and tracking in the overall low, slow and small target fast recognition and tracking process without significantly increasing the hardware and software overhead, and providing a basis for differentiated parameter settings;

[0022] (2) Make full use of big data information, perform multi-view data association and three-dimensional positioning for each moving target based on current moving target information and historical moving target information, determine the associated video device set for each moving target in real time, and support anomaly monitoring based on simple quantitative calculation on the basis of the associated video device set, sensitively detect the unexpected trajectory changes, disappearance and appearance of moving targets within the target range, and supplement the robustness of identification. Attached Figure Description

[0023] The accompanying drawings, which are provided to further illustrate the invention and form part of this application, are not intended to unduly limit the invention. In the drawings:

[0024] Figure 1 This is a schematic diagram of the fast identification and tracking method for small, slow targets provided by the present invention. Detailed Implementation

[0025] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions are only used to explain the present invention and are not intended to limit the present invention.

[0026] As attached Figure 1 As shown, the method includes the following steps:

[0027] Step S1: Construct a video device array targeting the target area; the video device array contains multiple video devices; the field of view of each video device in the video device array may overlap or not overlap; the field of view of all video devices in the video device array covers the target area;

[0028] Preferred method: Calibrate and initialize the video device array and each video device therein; specifically: perform intrinsic parameter calibration on each video device to determine its intrinsic parameter matrix, such as focal length, principal point coordinates, and distortion coefficients; perform extrinsic parameter calibration on the video device array;

[0029] Preferred method: Use a calibration board such as a checkerboard to complete the intrinsic parameter calibration; perform extrinsic parameter calibration on the video device array, place one or more obvious calibration objects or move one or more feature points in the common field of view of all video devices, i.e., the target area, calculate the multi-view correspondence of each relative to the calibration object, derive the view relationship between the video devices, select a fixed geographical coordinate point as the origin of the world coordinate system, uniformly transform the extrinsic parameters of all video devices to this coordinate system, and further use rotation matrix and translation vector to characterize the relative position and attitude relationship between the video devices to complete the extrinsic parameter calibration;

[0030] Step S2: Each video device has a first state and a second state, and performs video stream acquisition and local processing independently and in parallel; each video device is stationary in the first state and performs real-time detection of moving targets within its field of view; when it is one of the advantageous video device groups for a specific moving target, it simultaneously enters and maintains the second state to perform high-precision identification and tracking of the specific moving target; the advantageous video device group for the specific moving target includes the video device that enters and maintains the second state.

[0031] Preferred method: Each video device uses calibrated intrinsic parameters to perform real-time lens distortion correction to ensure image geometric accuracy; a background segmentation model is used to extract foreground moving pixel blocks; the foreground blocks are clustered and filtered to obtain candidate regions for moving targets; the candidate regions are input into the coarse classification module to obtain a coarse classification of the moving targets;

[0032] Preferably, the coarse classification module only needs to determine whether the moving target is a low-speed, small target; the low-speed, small targets identified by classification include predetermined types, such as birds, airplanes, and drones, and do not require fine classification; for confirmed moving targets, their two-dimensional image coordinates, appearance features, timestamps, and video device IDs are extracted as moving target information.

[0033] Step S3: Based on the current and historical moving target information, perform multi-view data association and 3D positioning for each moving target, and determine the associated video device set for each moving target in real time; the associated video device set consists of video devices capable of effectively observing the moving target; specifically, all video devices send the moving target information to the monitoring server in real time; the monitoring server performs time synchronization, alignment, and data association on the moving target information; wherein: the moving target information includes image coordinates, features, timestamps, etc.; the timestamp alignment of all moving target information can be performed to address network latency and video array node calibration;

[0034] The multi-view data association specifically involves: performing geometric constraint association and / or appearance feature association; and assigning a unique identifier to each moving target after association.

[0035] The geometric constraint association is specifically performed as follows: using calibrated extrinsic parameters, the image coordinates of the target in video device A are projected backwards into the coordinate system to form an observation ray. It is determined whether the observation ray passes near the target point on the imaging plane of video device B. If the observation rays of multiple video devices for the same moving target intersect at a point (a local area) in the spatial coordinate system, the association is successful, and the multiple video devices are set as the associated video device set for the same moving target. The associated video device set consists of video devices that can effectively observe the moving target. Using the moving target information and stereo vision triangulation of the associated video device set, the three-dimensional coordinates (X, Y, Z) of the moving target in the coordinate system are solved. The associated video device set can be filtered using the length of the observation ray or the two-dimensional image coordinate information of the moving target, so that the one with the shorter length of the observation ray or the larger span of the two-dimensional image coordinates of the moving target (the size of the moving target's border) enters the associated video device set of the moving target.

[0036] Further, time-based geometric constraint filtering is performed, specifically: the historical 3D coordinates of the moving target are used to predict its position coordinates, and its next-moment position coordinates are obtained as the predicted position. The predicted position is then projected onto the image plane of each video device in the associated video device set to obtain the predicted pixel coordinates. For each associated video device, its predicted pixel coordinates are compared with its current 2D image coordinates. If the distance between the two is less than a distance threshold, the associated video device is kept in the associated video device set of the moving target; otherwise, it is removed from the associated video device set. The distance threshold here can be set relatively loosely.

[0037] The process of associating appearance features specifically involves: calculating the cosine similarity or Euclidean distance between the appearance feature vectors of the moving target detected by different video devices; if the similarity or distance obtained by multiple video devices is small, the association is successful, and the multiple video devices are set as the associated video device set for the same moving target; the appearance feature vector of the moving target can be obtained by the video device in the first state or further calculated by the monitoring server based on the two-dimensional image coordinates of the moving target;

[0038] Alternative: When associating based on both geometric constraints and appearance features, perform joint probabilistic data association using JPDA to improve association accuracy in cases of occlusion or field of view edges;

[0039] Step S4: In real time, determine the dominant video device group for each moving target. When a change in the dominant video device group is detected, send a state switching message to the current dominant video device group and the new dominant video device group. The corresponding video device in the current dominant video device group closes its second state and maintains its first state, while the corresponding video device in the new dominant video device group opens its second state and maintains its first state. Clearly, the corresponding video device is the part that has changed. That is, the dominant video device group is used for high-precision identification and tracking of the moving target, with each video device detecting moving targets within its detection range. Specifically: for a known global moving target, select video devices from all its associated video device sets to form a dominant video device group. The dominant video device group is used for high-precision identification and tracking of the moving target. The dominant video device group includes a primary dominant video device and secondary dominant video devices. The primary dominant video device is used for high-confidence identification and image forensics; the secondary dominant video devices are used in conjunction with the primary dominant video device to perform stereoscopic vision triangulation and output high-precision three-dimensional coordinates (X, Y, Z) in real time.

[0040] Clearly, a second state can be set independently for each moving target, and its entry and maintenance can be managed. Alternatively, a unified second state can be set for all moving targets, and its entry and maintenance can be managed accordingly. In the case of independent settings, the entry and maintenance of the second state associated with the moving target, or its deactivation, is based on a state switching message. In the case of a unified setting, when the video device has not entered the second state due to any moving target, it enters and maintains the second state upon receiving a state switching message. While maintaining the second state, if a state switching message is received requiring the second state to be deactivated, only the second state setting maintained by that moving target needs to be deactivated, while the second state settings maintained by other moving targets remain. If the video device maintains the second state only for one moving target, the second state can be deactivated directly. Each moving target can be managed using a state list corresponding to the moving target.

[0041] Replaceable: Both the primary and secondary dominant video devices are used for high-confidence identification and image forensics, as well as high-precision 3D coordinates; the primary and secondary dominant video devices are redundant to each other.

[0042] Replaceable: The superior video device group contains three superior video devices, all of which are used for high-confidence recognition and image forensics, as well as high-precision three-dimensional coordinates; the recognition results of the three devices are voted on to determine the final recognition result;

[0043] The real-time determination of the dominant video device group for each moving target specifically involves: calculating the advantage score of each video device in the associated video devices based on the size of the moving target, its centering degree, image clarity, and video device resolution; selecting the two video devices with the highest advantage scores to form the dominant video device group, where the one with the highest advantage score is the primary dominant video device, and the one with the second highest advantage score is the secondary dominant video device; real-time monitoring of video frames from all video devices; re-determining the dominant video device group when the determination conditions are met; and sending a state switching message to the current dominant video device group and the new dominant video device group when a change in the dominant video device group is detected. In the current dominant video device group, the corresponding video device closes its second state and remains in its first state, while in the new dominant video device group, the corresponding video device opens its second state and remains in its first state. The determination conditions are reaching a preset number of frames or the moving target approaching the edge of the field of view of the corresponding video device in the dominant video device group. Therefore, the first state is the constant state of the video device.

[0044] The advantage score of each video device k in the associated video devices is calculated based on the size of the moving target, the centering degree, the image sharpness, and the resolution of the video device. Specifically, the size score is calculated based on the following formulas (1)-(3). Centered rating and clarity rating and determine the resolution score. Perform linear or nonlinear normalization to obtain the corresponding normalized value. , , , ; Calculate the advantage score based on the normalized value and formula (4) ;in: This is the scoring coefficient, a preset value with a sum of 1. It can be adjusted according to specific application scenarios; increasing the scoring coefficient will correspondingly increase the scoring coefficient to improve recognition accuracy. and Improving stability will improve , These are relatively fixed static values; and These are the border width and border height of the moving target in the video device k, respectively; It is the size threshold of the moving target, which is a preset value; It is a parameter that controls the rate at which the score decays; it is an adjustable hyperparameter. and These are the coordinates of the center point of the bounding box of the moving target; and These are the coordinates of the center point of the image border area; , , These are sharpness, brightness, and contrast. It refers to resolution;

[0045] (1);

[0046] (2);

[0047] (3);

[0048] (4);

[0049] The superior video device group is used for high-precision identification and tracking of moving targets; specifically: the main superior video device uses an accurate identification model to perform fine classification of moving targets and tracks them based on the fine classification results; it records a high-definition video stream of the tracking process and extracts the clearest close-up image of the target for post-event evidence collection and analysis; for example: using deep learning models such as YOLOv8 and Faster R-CNN for fine classification to identify that it is a DJI Mavic 3 rather than just a drone;

[0050] Preferably, the method further includes: monitoring the server to perform anomaly monitoring based on the changes in the associated video device set cls for each moving target; specifically, calculating the degree of set change based on the following formula (5). When the change in the set exceeds the change threshold, an anomaly is determined to have occurred; after an anomaly occurs, an alarm is triggered; in response to the anomaly alarm, the monitoring server switches all video devices in the set of associated video devices of the moving target to enter the second state; where: cls is the current set of associated video devices, and becls is the previous set of associated video devices; This determines the number of elements in the set, which is the set size; the threshold for the degree of change is a preset value; for example, it can be set to 5~50%; of course, the setting of this threshold is related to the relationship between the target area and the size of the video device array, the size of each video device and the size of the video device array; it is also related to the sensitivity of the abnormal alarm.

[0051] (5);

[0052] To further improve the sensitivity of anomaly detection, the method further includes: the monitoring server determining the sequence of associated video devices for each moving target based on the set of associated video devices for each moving target; and performing real-time anomaly detection based on changes in the sequence of associated video devices.

[0053] The determination of the associated video device sequence for each moving target based on the associated video device set for each moving target specifically involves arranging each video device in the associated video device set from largest to smallest based on the magnitude of the advantage score to form the associated video device sequence.

[0054] The real-time anomaly monitoring based on changes in the associated video device sequence specifically involves: vectorizing the associated video device sequence to construct a feature vector, where each element of the feature vector corresponds to a fixed video device in the video array; and placing the position of the video device in the associated video device sequence into the feature vector. The corresponding positions are used to obtain the instantiated feature vector; missing values ​​are set to default values, such as: maximum value or 0; missing values ​​indicate that the video device does not appear in the current associated video device sequence; the sequence order variability and sequence entropy are calculated based on the feature vector, and anomaly detection is performed based on the sequence order variability and dispersion.

[0055] The anomaly monitoring based on sequence order variability and dispersion is specifically: calculating a first sequence order variability and a second sequence order variability to measure the overall structural change of the reaction sequence, and performing anomaly monitoring based on the first and / or second sequence order variability.

[0056] The calculation of the order change degree of the first sequence Specifically, the degree of change in the order of the first sequence is calculated based on the following formulas (6)-(8). ;in: It is the current feature vector. It is the previous feature vector; The monitoring period is k; k is the video device number. ; When ρ = 1, it indicates no change; ρ = -1 indicates the exact opposite; ρ ≈ 0 indicates no monotonic relationship; when ρ = 1, = 0, when ρ = -1, =1; obviously, The closer to 0, the smaller the degree of order variation; the closer to 1, the larger the degree of order variation; K is the number of video devices; it can also be the maximum value of the identifier;

[0057] (6);

[0058] (7);

[0059] (8);

[0060] The calculation of the second sequence order change degree Specifically, the degree of change in the order of the second sequence is calculated based on the following formulas (9)-(10). ;

[0061] (9);

[0062] (10);

[0063] Anomaly monitoring is performed based on the degree of change in the first sequence and / or the degree of change in the second sequence. When the degree of change in the first sequence is greater than or equal to the first degree of change threshold and / or the degree of change in the second sequence is greater than or equal to the second degree of change threshold, an anomaly is determined to have occurred. In response to the anomaly, the monitoring server switches all video devices in the associated video device set of the moving target to enter the second state. The degree of change threshold can be determined and preset based on historical data. For example, the first degree of change threshold is set to 0.05~0.3. The setting of the second degree of change threshold is related to the arrangement of the video device array, the field of view size, the speed of movement of small targets, etc., and can be modulated to a unified, experience-based relative threshold (such as 0.4~0.7) based on simulation data for judgment.

[0064] Preferred approach: The variability threshold is determined based on the recognition results of the superior video device group on the moving target; when the recognition results of the moving target show that its speed and trajectory change significantly, a relatively large variability threshold is set; conversely, a smaller variability threshold can be set. In other words, specific parameters in the recognition device and its recognition strategy can be customized for each moving target, thereby providing differentiated customized recognition and tracking in the overall process of fast recognition and tracking of slow, small targets without significantly increasing hardware and software overhead; in addition, when a moving target is found to have an unexpected trajectory change, disappearance, or appearance within the target range, a drastic change in the sequence or set will occur, thereby supporting anomaly monitoring based on simple quantitative calculations on the basis of the associated video device set, sensitively detecting situations such as unexpected trajectory changes, disappearance, and appearance of moving targets within the target range, supplementing the robustness of recognition; of course, a tiered threshold judgment can be provided for the variability and degree of change to classify different anomaly levels, which will not be detailed here;

[0065] Preferred method: Determine the dominant video device group for each moving target in real time based on the first cycle; perform anomaly monitoring based on changes in the sequence of associated video devices in real time based on the second cycle; wherein: the first cycle is less than or equal to the second cycle;

[0066] More importantly, for small, slow-moving targets in bee colonies, the system's dynamic advantage video device set and its combination strategy make it naturally suitable for the highly dynamic, multi-target computing environment of bee colonies. The front-end processing capability of each video device node can process multiple targets within the field of view in parallel, greatly alleviating the computing pressure on the central monitoring server. It can intelligently and dynamically allocate optimal computing resources to different subsets or key individuals in the bee colony, ensuring the most comprehensive perception and the most accurate tracking even in areas with the highest bee colony density. This enables multi-granular, integrated monitoring of the overall macroscopic behavior of the bee colony and the microscopic movement of individuals, providing solid and comprehensive information support for subsequent early warning and decision-making responses.

[0067] Based on the same inventive concept, the present invention also provides a fast identification and tracking system for low, slow, and small targets. The system is used to implement the above-mentioned fast identification and tracking method for low, slow, and small targets. The system includes a monitoring server and a video device array. The video device array includes multiple video devices, and each video device is connected to the monitoring server via wired or wireless means.

[0068] Based on the same inventive concept, the present invention also provides a server for fast identification and tracking of low, slow, and small targets, wherein the server is equipped with the aforementioned monitoring server for implementing the aforementioned method for fast identification and tracking of low, slow, and small targets; the monitoring server is a cloud server.

[0069] Based on the same inventive concept, the present invention also provides a device for rapid identification and tracking of low, slow and small targets, the device being used to implement the above-mentioned method for rapid identification and tracking of low, slow and small targets;

[0070] Based on the same inventive concept, the present invention also provides a fast identification and tracking platform for low, slow and small targets, the platform being used to implement the above-mentioned fast identification and tracking method for low, slow and small targets;

[0071] The aforementioned system achieves precise correlation and fusion of observation data in three-dimensional space through accurate spatiotemporal calibration and multi-view geometric constraints, fundamentally ensuring the uniqueness of target identity and the continuity of trajectory, greatly reducing the probability of false correlation. By introducing a dynamic advantage camera selection strategy, only a few cameras with the best viewpoints undertake the high-load fine recognition and localization tasks, while other cameras remain in a low-power standby state, thereby doubling the computational efficiency and making the deployment of large-scale camera arrays possible. Finally, the system has powerful self-sensing and self-diagnostic capabilities. By transforming the abstract sequence of correlated video devices into quantifiable feature vectors and performing real-time time series analysis, the system can keenly capture abnormal changes in the tracking link, achieving a leap from passive monitoring to active early warning, significantly improving the system's reliability and robustness. In summary, this method, through deep fusion, intelligent collaboration, and resource optimization, constructs an efficient, reliable, and scalable low-speed small target recognition and tracking solution.

[0072] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including assembly or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to said program, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.

[0073] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0074] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0075] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0076] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for fast identification and tracking of small, slow targets, characterized in that, The method includes: Step S1: Construct a video device array for the target area; the video device array contains multiple video devices; the field of view of each video device in the video device array may or may not overlap; Step S2: Each video device has a first state and a second state, and performs video stream acquisition and local processing independently and in parallel; each video device is stationary in the first state and performs real-time detection of moving targets within its field of view; when it is one of the dominant video device groups for a moving target, it enters and maintains the second state to perform high-precision identification and tracking of the moving target. Step S3: All video devices send the current moving target information to the monitoring server in real time; based on the current moving target information and historical moving target information, multi-view data association is performed on each moving target to determine the associated video device set for each moving target in real time; the associated video device set is the set of video devices that can effectively observe the moving target. Step S4: Based on the associated video device set, determine the dominant video device group for each moving target in real time. When a change in the dominant video device group is detected, send a state switching message to the current dominant video device group and the new dominant video device group. The corresponding video device in the current dominant video device group closes the second state and maintains the first state for the changed moving target. The corresponding video device in the new dominant video device group opens the second state and maintains the first state for the changed moving target.

2. The method for fast identification and tracking of low-speed, small targets according to claim 1, characterized in that, The moving target is identified and tracked with high precision using a group of dominant video devices. In the first state, each video device detects moving targets within its detection range. Specifically, for a known global moving target, video devices are selected from all associated video devices to form a group of dominant video devices. This group of dominant video devices is used for high-precision identification and tracking of the moving target. The group of dominant video devices includes a primary dominant video device and secondary dominant video devices. The primary dominant video device is used for high-confidence identification and image forensics. The secondary dominant video devices are used in conjunction with the primary dominant video device to perform stereoscopic vision triangulation and output high-precision three-dimensional coordinates in real time.

3. The method for fast identification and tracking of low-speed, small targets according to claim 2, characterized in that, In each video device, a unified second state is set for all moving targets, and entry and maintenance are managed. Under the unified setting, when the video device enters the second state upon receiving a state switching message, it first enters the second state and maintains it. When maintaining the second state, if a state switching message is received and it is necessary to close the second state for a moving target, only the second state setting maintained by the video device for that moving target needs to be closed, while the second state settings maintained by other moving targets are retained. If the video device maintains the second state only for one moving target, the second state can be closed directly.

4. The method for fast identification and tracking of small, slow targets according to claim 2, characterized in that, Each video device independently sets a second state for each moving target and manages its entry and maintenance. When set independently, the second state associated with the moving target is entered or maintained, or closed, based on the state switching message. When the second state is maintained, the superior video device group performs high-precision identification and tracking of the moving target.

5. The method for fast identification and tracking of slow, small targets according to claim 4, characterized in that, The method for determining the dominant video device group for each moving target in real time is as follows: based on the size of the moving target, the centering degree, the image clarity and the resolution of the video device, the advantage score of each video device in the associated video device is calculated, and the two video devices with the highest advantage scores are selected to form the dominant video device group, wherein: the one with the highest advantage score is the primary dominant video device, and the one with the second highest advantage score is the secondary dominant video device.

6. The method for fast identification and tracking of small, slow targets according to claim 5, characterized in that, in: The moving target information includes one or more of image coordinates, features, and timestamps; the monitoring server aligns the timestamps of all received moving target information.

7. The method for fast identification and tracking of small, slow targets according to claim 5, characterized in that, The monitoring server determines the sequence of associated video devices for each moving target based on the set of associated video devices for each moving target; and performs real-time anomaly monitoring based on changes in the sequence of associated video devices.

8. A server for fast identification and tracking of small, slow targets, characterized in that, The method includes a processor coupled to a memory, the memory storing program instructions, which, when executed by the processor, implement the fast identification and tracking method for slow, small targets as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, Includes a program that, when run on a computer, causes the computer to perform the method for fast identification and tracking of slow, small targets as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Multiresolution large visual field angle high precision photogrammetry apparatus

    CN105066962A

  • Multi-camera cooperative visual tracking system and method for invading animals of transformer substation

    CN119762529A