Two-way moving target trajectory recognition analysis device and method based on video stream
Patent Information
- Application Number
- CN202610554278.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-24
- Publication Date
- 2026-08-28
AI Technical Summary
[0003]本申请提供了基于视频流的双向运动目标轨迹识别分析装置及其方法,用于解决现有技术在复杂场景下难以有效处理目标遮挡和双向运动方向误判,轨迹识别的准确性和可靠性不足的技术问题
本申请提供的基于视频流的双向运动目标轨迹识别分析装置及其方法,涉及图像处理技术领域,通过视频流数据采集和目标检测获取目标边界框、物种类型和置信度,分析帧间关联得到运动轨迹并判断遮挡情况,并自适应调整置信度搜索范围步长,恢复目标运动轨迹,并根据虚拟门控线判断目标的运动方向,确保高精度的目标计数和运动方向判定,解决了现有技术在复杂场景下难以有效处理目标遮挡和双向运动方向误判,轨迹识别的准确性和可靠性不足的技术问题,实现了通过遮挡分析和虚拟门控线判定,提高复杂环境下的轨迹识别的准确性和可靠性的技术效果。
Smart Images

Figure CN122657782A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically to a device and method for bidirectional moving target trajectory recognition and analysis based on video streams. Background Technology
[0002] In the field of modern video surveillance and analysis, accurate identification and analysis of target motion trajectories are widely used in various scenarios such as traffic management, security monitoring, and ecological monitoring. However, target tracking in complex environments faces many challenges, such as target occlusion, changes in lighting, and background interference. These problems seriously affect the accuracy and continuity of trajectory recognition. Although existing technologies have solved some of these problems to a certain extent, they still have shortcomings in the trajectory recognition of bidirectional moving targets, especially in cases of frequent target occlusion and complex movement directions, making it difficult to achieve efficient and accurate trajectory reconstruction and updating. Summary of the Invention
[0003] This application provides a bidirectional moving target trajectory recognition and analysis device and method based on video stream, which solves the technical problems of existing technologies in handling target occlusion and misjudgment of bidirectional motion direction in complex scenes, and the insufficient accuracy and reliability of trajectory recognition.
[0004] The first aspect of this application provides a bidirectional moving target trajectory recognition and analysis device based on video stream. The device includes: a target detection module, used to collect video stream data of a target area, perform video frame target detection on the video stream data, and obtain first target detection data, the first target detection data including target bounding box, species type, and confidence level; an occlusion type analysis module, used to obtain the target motion trajectory by analyzing the inter-frame correlation of the first target detection data, determine whether the target motion trajectory is occluded, perform occlusion type label analysis on the occluded video frames, and adaptively output a confidence search range step size for relocalization according to the occlusion type label; a motion trajectory restoration module, used to reconstruct the occluded target motion trajectory according to the confidence search range step size, and output the restored target motion trajectory as the target trajectory recognition result; and a target trajectory update module, used to update the target trajectory recognition result by setting a first virtual gate line and a second virtual gate line, and output the updated target trajectory recognition result.
[0005] A second aspect of this application provides a bidirectional moving target trajectory recognition and analysis method based on video streams. The method includes: acquiring video stream data of a target region; performing video frame target detection on the video stream data to obtain first target detection data, the first target detection data including target bounding boxes, species type, and confidence level; obtaining the target motion trajectory by analyzing the inter-frame correlation of the first target detection data; determining whether the target motion trajectory is occluded; performing occlusion type label analysis on the occluded video frames; adaptively outputting a confidence search range step size for relocalization according to the occlusion type label; reconstructing the occluded target motion trajectory according to the confidence search range step size; outputting the reconstructed target motion trajectory as the target trajectory recognition result; updating the target trajectory recognition result by setting a first virtual gating line and a second virtual gating line; and outputting the updated target trajectory recognition result.
[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application provides a bidirectional moving target trajectory recognition and analysis device and method based on video stream, relating to the field of image processing technology. It acquires target bounding boxes, species types, and confidence levels through video stream data acquisition and target detection; analyzes inter-frame correlations to obtain motion trajectories and determine occlusion; adaptively adjusts the confidence search range step size to recover the target motion trajectory; and determines the target's motion direction based on virtual gating lines, ensuring high-precision target counting and motion direction determination. This solves the technical problems of existing technologies in effectively handling target occlusion and misjudgment of bidirectional motion directions in complex scenes, resulting in insufficient accuracy and reliability of trajectory recognition. It achieves the technical effect of improving the accuracy and reliability of trajectory recognition in complex environments through occlusion analysis and virtual gating line determination. Attached Figure Description
[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 A schematic diagram of the structure of the bidirectional moving target trajectory recognition and analysis device based on video stream provided in the embodiments of this application; Figure 2 A schematic flowchart of the bidirectional moving target trajectory recognition and analysis method based on video stream provided in the embodiments of this application.
[0009] Figure labeling: Target detection module 11, Occlusion type analysis module 12, Motion trajectory restoration module 13, Target trajectory update module 14. Detailed Implementation
[0010] This application provides a bidirectional moving target trajectory recognition and analysis device and method based on video stream, which solves the technical problems of existing technologies in handling target occlusion and misjudgment of bidirectional motion direction in complex scenes, and the insufficient accuracy and reliability of trajectory recognition.
[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0012] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, apparatus, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.
[0013] Example 1 like Figure 1 As shown, this application provides a bidirectional moving target trajectory recognition and analysis device based on video stream, the device comprising: The target detection module 11 is used to collect video stream data of the target area, perform video frame target detection on the video stream data, and obtain first target detection data, which includes target bounding box, species type and confidence level.
[0014] Specifically, the target detection module 11 of this application is responsible for extracting targets from the video stream data of the target region and providing the core data required for subsequent analysis. Its main function is to detect, identify and extract key information of targets in video frames through deep learning algorithms, including the target's bounding box, species type and confidence level.
[0015] First, real-time video stream data is acquired from the camera equipment in the monitored area. The video stream is a continuous sequence of images containing all dynamic information within the monitored area, such as the movement trajectories of fish. Because fishway monitoring typically faces complex underwater environments, such as turbid water currents, strong light reflections, and obstructions from bubbles and other objects, the target detection system needs to be highly robust. Therefore, the video stream acquisition needs to meet certain frame rate and resolution requirements to clearly capture the movement details and features of the target.
[0016] Next, to extract useful information from the video stream, the video frame data is decoded and preprocessed, including noise removal, contrast adjustment, and necessary color correction to optimize image quality and enhance the distinction between the target and the background. Subsequently, deep learning algorithms are used to perform target detection on the preprocessed video frames. For example, the YOLOv8 model based on deep learning can be used. YOLOv8 is an efficient and accurate model in the field of target detection, capable of quickly and accurately identifying and locating targets in real-time video stream data. It processes each frame of the image through a convolutional neural network (CNN) to generate all target location and category information within that frame.
[0017] In its implementation, YOLOv8 generates a bounding box for each target. The bounding box marks the target's position within the video frame and is typically defined by four parameters: the top-left corner coordinates (x1, y1) and the bottom-right corner coordinates (x2, y2). These four coordinates determine the target's precise location. Each target's bounding box not only marks its position in the image but is also output along with its category and confidence score. The target category represents the species type to which the target belongs, such as carp or crucian carp. The target detection model performs classification based on a pre-trained dataset, distinguishing different species of fish and other related objects.
[0018] In addition, the YOLOv8 model provides a confidence score for each target. The confidence score is an evaluation of the reliability of the detection result by the target detection algorithm, representing the probability that the target was correctly detected, with a value ranging from 0 to 1. A higher confidence score indicates a greater likelihood that the target was accurately identified; conversely, a lower confidence score indicates uncertainty in the detection result. In practical applications, the confidence score is used to filter out more reliable detection results. Only when the confidence score exceeds a predetermined threshold is the detection result considered valid and proceeds to subsequent target tracking, motion analysis, and other processing steps.
[0019] The final output is the first target detection data, which includes bounding boxes, species types, and confidence scores for multiple targets. This data will provide basic information for subsequent target tracking, occlusion type analysis, and motion trajectory reconstruction modules, and is the key input data for the entire system operation.
[0020] The occlusion type analysis module 12 is used to obtain the target motion trajectory by analyzing the inter-frame correlation of the first target detection data, determine whether the target motion trajectory is occluded, perform occlusion type label analysis on the video frames with occlusion, and adaptively output the confidence search range step size for relocalization according to the occlusion type label.
[0021] Optionally, the main function of the occlusion type analysis module 12 in this application is to analyze the first target detection data output by the target detection module 11, identify the occlusion situation in the target motion trajectory, and perform detailed occlusion type analysis on the video frames with occlusion. Based on the occlusion type, the confidence search range step size used for target relocalization is adaptively adjusted, thereby providing support for subsequent target trajectory reconstruction.
[0022] First, the system receives initial target detection data, including target bounding boxes, species types, and confidence levels. Using this data, the occlusion type analysis module 12 analyzes the target detection results in each frame and performs continuous target tracking based on inter-frame target correlations. This process can be implemented using multi-target tracking algorithms, such as the ByteTrack algorithm. The ByteTrack algorithm incorporates both high-confidence and low-confidence detection results into the trajectory association process, uses a Kalman filter for target state prediction, and combines an intersection-over-union (IoU) matching strategy to associate the detection results with existing trajectories, thereby generating a continuous motion trajectory for each target. This algorithm can maintain trajectory continuity in complex scenes, such as target crowding and occlusion, providing a foundation for subsequent occlusion analysis.
[0023] After generating the target's motion trajectory, the module further determines whether occlusion exists within the trajectory. Occlusion refers to the phenomenon where a target is partially or completely obscured by other targets or background objects during its movement, which may lead to the loss of target detection or trajectory breakage. Occlusion determination is achieved by analyzing the continuity of the trajectory, changes in the target bounding box, and fluctuations in confidence level. For example, a sudden decrease in the area of the target bounding box or a significant drop in confidence level may indicate that the target has entered an occlusion state. Furthermore, the module analyzes changes in the shape of the target bounding box and its relative positional relationship with other targets to determine the specific type of occlusion.
[0024] Once occlusion is confirmed, the occlusion type analysis module 12 further analyzes the type of occlusion. The core of this analysis lies in determining the nature of the occlusion based on changes in the target's trajectory, including whether it is partial or complete occlusion. Partial occlusion typically occurs when parts of the target are obscured by objects or water flow, but the target is still visible in the image. Complete occlusion, on the other hand, means the target is completely undetectable in the current frame, possibly because it is completely blocked or outside the field of view. Occlusion type analysis can be based on changes in the shape and size of the target's bounding box and its relative position to other targets. For example, the type of occlusion can be determined by calculating the overlapping area and shape change rate of the bounding box. This process requires combining the target's motion model and prior scene knowledge to improve the accuracy of occlusion type determination. For instance, for rapid occlusion, such as a target being briefly occluded due to instantaneous changes in water flow or rapid swimming, information such as the target's speed and acceleration can be used to infer that the target is still on its original trajectory. For static occlusion, it can be identified that the target cannot be detected at a certain location for an extended period.
[0025] After completing the occlusion type analysis, the confidence search range step size for target relocalization is adaptively adjusted based on the occlusion type. The confidence search range step size refers to the confidence threshold range and search step size used by the system to re-detect the target after occlusion is removed. Adjusting this parameter is crucial for target relocalization, especially in complex scenes where the target may enter and exit occlusion states multiple times within a short period. The adaptive adjustment is based on factors including occlusion type, target movement speed, scene complexity, and target category. For example, for a completely occluded target, the system needs to expand the confidence search range and increase the search step size to quickly relocalize the target after occlusion is removed; while for partially occluded targets, a smaller confidence search range and shorter search step size can be maintained because some features of the target are still visible, making relocalization relatively easier.
[0026] By adaptively adjusting the confidence search range step size, the accuracy and efficiency of target relocation can be effectively improved, providing reliable data support for subsequent target trajectory reconstruction and motion direction determination.
[0027] Furthermore, when the occlusion type analysis module 12 obtains the target motion trajectory by analyzing the inter-frame correlation of the first target detection data, it is also used to perform the following steps: P21: Predict the first target detection data to obtain the predicted position of the video frame; P22: Analyze the target bounding box of the video frame in the first target detection data and perform association matching with the predicted position of the video frame, and perform trajectory association on the successfully matched target bounding boxes to obtain the target motion trajectory; wherein, successful association matching includes the overlapping pixel area ratio of the target bounding box of the video frame and the predicted position of the video frame being greater than a preset ratio.
[0028] It should be understood that the occlusion type analysis module 12 can further refine the specific process of trajectory generation during the generation of the target motion trajectory, so as to improve the system's ability to stably track the target motion in complex environments.
[0029] Specifically, the system first predicts the target's position in the video frame based on the first target detection data. This predicted position is calculated using the target's motion trajectory in the previous few frames. Specifically, the system can use common trajectory prediction algorithms such as the Kalman filter, combined with the target's velocity and direction, to estimate the target's position in the current frame. Since the target's motion is usually relatively smooth and continuous, the predicted position based on information from previous frames can generally reflect the target's actual position quite accurately, especially when the target is partially occluded or temporarily disappears. This method helps reduce interruptions or false positives in target detection.
[0030] After obtaining the predicted target location, the target bounding box in each frame of the first target detection data can be further analyzed to perform association matching with the predicted location of that frame. Specifically, this process determines whether the target matches the predicted location in the current frame by comparing the spatial overlap between the target bounding box and the predicted location in the video frame. The criterion for association matching is that the percentage of overlapping pixel area between the target bounding box and the predicted location in the video frame is greater than a preset threshold percentage. Here, the percentage of overlapping pixel area refers to the ratio between the area of the intersection region of the target bounding box and the predicted location in the video frame and the total area of the target bounding box. If this ratio exceeds the preset threshold, the target bounding box is considered to have successfully matched the predicted location, indicating that the prediction of the target location is accurate and the target's position in the current frame matches the previously predicted trajectory.
[0031] For example, if the preset threshold is 50%, the target bounding box is considered to have successfully matched the predicted location only when the overlapping pixel area between the target bounding box and the predicted location exceeds 50%. Successfully matched target bounding boxes are then incorporated into the target's motion trajectory as part of the target's continuous motion, further enhancing the trajectory's stability and accuracy.
[0032] If the percentage of overlapping pixels between the target bounding box and the predicted position in the video frame does not reach the preset threshold, the occlusion type analysis module 12 will consider that the target does not match the predicted position, which may indicate that the target has been occluded, has an error, or has external interference. The system will then enter the corresponding fault tolerance mechanism and try to repair the trajectory in other ways.
[0033] Through this association and matching process, the occlusion type analysis module can maintain accurate tracking of the target trajectory even when the target is partially occluded or moving rapidly. Even if the target is not fully detected in some frames due to occlusion or other reasons, the module can still recover the target's position through prediction and matching mechanisms and incorporate it into the continuous target motion trajectory. Finally, the successfully matched target bounding boxes will be associated with the previous trajectory data to form a continuous target motion trajectory.
[0034] This method is particularly suitable for complex underwater environments and fishway monitoring scenarios, where targets may temporarily disappear or be partially obscured due to environmental factors such as water flow or crowd occlusion. Traditional single-frame association methods may not be effective in handling such situations. However, by predicting and matching target positions, the occlusion type analysis module 12 can improve adaptability in dynamically changing environments and ensure the continuity and accuracy of target trajectories.
[0035] Furthermore, when performing occlusion type label analysis on video frames with occlusion, the occlusion type analysis module 12 is also used to perform the following steps: P23: Collect image feature vectors for video frames with occlusion, wherein when a preset number of consecutive video frames return as unsuccessful association matching, they are marked as video frames with occlusion; P24: Obtain occlusion type labels by identifying the image feature vectors through an occlusion type classifier, wherein the occlusion type classifier is obtained by training through known occlusion type samples and corresponding image feature vector samples, wherein the known occlusion type samples include at least water turbidity, reflective occlusion, floating bubbles, and crowding occlusion.
[0036] Specifically, in the occlusion type analysis module 12, the process of performing occlusion type label analysis on video frames with occlusion can be further refined to ensure continuous tracking of the target.
[0037] Specifically, when a predetermined number of consecutive video frames fail to match the target location (i.e., the overlap between the target's bounding box and the predicted location is less than a predetermined threshold), these video frames are marked as occluded. At this point, image feature vector acquisition is performed to analyze the occluded video frames and extract feature vectors that characterize the current occlusion situation. These feature vectors may include the shape, size, and positional variations of the target bounding box, as well as information such as contrast and texture with the surrounding environment. These features form the basis of occlusion type analysis and help the classifier distinguish between different occlusion types.
[0038] Subsequently, an occlusion type classifier is used to identify the collected image feature vectors, thereby obtaining occlusion type labels. The occlusion type classifier is a pre-trained model capable of identifying the specific type of occlusion based on the input image feature vector. The classifier is trained based on known occlusion type samples and corresponding image feature vector samples. These known occlusion type samples include at least common occlusion conditions such as water turbidity, reflective occlusion, floating bubbles, and crowding occlusion. Through extensive training with numerous samples, the occlusion type classifier learns the representation of different occlusion types in image feature vectors, thus accurately identifying occlusion types in practical applications.
[0039] Known occlusion types include several typical occlusion scenarios, such as turbid water, reflective occlusion, floating bubbles, and crowding occlusion. Turbid water occlusion typically occurs in turbid water with high particulate matter content, blurring target boundaries and making accurate identification difficult. Reflective occlusion occurs on the water surface or other reflective surfaces; reflected light interferes with target detection, especially under strong lighting conditions. Floating bubbles occlusion refers to floating bubbles in the water potentially obscuring the target and affecting its visibility. Crowding occlusion occurs when multiple targets or objects obscure each other, commonly seen in environments with dense fish populations, making it impossible to accurately distinguish each target.
[0040] By analyzing these image feature vectors, the occlusion type classifier can accurately determine the type of occlusion in video frames. The classifier compares extracted features, such as color distribution, texture patterns, and edge information, with occlusion type samples obtained during training to identify the specific occlusion type. For example, if the image features exhibit a blurry or grainy pattern, the classifier might classify it as turbid water occlusion; if the image has strong reflected light, it might be classified as reflective occlusion.
[0041] This method, which uses image feature vector analysis and classification, can effectively identify and classify different types of occlusion, providing the system with more accurate target tracking information in complex environments. Once a specific occlusion type is identified, the system can take corresponding processing measures based on the different occlusion conditions. For example, in turbid water conditions, the search range can be increased to recover the target; while in cases of reflective occlusion, specific anti-reflection processing can be performed on the image.
[0042] Furthermore, when the occlusion type analysis module 12 adaptively outputs the confidence search range step size for relocation according to the occlusion type label, it is also used to perform the following steps: P25: Extract the occlusion time features, occlusion area features, and occlusion stability features of the occlusion type label; P26: Construct an occlusion weight vector based on the occlusion time features, occlusion area features, and occlusion stability features; P27: Input the occlusion weight vector into the adaptive search mapping model for analysis to determine the confidence threshold adjustment step size and the search range radius step size; P28: Output the confidence search range step size according to the confidence threshold adjustment step size and the search range radius step size.
[0043] Optionally, the process of the occlusion type analysis module 12 adaptively outputting the confidence search range step size for target relocation based on the occlusion type label can be further refined to ensure that the target can be relocated efficiently and accurately in complex occlusion environments.
[0044] After identifying the occlusion type label, the system first extracts features related to the occlusion type label. These features include occlusion time features, occlusion area features, and occlusion stability features. Occlusion time features refer to the duration for which the target is occluded in the video stream. For shorter occlusion periods, the system can assume the target will recover quickly; however, for longer occlusion periods, a larger search is required to recover the target. The system determines the likelihood of target recovery by identifying the duration of the occlusion.
[0045] Occlusion area features are used to distinguish between complete occlusion and partial occlusion. Complete occlusion means that the target is completely invisible in the current frame, while partial occlusion means that the target is occluded in a part of the image, but part of the target's outline can still be seen in other areas of the image. Complete occlusion usually requires a larger search area because the target may disappear completely, while the search area for partial occlusion can be appropriately reduced.
[0046] Occlusion stability refers to the persistence and instability of occlusion, specifically measured by flicker rate. Flicker rate refers to how frequently the bounding box of a target disappears and reappears in the image after the target is occluded. A high flicker rate usually indicates that the occlusion is caused by changes in the environment, such as water flow or other dynamic factors, and the system may need to make more dynamic adjustments to cope with these changes.
[0047] Subsequently, an occlusion weight vector is constructed based on the extracted occlusion time features, occlusion area features, and occlusion stability features. This vector integrates these features, reflecting the severity and type of occlusion. For example, complete occlusion may have a higher weight, indicating that a larger search area is needed; while partial occlusion has a lower weight, indicating that the search area can be appropriately reduced. Similarly, longer periods of occlusion will result in higher weights, while short periods of occlusion will have lower weights; high flicker rate occlusion will also increase the weight, indicating that the impact of occlusion is greater and a more flexible search strategy is needed.
[0048] Next, the constructed occlusion weight vector is input into an adaptive search mapping model for analysis. This adaptive search mapping model is a pre-trained model that determines the appropriate confidence threshold adjustment step size and search range radius step size based on the input occlusion weight vector. The model can be trained on a large amount of sample data under different occlusion conditions, thus learning the impact of different combinations of occlusion features on target relocalization requirements. By analyzing the occlusion weight vector, the model can output the most suitable confidence threshold adjustment step size and search range radius step size for the current occlusion situation.
[0049] Finally, the step size and search range radius step size are adjusted according to the confidence threshold output by the adaptive search mapping model to generate the final confidence search range step size. This confidence search range step size will be used in the subsequent target relocalization process to ensure that the target can be quickly and accurately redetected after the occlusion is removed. The confidence search range step size will be dynamically adjusted according to the type, time, area, and stability characteristics of the occlusion to adapt to different occlusion scenarios.
[0050] Furthermore, in the occlusion type analysis module 12, the adaptive search mapping model is trained on the occlusion relocation accuracy by using historical video stream data samples and the solution space of the confidence threshold adjustment step size and the search range radius step size under different occlusion type labels; until the optimal solutions of the confidence threshold adjustment step size and the search range radius step size are obtained for different occlusion type labels while meeting the target accuracy, thus obtaining the adaptive search mapping model. The confidence threshold adjustment step size includes a positive confidence threshold adjustment step size and a negative confidence threshold adjustment step size.
[0051] It should be understood that in the occlusion type analysis module 12, the training and optimization of the adaptive search mapping model is the core part of achieving efficient target relocalization. To ensure the system can accurately recover the target under different occlusion conditions, the model needs continuous feedback training using historical video stream data and occlusion type labels to optimize the search strategy, especially the confidence threshold adjustment step size and the search range radius step size. This process is completed through feedback training, ensuring that the adaptive search mapping model can adapt to various occlusion scenarios and provide the optimal relocalization solution.
[0052] First, the adaptive search mapping model is trained based on historical video stream data samples. These historical samples contain detailed data on how the target is recovered under different occlusion types. Each video stream sample contains different types of occlusion labels, such as turbid water, reflective occlusion, floating bubbles, and crowded occlusion, which provide the model with various target occlusion scenarios in real-world environments.
[0053] During model training, the solution space of the confidence threshold adjustment step size and the solution space of the search range radius adjustment step size are key components that need to be adjusted. The solution space of the confidence threshold adjustment step size includes two types: positive and negative. The positive confidence threshold adjustment step size is mainly used to gradually increase the confidence threshold for accepting a target when the confidence of target detection is low, allowing the system to consider a wider range of target detection possibilities. Conversely, the negative confidence threshold adjustment step size is used to reduce the confidence threshold required for target determination when the confidence of a target is high, preventing premature rejection of reliable targets.
[0054] Specifically, the model undergoes simulated training using historical samples and corresponding occlusion type labels, continuously adjusting the confidence threshold adjustment step size and the search range radius step size, and providing feedback based on the detection results after each adjustment. Through this feedback mechanism, the model gradually learns during training which adjustment step size best ensures high accuracy for the target under different types of occlusion.
[0055] During training, the system will adjust the search strategy based on the occlusion type label of each sample. For example, in turbid water conditions, the boundary of the target may be blurred, and the step size of the positive confidence threshold adjustment may need to be larger to avoid the system misjudging the target as disappeared. In the case of crowded occlusion, the step size of the negative confidence threshold adjustment may need to be smaller, because the presence of multiple targets may cause one target to be misjudged as disappeared, requiring a higher confidence level to maintain target tracking.
[0056] Each adjusted model receives feedback by comparing its performance to the preset target accuracy, i.e., the precision of target relocalization after occlusion. If the system achieves high target relocalization accuracy under specific occlusion types, the adjustment in this step is confirmed as an optimization step size and applied to the next round of training. This feedback training process continues until the model finds the optimal confidence threshold adjustment step size and search range radius step size under different occlusion type labels. Optimizing the confidence threshold adjustment step size ensures that the system can effectively identify and recover targets under various occlusion conditions, while adjusting the search range radius step size allows the system to determine the size of the search area based on the target's specific location and occlusion situation, further improving the efficiency and accuracy of target relocalization.
[0057] Through these training and feedback processes, the adaptive search mapping model will eventually obtain an optimal solution, namely the optimal solution for the confidence threshold adjustment step size and the optimal solution for the search range radius step size under different occlusion types. These optimal solutions ensure that the system can perform target relocalization with the most appropriate search range and confidence threshold when facing different types of occlusion, thereby ensuring the stability and high accuracy of the target tracking process.
[0058] The motion trajectory restoration module 13 is used to restore the motion trajectory of the occluded target according to the confidence search range step size, and output the restored target motion trajectory as the target trajectory recognition result.
[0059] Specifically, the core task of the motion trajectory restoration module 13 in this application is to use the adjusted confidence search range step size to reconstruct the motion trajectory of the target through prediction and inference, so as to ensure that the trajectory of the target can continuously and accurately reflect its motion state.
[0060] First, after the occlusion type analysis module 12 completes the determination of target occlusion and optimizes the confidence search range step size through the adaptive search mapping model, the motion trajectory reconstruction module 13 begins to reconstruct the target's motion trajectory. The target's motion in the video stream is usually continuous. Even if the target cannot be detected in some frames due to occlusion, the system can still recover the target's trajectory through historical data and prediction methods.
[0061] The reconstruction of the motion trajectory primarily relies on the confidence search range step size. This step size controls the size of the search area used by the system for target relocalization during target occlusion. When the target is completely or partially occluded, the system needs to dynamically adjust the search range based on the occlusion characteristics to ensure the target's position in the current frame is found as accurately as possible. By adjusting the confidence search range step size, the system searches for the target within the defined area and determines whether the position matches the target's historical trajectory.
[0062] During the restoration process, the system combines information such as the target's position, velocity, and direction of motion in historical frames to infer the target's possible position during the occlusion period. This process can be based on physical motion models, such as uniform linear motion or accelerated motion, to make predictions. By combining the target's previous motion trajectory, the system estimates the target's possible position in the current frame and matches it with candidate positions within the search range. Specifically, when the system cannot find the exact position of the target in the video frame, the motion trajectory restoration module 30 will use the target's velocity, direction of motion, and historical position to calculate the most likely position, and assign a confidence score to this position. Based on the confidence score, the system will then lock the most probable target position, recover the trajectory data lost during the occlusion period, and ensure the continuity of the target's motion trajectory.
[0063] Next, during the reconstruction process, the motion trajectory reconstruction module 13 not only relies on the confidence search range step size but also needs to consider the target's motion inertia. For example, if the target's velocity was high in the previous frame, the system will predict the target's possible position in the next frame based on the physical model and expand or shrink the search range accordingly. Furthermore, the system will adjust the trajectory reconstruction strategy for different types of occlusion. For completely occluded targets, a wider search range and a longer historical trajectory dependency are required; while for partially occluded targets, the system relies more on the target detection information of the current frame to recover the target's position more quickly.
[0064] After reconstructing the target's motion trajectory, the reconstructed trajectory is output and used as the target trajectory recognition result. This output not only provides a foundation for continuous target tracking but also provides necessary trajectory data for subsequent data analysis, counting, and direction determination.
[0065] The target trajectory update module 14 is used to update the target trajectory recognition result by setting a first virtual gate line and a second virtual gate line, and output the updated target trajectory recognition result.
[0066] Furthermore, when the target trajectory update module 14 updates the target trajectory recognition result by setting the first virtual gating line and the second virtual gating line, it is also used to perform the following steps: P41: Obtain the first and second crossing timing records of the first and second virtual gating lines; P42: Construct a finite state machine model, use the finite state machine model to identify the first and second crossing timing records, and mark the target motion direction of the target trajectory identification result; P43: Update the target trajectory identification result according to the target motion direction, and output the updated target trajectory identification result.
[0067] It should be understood that the main function of the target trajectory update module 14 in this application is to ensure the accuracy of the target's movement direction and counting results by updating the target trajectory recognition results. The operation of this module not only depends on the target's movement trajectory, but also incorporates the crossing timing of the virtual gating line to dynamically adjust the target's movement direction and counting.
[0068] First, the target trajectory update module 40 begins collecting and recording the timing data of the target crossing the line using the first and second virtual gating lines. The first and second virtual gating lines are typically placed at key locations in the target's trajectory, serving as reference lines for determining the target's direction of movement. Whenever the target crosses these two gating lines, the target trajectory update module 14 records the timing data of the crossing. This timing data includes the timestamp of the target crossing the line, the target's direction of movement, and whether the target successfully crossed the virtual gating lines—that is, the first and second crossing timing data—which serve as crucial evidence for subsequent direction of movement determination.
[0069] Subsequently, a finite state machine (FSM) model is constructed and used to identify the timing records of the first and second virtual gate crossings, thereby marking the target's motion direction based on the target trajectory identification results. The finite state machine model is a logic model based on state transitions, capable of accurately determining the target's motion direction according to the order and timing of the target's crossings of virtual gate lines. For example, if the target crosses the first virtual gate line first and then the second, it is determined that the target is moving upwards; conversely, if the target crosses the second virtual gate line first and then the first, it is determined that the target is moving downwards. Furthermore, the finite state machine model can also identify special cases, such as the target crossing the gate line multiple times within a short period (U-turn behavior), and handle these cases specially according to preset rules to avoid misjudgments and duplicate counting.
[0070] Finally, the target trajectory recognition results are updated based on the identified target movement direction. By using the crossing time sequence of the virtual gating lines, the system can accurately determine the movement direction of each target and update the target count accordingly. For example, for a target moving upwards, the count is incremented in the upward counter; for a target moving downwards, the count is incremented in the downward counter. Simultaneously, the updated count results and target movement direction information are integrated into the target trajectory recognition results to form complete output data. This data includes not only the target's trajectory but also detailed information such as the target's movement direction, count information, and timestamps of crossing the virtual gating lines, providing comprehensive data support for subsequent analysis and applications.
[0071] Furthermore, the target trajectory update module 14 is also used to perform the following steps: P42-1: The finite state machine model includes a direction determination rule and a set of states for direction determination. The set of states includes a standby state, a first crossing state, a second crossing state, and a completion state. P42-2: The finite state machine model identifies the first crossing time sequence data and the second crossing time sequence data according to the direction determination rule to obtain the target motion direction of the target trajectory identification result.
[0072] Specifically, the target trajectory update module 14 can be further refined to determine the direction of the target's movement using virtual gating lines and finite state machine models, and update the target's trajectory using this information.
[0073] First, a finite state machine model is constructed, which includes direction determination rules and a set of states used for direction determination. The state set includes a standby state, a first crossing state, a second crossing state, and a completion state. The standby state indicates that the system is in its initial state and has not yet detected the target crossing any virtual gate lines; the first crossing state indicates that the target has crossed the first virtual gate line; the second crossing state indicates that the target has crossed the second virtual gate line; and the completion state indicates that the target has completely crossed both virtual gate lines, completing one direction determination.
[0074] Next, the finite state machine model identifies the timing records of the first and second virtual gate crossings according to the direction determination rules. Specifically, when the centroid of the target trajectory first crosses the first virtual gate line, the model records the timestamp of this event and transitions from the standby state to the first crossing state. Subsequently, when the centroid of the target trajectory crosses the second virtual gate line, the model records the timestamp of this event and transitions from the first crossing state to the second crossing state. At this point, the model determines the target's direction of motion based on the order and timing of the target crossing the two virtual gate lines. For example, if the target crosses the first virtual gate line first and then the second virtual gate line, it is determined that the target is moving upwards; conversely, if the target crosses the second virtual gate line first and then the first virtual gate line, it is determined that the target is moving downwards.
[0075] Furthermore, the finite state machine model considers special cases, such as a target crossing a gate line multiple times within a short period (U-turn behavior). In such cases, the model handles these special situations according to pre-defined rules, such as canceling false counts, to ensure the accuracy of direction determination. After completing the direction determination, the model transitions to the completed state and records the determined target motion direction as part of the target trajectory recognition result. This process not only improves the accuracy of target motion direction determination but also provides reliable data support for subsequent target trajectory updates and counting.
[0076] Furthermore, when updating the target trajectory recognition result, the target trajectory update module 14 is also used to perform the following steps: P43-1: Set a first virtual gating line and a second virtual gating line, the first virtual gating line and the second virtual gating line include a preset spacing, and configure a preset time window according to the preset spacing; P43-2: Analyze the first crossing time sequence record data and the second crossing time sequence record data to determine whether the duration between the sequential triggering of the first crossing state and the second crossing state is within the preset time window; P43-3: If it is within the preset time window, mark the identified target movement direction as valid; P43-4: If it is not within the preset time window, mark the identified target movement direction as invalid.
[0077] Optionally, the target trajectory update module 14 can further update the counting of the target trajectory recognition results by setting the spacing and time window of the virtual gate lines and analyzing the crossing time sequence data to ensure the accuracy of the target's movement direction and counting.
[0078] First, a first virtual gating line and a second virtual gating line are set, and a preset spacing is configured for these two gating lines. This spacing is configured according to the actual application scenario and the target tracking requirements. Based on the preset spacing, a preset time window is further configured. The preset time window is set based on the expected crossing time of the target between the two virtual gating lines and is used to determine whether the target's crossing behavior is logical. For example, if the preset spacing is 1 meter and the target's average speed is 1 meter per second, the preset time window can be set between 0.5 seconds and 1.5 seconds to allow for some speed fluctuation.
[0079] Next, the timing data of the first and second virtual gate crossings are analyzed to determine whether the time between the target's sequential triggering of the first and second virtual gate crossing states falls within a preset time window. Specifically, the time difference between the target's crossing from the first virtual gate line to the second virtual gate line is calculated and compared with the preset time window. If the time difference is within the preset time window, it indicates that the target's crossing behavior is as expected, and the identified target movement direction is marked as valid. For example, if the time difference between the target's crossing from the first and second virtual gate lines is 1 second, and the preset time window is 0.5 to 1.5 seconds, then the target's movement direction is marked as valid.
[0080] Conversely, if the time difference is not within the preset time window, it indicates that the target's crossing behavior may be interfered with or illogical, and the identified target movement direction will be marked as invalid. For example, if the time difference is 2 seconds, which exceeds the preset time window, the target's movement direction will be marked as invalid to avoid false counts caused by abnormal crossing behavior.
[0081] Through the above steps, the target trajectory update module 14 can accurately update the target trajectory recognition results. This process not only improves the accuracy of target motion direction determination, but also effectively filters abnormal crossing behavior through the preset time window mechanism, further ensuring the reliability of the count and providing effective support for data analysis and application.
[0082] In summary, the embodiments of this application have at least the following technical effects: This application effectively recovers the motion trajectory of occluded targets through occlusion type analysis and motion trajectory reconstruction, ensuring the continuity and integrity of the trajectory and significantly solving the problem of trajectory interruption in existing technologies. Utilizing virtual gating lines and finite state machine models, it accurately determines the bidirectional motion direction of the target, avoiding misjudgment and repeated counting, thereby improving the reliability of direction determination. Combined with adaptive adjustment of the confidence search range step size, it optimizes the target relocation process, improving the efficiency and accuracy of relocation, especially performing well in complex occlusion environments. The modular design allows the system to flexibly respond to different occlusion types and motion patterns in various scenarios, significantly improving the system's adaptability and robustness. By adaptively adjusting the confidence threshold and search range, it effectively reduces the false detection and false negative rates caused by occlusion or environmental interference, improving the overall accuracy of trajectory recognition.
[0083] The technology achieves the goal of improving the accuracy and reliability of trajectory recognition in complex environments through occlusion analysis and virtual gating line determination.
[0084] Example 2 Based on the same inventive concept as the bidirectional moving target trajectory recognition and analysis device based on video streams in the foregoing embodiments, such as Figure 2 As shown, this application provides a method for bidirectional moving target trajectory recognition and analysis based on video streams. The apparatus and method embodiments in this application are based on the same inventive concept. The method includes: Video stream data of the target area is collected, and video frame target detection is performed on the video stream data to obtain first target detection data, which includes target bounding box, species type, and confidence score. The target motion trajectory is obtained by analyzing the inter-frame correlation of the first target detection data, and it is determined whether the target motion trajectory is occluded. For video frames with occlusion, occlusion type label analysis is performed, and the confidence search range step size for relocalization is adaptively output according to the occlusion type label. The target motion trajectory with occlusion is reconstructed according to the confidence search range step size, and the reconstructed target motion trajectory is output as the target trajectory recognition result. The target trajectory recognition result is updated by setting a first virtual gate line and a second virtual gate line, and the updated target trajectory recognition result is output.
[0085] Furthermore, the target trajectory recognition result is updated by setting a first virtual gating line and a second virtual gating line, the method including: Acquire the first and second crossing timing records of the first and second virtual gating lines; construct a finite state machine model, use the finite state machine model to identify the first and second crossing timing records, and mark the target motion direction of the target trajectory identification result; update the target trajectory identification result according to the target motion direction, and output the updated target trajectory identification result.
[0086] Furthermore, the finite state machine model includes direction determination rules and a set of states for direction determination, the set of states including a standby state, a first crossing state, a second crossing state, and a completed state; the finite state machine model identifies the first crossing time sequence record data and the second crossing time sequence record data according to the direction determination rules to obtain the target motion direction of the target trajectory identification result.
[0087] Furthermore, a first virtual gating line and a second virtual gating line are set, the first virtual gating line and the second virtual gating line having a preset spacing, and a preset time window is configured according to the preset spacing; the timing data of the first and second crossing are analyzed to determine whether the duration between the sequential triggering of the first crossing state and the second crossing state is within the preset time window; if it is within the preset time window, the identified target movement direction is marked as valid; if it is not within the preset time window, the identified target movement direction is marked as invalid.
[0088] Furthermore, the target motion trajectory is obtained by analyzing the inter-frame correlation of the first target detection data, the method including: The first target detection data is predicted to obtain the predicted position of the video frame; the target bounding box of the video frame in the first target detection data is analyzed and associated with the predicted position of the video frame, and the target bounding box with successful association is associated with the trajectory to obtain the target motion trajectory; wherein, successful association includes the overlapping pixel area ratio of the target bounding box of the video frame and the predicted position of the video frame being greater than a preset ratio.
[0089] Furthermore, occlusion type label analysis is performed on video frames with occlusion, including the following methods: Image feature vectors are collected from video frames with occlusion. When a preset number of consecutive video frames return as unsuccessful association matching, they are marked as occluded video frames. The occlusion type label is obtained by identifying the image feature vectors through an occlusion type classifier. The occlusion type classifier is trained by known occlusion type samples and corresponding image feature vector samples. The known occlusion type samples include at least water turbidity, reflective occlusion, floating bubbles, and crowding occlusion.
[0090] Furthermore, the confidence search range step size for relocalization is adaptively output according to the occlusion type label, and the method includes: Extract the occlusion time features, occlusion area features, and occlusion stability features of the occlusion type label; construct an occlusion weight vector based on the occlusion time features, occlusion area features, and occlusion stability features; input the occlusion weight vector into an adaptive search mapping model for analysis to determine the confidence threshold adjustment step size and the search range radius step size; output the confidence search range step size according to the confidence threshold adjustment step size and the search range radius step size.
[0091] Furthermore, the adaptive search mapping model is trained by using historical video stream data samples and adjusting the step size solution space and the search range radius step size solution space based on confidence threshold under different occlusion type labels to improve occlusion relocation accuracy; until the optimal solutions for confidence threshold adjustment step size and search range radius step size are obtained for different occlusion type labels while meeting the target accuracy, thus obtaining the adaptive search mapping model.
[0092] Furthermore, the confidence threshold adjustment step size includes a positive confidence threshold adjustment step size and a negative confidence threshold adjustment step size.
[0093] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0094] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0095] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.
Claims
1. A bidirectional moving target trajectory recognition and analysis device based on video stream, characterized in that, The device includes: The target detection module is used to collect video stream data of the target area, perform video frame target detection on the video stream data, and obtain first target detection data, which includes target bounding box, species type and confidence level; The occlusion type analysis module is used to obtain the target motion trajectory by analyzing the inter-frame correlation of the first target detection data, determine whether the target motion trajectory is occluded, perform occlusion type label analysis on the video frames with occlusion, and adaptively output the confidence search range step size for relocalization according to the occlusion type label. The motion trajectory restoration module is used to reconstruct the motion trajectory of the occluded target according to the confidence search range step size, and output the restored target motion trajectory as the target trajectory recognition result. The target trajectory update module is used to update the target trajectory recognition result by setting a first virtual gate line and a second virtual gate line, and output the updated target trajectory recognition result.
2. The bidirectional moving target trajectory recognition and analysis device based on video stream according to claim 1, characterized in that, When the target trajectory update module updates the target trajectory recognition result by setting a first virtual gating line and a second virtual gating line, it is also used for: Acquire the first and second crossing timing record data of the first and second virtual gating lines; A finite state machine model is constructed, and the first and second crossing time sequence records are identified using the finite state machine model. The target motion direction of the target trajectory identification result is marked. The target trajectory recognition result is updated by counting according to the target's direction of motion, and the updated target trajectory recognition result is output.
3. The bidirectional moving target trajectory recognition and analysis device based on video stream according to claim 2, characterized in that, In the target trajectory update module, the finite state machine model includes direction determination rules and a set of states for direction determination. The set of states includes a standby state, a first crossing state, a second crossing state, and a completion state. The finite state machine model identifies the first and second crossing time sequence records based on the direction determination rules to obtain the target motion direction of the target trajectory identification result.
4. The bidirectional moving target trajectory recognition and analysis device based on video stream according to claim 3, characterized in that, When updating the target trajectory recognition result, the target trajectory update module is further configured to: Set a first virtual gate control line and a second virtual gate control line, the first virtual gate control line and the second virtual gate control line include a preset interval, and configure a preset time window according to the preset interval; Analyze the first and second cross-line timing records to determine whether the duration between the sequential triggering of the first and second cross-line states falls within the preset time window. If it falls within the preset time window, the identified target motion direction will be marked as valid; If the target motion direction is not within the preset time window, the identified target motion direction will be marked as invalid.
5. The bidirectional moving target trajectory recognition and analysis device based on video stream according to claim 1, characterized in that, When the occlusion type analysis module obtains the target motion trajectory by analyzing the inter-frame correlation of the first target detection data, it is also used for: Predict the first target detection data to obtain the predicted position of the video frame; The target bounding boxes of video frames in the first target detection data are analyzed and matched with the predicted positions of the video frames. The target bounding boxes that are successfully matched are then used to obtain the target motion trajectory through trajectory association. Successful association matching includes situations where the percentage of overlapping pixel area between the target bounding box of the video frame and the predicted position of the video frame is greater than a preset percentage.
6. The bidirectional moving target trajectory recognition and analysis device based on video stream according to claim 5, characterized in that, When performing occlusion type label analysis on video frames with occlusion, the occlusion type analysis module is also used for: Image feature vectors are collected from video frames that are occluded. When a preset number of consecutive video frames return as unsuccessful association matching, they are marked as occluded video frames. The occlusion type label is obtained by identifying the image feature vector through an occlusion type classifier. The occlusion type classifier is trained by known occlusion type samples and corresponding image feature vector samples. The known occlusion type samples include at least water turbidity, reflective occlusion, floating bubbles, and crowding occlusion.
7. The bidirectional moving target trajectory recognition and analysis device based on video stream according to claim 6, characterized in that, When the occlusion type analysis module adaptively outputs the confidence search range step size for relocation according to the occlusion type label, it is also used for: Extract the occlusion time features, occlusion area features, and occlusion stability features of the occlusion type label; An occlusion weight vector is constructed based on the occlusion time characteristics, occlusion area characteristics, and occlusion stability characteristics. The occlusion weight vector is input into the adaptive search mapping model for analysis to determine the confidence threshold adjustment step size and the search range radius step size. Adjust the step size and search range radius step size according to the confidence threshold, and output the confidence search range step size.
8. The bidirectional moving target trajectory recognition and analysis device based on video stream according to claim 7, characterized in that, In the occlusion type analysis module, the adaptive search mapping model is trained to improve the accuracy of occlusion relocation by adjusting the step size solution space and the search range radius step size solution space based on the confidence threshold under different occlusion type labels using historical video stream data samples. The adaptive search mapping model is obtained by obtaining optimal solutions for the confidence threshold adjustment step size and the search range radius step size for different occlusion types of labels while satisfying the target accuracy.
9. The bidirectional moving target trajectory recognition and analysis device based on video stream according to claim 7, characterized in that, In the occlusion type analysis module, the confidence threshold adjustment step size includes a positive confidence threshold adjustment step size and a negative confidence threshold adjustment step size.
10. A bidirectional moving target trajectory recognition and analysis method based on video stream, characterized in that, The method includes: Collect video stream data of the target area, perform video frame target detection on the video stream data, and obtain first target detection data, which includes target bounding box, species type and confidence level; The target motion trajectory is obtained by analyzing the inter-frame correlation of the first target detection data. It is then determined whether the target motion trajectory is occluded. For video frames with occlusion, occlusion type label analysis is performed, and the confidence search range step size for relocalization is adaptively output according to the occlusion type label. The target motion trajectory with occlusion is reconstructed according to the confidence search range step size, and the reconstructed target motion trajectory is output as the target trajectory recognition result. The target trajectory recognition result is updated by setting a first virtual gating line and a second virtual gating line, and the updated target trajectory recognition result is output.