Open environment dynamic target tracking and recognition algorithm for short video platform

By combining deep learning and traditional computer vision technology, an efficient feature extraction and fusion network is built, and the accuracy and real-time problems of object detection and tracking in an open environment are solved, and an object detection system with high accuracy and strong generalization capabilities is achieved.

CN119942409APending Publication Date: 2025-05-06SPACE VISION (CHONGQING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510068141.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In an open environment, especially short video platforms, it is difficult to achieve high accuracy and real-time object detection and tracking under complex backgrounds, multi-variable target patterns and changes in lighting conditions.

Method used

The combination of deep learning technology and traditional computer vision technology is adopted to build an efficient feature extraction and fusion network, generate accurate candidate boxes, use the Hungarian graph matching algorithm to perform matching and classification prediction, and automatically identify unknown categories through unsupervised and semi-supervised learning to achieve continuous optimization of the model.

Benefits of technology

It significantly improves the accuracy and real-timeness of object detection in complex environments, improves the mean average mAP and F1 scores, and enhances the generalization ability and scalability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
Patent Text Reader

Abstract

The invention relates to an open environment dynamic target tracking and recognition algorithm for a short video platform, relates to the technical field of video processing, and aims to improve the accuracy and real-time performance of target detection and tracking in the short video platform. The method comprises the following steps: fusing visual features and position coding information extracted by a deep learning convolutional neural network to carry out complex background processing; a method for generating a random candidate detection frame based on Gaussian distribution is introduced, and the precision and the real-time performance of target tracking are optimized in combination with motion information and a context sensing strategy; the influence of illumination change on feature extraction is reduced by adopting color space conversion, and the stable performance under the illumination condition change is ensured through an efficient feature extraction mechanism; and automatic identification of novel target categories is realized by using unsupervised and semi-supervised learning methods. According to the method, the average precision mean value mAP and the F1 score are obviously superior to those in the prior art, the method is particularly excellent in the aspects of complex background processing, illumination condition changes and the like, and the efficiency and the accuracy of target detection and tracking are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing technology, and in particular to an open environment dynamic target tracking and recognition algorithm for a short video platform. Background Art

[0002] The widespread popularity of short video applications has led to a significant increase in users' demand for diversified video content. Although existing target detection and tracking technologies perform well in closed environments, they often perform poorly in open environments, especially in challenging scenarios such as short videos. The complex backgrounds in short videos, the changing shapes of targets, and the constantly changing lighting conditions are all factors that make it difficult for tracking and recognition accuracy and real-time performance to meet high standards. These challenges highlight the need for more advanced algorithms to better adapt to dynamic and unpredictable environments and meet expectations for high-quality video content; Traditional dynamic target tracking and recognition methods mainly include methods based on feature point matching, methods based on optical flow, and methods based on deep learning. The following are several typical existing technologies and their shortcomings: Complex background processing: When dealing with complex backgrounds, the Mean-Shift algorithm is easily affected by similar color areas in the background, resulting in tracking drift. The CamShift algorithm relies on the color histogram, so when there are similar colors to the target in the background, tracking is likely to fail. The particle filter algorithm may produce a large number of false positives in complex backgrounds due to background noise. Although deep learning methods have advantages in feature extraction, they may still be disturbed in extremely complex backgrounds; Targets with variable shapes: The Mean-Shift algorithm responds slowly and is difficult to adapt to rapid changes in target shape and size. The CamShift algorithm has poor adaptability to rapid changes in target shape and needs to reinitialize tracking. For rapid and irregular changes in targets, the particle filter algorithm requires a large number of particles to maintain tracking, which has high computational costs. Deep learning models require a large amount of data to train, and may require continuous online learning to adapt to the variable shapes of targets. Changes in lighting conditions: Light changes can affect color distribution. The Mean-Shift algorithm is more sensitive to lighting changes, resulting in unstable tracking. The CamShift algorithm's performance degrades under lighting changes. Light changes can affect feature extraction, and the particle filter algorithm may need to be reinitialized in this case. Deep learning models are somewhat robust to lighting changes, but extreme lighting changes may still affect model performance. Compared with the existing technology, this invention comprehensively utilizes the advantages of deep learning technology and traditional computer vision technology, and effectively responds to the challenges of target detection in open environments by building an efficient feature extraction and fusion network, an accurate candidate box generation and detection mechanism, a stable matching and classification prediction network, and a flexible unknown type recognition and model optimization strategy. In particular, it shows higher accuracy and robustness in complex background processing, variable target morphology, and changes in lighting conditions. Summary of the invention

[0003] This paper proposes a dynamic target tracking and recognition algorithm in an open environment for short video platforms, aiming to improve the accuracy and real-time performance of target detection and tracking in dynamic scenes. The algorithm combines the advantages of deep learning technology and traditional computer vision technology to meet the challenges of target detection in open environments; To achieve the above object, the present invention provides a dynamic target tracking and recognition algorithm in an open environment, comprising the following steps: S1 uses a deep learning convolutional neural network architecture to effectively extract rich visual features from video frames. These features cover multiple dimensions such as color information, delicate texture details, and the shape and outline of objects. Secondly, position information is added to each pixel to ensure that the model can understand the positional relationship of the features in the image. The feature and position information are fused to generate a comprehensive feature embedding. This embedding contains rich semantic information, laying the foundation for subsequent target detection and tracking. S2 generates multiple random candidate detection frames based on Gaussian distribution and statistical models. These candidate frames cover areas that may contain targets and provide candidate regions for the generation of feature detection frames. Using the feature embedding and candidate frames generated by the encoder, more accurate feature detection frames are generated through sampling and feature mapping. These detection frames are used for preliminary positioning of targets; S3 uses the Hungarian graph matching algorithm to evaluate the matching degree between feature detection boxes and known detection categories and screen out candidate targets with high matching degree. Based on the successfully matched feature detection boxes, a classifier (including fully connected layers or convolutional layers) is used to predict the category, location and confidence of the target; S4 automatically identifies targets of unknown categories by analyzing background frames and feature clustering, combining unsupervised learning and semi-supervised learning methods. When the model encounters new categories, it continuously updates model parameters through incremental learning mechanisms, including online learning and continuous learning, so that new data can adapt to the model and continuously enhance the model's adaptability and generalization capabilities.

[0004] Preferably, in said S1, it also includes: Select a pre-trained deep learning model as the backbone network, such as ResNet50, VGG16, etc. These models have been pre-trained on the large dataset ImageNet and can provide rich initial weight parameters, which not only speeds up the subsequent training process, but also provides a high performance starting point for the model. Adjust the input video frame to a resolution suitable for the backbone network input size (1080p) to reduce the amount of calculation and increase the processing speed. Use bilinear interpolation or nearest neighbor interpolation technology to flexibly adjust the image size. In order to improve the robustness of feature extraction, the RGB format can be converted to other more suitable color spaces, HSV or LAB color space, which can reduce the impact of illumination changes on the detection results. Apply a Gaussian filter or a median filter to remove noise from the video frame and normalize the image pixel values ​​to limit them to the interval [0, 1] for subsequent processing. Send the pre-processed video frame to the selected backbone network, and extract the output of the intermediate layer (conv4_x or conv5_x layer of ResNet50) as the feature map. These feature maps contain information in multiple dimensions, such as color information, texture details, and shape contours, laying a solid foundation for subsequent target detection. In order for the model to better understand the positional relationship of the features, position encoding information is added on the basis of the feature map. It is generated by fixed or learned methods to ensure that each feature point carries its spatial position information. The visual features obtained above are combined with the position encoding to form a comprehensive feature embedding. This embedding contains rich semantic information and retains the relative position of the object in the image, preparing for the next step of target detection.

[0005] Preferably, in said S2, further comprising: Based on Gaussian distribution and statistical models, multiple random candidate detection frames are generated. These candidate frames cover the areas that may contain the target and provide candidate regions for the generation of feature detection frames. The size, scale and position of the candidate regions are adjusted to adapt to the variable forms of the targets in the video. The features extracted by deep learning are mapped to these candidate regions to generate more accurate feature detection frames. Based on target detection, the target's motion information and context information are used to continuously track the target through a tracking algorithm (Kalman filter or particle filter).

[0006] Preferably, in said S3, further comprising: Based on the generated comprehensive feature embedding, a region proposal network (RPN) or other methods are used to generate multiple randomly distributed candidate detection boxes. These candidate boxes cover the areas that may contain the target and provide candidate areas for the generation of feature detection boxes. Non-maximum suppression is performed on the generated candidate boxes to remove redundant boxes with high overlap and only retain the detection boxes that are most likely to contain the target. This step helps to reduce the subsequent computational burden and improve efficiency. For each detection box processed by NMS, the similarity score (IoU - Intersection over Union) between it and the known category is calculated. Based on these scores, a cost matrix is ​​constructed, where the rows represent the detection boxes in the current frame and the columns represent the tracked targets in the previous frame. The classic Hungarian algorithm (also known as the Kuhn-Munkres algorithm) is used to solve the best matching problem of the bipartite graph. The algorithm can find a matching solution that minimizes the total cost in polynomial time, that is, find the most suitable corresponding target for the detection box in each frame. Based on the results of the Hungarian algorithm, candidate targets with high matching are screened. For those detection boxes that fail to match successfully, they are considered to be new targets or false detections. The successfully matched feature detection box is input into the classification prediction network (fully connected layer or convolutional layer) to predict the specific category, location and confidence score of the target. In this way, not only can the identity of the target be confirmed, but also the possibility of its existence can be quantitatively evaluated.

[0007] Preferably, in said S4, further comprising: Optimize model performance and avoid overfitting through regularization techniques, hyperparameter adjustment and other methods. When evaluating multi-classification models, use precision, recall and F1 scores to measure performance, reflecting prediction accuracy, comprehensive recognition and overall performance respectively. By combining these indicators, ensure that the model shows good balance and robustness when processing data of various categories, thereby improving the reliability and accuracy of the model in practical applications. Integrate the collected user feedback into the system to further fine-tune and optimize the algorithm. The model iterates itself based on the evaluation results and user feedback to continuously improve the accuracy of detection and classification.

[0008] Compared with the prior art, the present invention has the following beneficial effects: 1. Improve detection accuracy: By combining deep learning and traditional computer vision technology, the present invention can accurately detect and track targets in complex open environments. The mean average precision (mAP) has been significantly improved to 78.2%, which is about 9.7 percentage points higher than the existing technology. The F1 score has also been significantly improved, from 0.75 to 0.84. 2. Enhanced generalization ability: The use of technologies such as random region generator and bipartite matching network enhances the model's generalization ability for unknown categories. The mean average precision (mAP) increased from 62.7% to 73.5%, an increase of 10.8%, and the F1 score increased from 0.72 to 0.81; 3. Improve real-time performance: The algorithm design takes real-time requirements into consideration. Through efficient feature extraction and matching mechanisms, it can respond quickly under limited computing resources. 4. Strong scalability: Through incremental learning, the present invention can continuously learn new categories and keep the model updated and upgraded. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative work: Figure 1 : Video processing and object detection flow chart, showing the various steps of video preprocessing and candidate box generation and detection; Figure 2 : Target recognition and model optimization flowchart, showing the various steps of matching and classification prediction and model optimization. DETAILED DESCRIPTION

[0010] Step 1: Video preprocessing Preprocess the input video, including adjusting the resolution, color space conversion, etc. The specific steps are as follows: S11: Resolution adjustment: Adjust the video resolution to 1080P to reduce the amount of calculation and increase the processing speed. The image resolution is flexibly adjusted by using bilinear interpolation technology or nearest neighbor interpolation method; S12: Color space conversion: Convert RGB format to grayscale or other color spaces more suitable for feature extraction, including HSV or LAB color space, to reduce the amount of calculation and improve the robustness of feature extraction; S13: Noise elimination: In order to improve the image quality of the video frame, Gaussian filtering and median filtering are used to effectively remove the noise; S14: Normalization processing: normalizing the image pixel values ​​to limit these values ​​to the range of 0 to 1; S15: Input the preprocessed video frames into the backbone network (ResNet50 or VGG16) for feature extraction to obtain feature maps. When building the model, choose the weights pre-trained on the large dataset ImageNet to initialize the backbone of the model, which can significantly speed up the training process and provide a higher starting point performance for the model. Select the output of the middle layer in the backbone network as the feature map, including the conv4_x layer or conv5_x layer of ResNet50; S16: Pass the feature map to the position encoding network to generate position encoding information. Position encoding is used to enhance the model's understanding of feature position information and ensure that the model can accurately locate the target. When generating position encoding, fixed or learned encoding methods are usually selected.

[0011] Step 2: Candidate box generation and detection S21: Use the region generator to generate multiple random detection boxes. Use sliding window technology or anchor box generation strategy to generate candidate region boxes. The sliding window method generates candidate boxes by sliding a fixed-size window on the image, and the anchor box method generates candidate boxes of different scales and proportions on the feature map. The number of generated candidate boxes can be adjusted according to the complexity of the video frame, and usually 1000 to 2000 candidate boxes are generated; S22: Input the feature embedding and random detection frames into the decoder (RPN network) to generate more accurate feature detection frames. These detection frames are used for preliminary positioning of the target. The feature detection frames are generated using the region proposal network, which consists of a convolutional layer and a fully connected layer. The convolutional layer is responsible for extracting local features of the image, while the fully connected layer is responsible for outputting the final detection frame coordinates. Non-maximum suppression is performed on the feature detection frames output by the region proposal network to eliminate redundant overlapping frames and retain only the detection frames with the highest scores as the final results; S23: Input the generated feature detection frame into a bipartite matching network to obtain matching information corresponding to a preset detection category. The bipartite matching network can use the Hungarian algorithm or other matching algorithms to evaluate the matching degree between the feature detection frame and the known detection category and screen out candidate targets with high matching degree.

[0012] Step 3: Matching and classification prediction S31: Use the Hungarian algorithm or IoU threshold for matching. The IoU threshold is usually set to 0.5 or higher. The successfully matched feature detection box is input into the classification prediction network (fully connected layer or convolutional layer) to predict the category, location and confidence of the target. The classification prediction network should be fully trained to ensure high-precision and high-confidence prediction results. When calculating the loss, the cross entropy loss function is used to evaluate the error of the classification task, and the smoothed L1 loss function is used to measure the error of the regression task; S32: Output target category, bounding box and score. The output result can be a JSON file or CSV file containing the category, bounding box coordinates and confidence. It can be used for subsequent tracking and recognition tasks, target trajectory generation, behavior analysis, etc. The detection results are superimposed on the original video frame to generate a video with annotated boxes for easy viewing and verification; S33: By analyzing the background frame and feature clustering, the K-means clustering algorithm is used to cluster the features in combination with unsupervised learning and semi-supervised learning methods. First, a series of feature vectors representing different regions or objects are obtained from the feature map after feature extraction. These feature vectors can be feature descriptors of a single pixel or local feature pools in a larger range. Randomly select k samples as the initial cluster centers, where k is the specified number of categories, which can be pre-set according to the specific application scenario, or the optimal k value can be automatically determined by methods such as the elbow rule and the silhouette coefficient. Calculate the distance between each feature vector and all cluster centers (commonly used Euclidean distance), and assign each feature vector to the cluster with the closest distance. This step constructs a preliminary clustering result. Recalculate the average value of all members in each cluster as the new cluster center. This step makes the cluster center closer to the actual data distribution. Repeat steps 3 and 4 until the cluster center no longer changes significantly or reaches the predetermined maximum number of iterations, and finally obtain a set of stable cluster centers and corresponding classification results. Unknown categories are identified based on clustering results, and verified and annotated in combination with contextual information and other sensor data for subsequent learning and tracking. Streaming learning technology is used to gradually optimize support vector machines or real-time gradient descent algorithms, dynamically adjust and update model parameters, so that the model can continuously adapt to new data distributions. At the same time, model snapshots are saved regularly to facilitate backtracking and restoring model status.

[0013] Step 4: Model optimization S41: Reduce feature sparsity by using L1 regularization and use L2 regularization to prevent weights from being too large. Combine grid search, random search, and Bayesian optimization methods to manually adjust hyperparameters such as learning rate, batch size, and optimizer to verify the model, optimize model performance, and avoid overfitting. In order to ensure that the model performs well on new data, divide the original data set into several subsets, randomly select one from the divided subset as the test set, and combine all other subsets as the training set. In this way, the model can be trained and tested multiple times, and finally the average performance indicator is taken as the evaluation result of the model; S42: Before using user feedback data for model training, the data needs to be cleaned and preprocessed to ensure data quality. Use the cleaned user feedback data to fine-tune the existing model. Fine-tuning is based on the original model, and a small number of parameters are updated for specific tasks or fields to adapt to the new data distribution; S43: Integrate the collected user feedback into the system, supplement the deficiencies of the model, and adjust and optimize the model. The annotated data provided by the user is a direct correction to the model prediction results, providing opinions or suggestions on the model performance, supplementing the deficiencies of the model, and improving the model's capabilities.

Claims

1. An open environment dynamic target tracking and recognition algorithm for short video platforms, characterized in that: The method comprises the following steps: S1: Use deep learning convolutional neural network architecture to extract rich visual features from video frames and add position information to generate comprehensive feature embeddings; S2: Generate multiple random candidate detection boxes based on Gaussian distribution and statistical models, use the feature embedding and candidate boxes generated by the encoder, and generate more accurate feature detection boxes through sampling and feature mapping; S3: The Hungarian graph matching algorithm is used to evaluate the matching degree between the feature detection box and the known detection category, and the candidate targets with high matching degree are screened out. The classifier is used to predict the category, location and confidence of the target. S4: Combine unsupervised learning and semi-supervised learning methods to automatically identify targets of unknown categories, and continuously update model parameters to adapt to new data through an incremental learning mechanism.

2. The method according to claim 1, characterized in that The S1 further comprises: Choose a pre-trained deep learning model as the backbone network; Resize the input video frame to a resolution suitable for the backbone network input size; Resize the image using bilinear interpolation or nearest neighbor interpolation techniques; Convert RGB format to other color spaces to reduce the impact of lighting changes; Apply a Gaussian filter or a median filter to remove noise from the video frame; Normalize the image pixel values.

3. The method according to claim 1, characterized in that: The S2 further includes: By adjusting the size, scale and position of the candidate region, it can adapt to the changing shapes of the target in the video; The target's motion information and context information are used to continuously track the target through a tracking algorithm.

4. The method according to claim 1, characterized in that: The S3 further includes: Perform non-maximum suppression on the generated candidate boxes to remove redundant boxes with high overlap; According to the results of the Hungarian algorithm, candidate targets with high matching degree are screened out; The successfully matched feature detection box is input into the classification prediction network to predict the specific category, location and confidence score of the target.

5. The method according to claim 1, characterized in that The S4 further comprises: Optimize model performance and avoid overfitting through regularization techniques, hyperparameter adjustment, and other methods; Integrate the collected user feedback into the system to further fine-tune and optimize the algorithm.