Target detection tracking method combining YOLOv8 detection algorithm and KCF tracking algorithm

By combining the YOLOv8 detection algorithm and the KCF tracking algorithm, real-time and accurate tracking of the target vehicle in the video stream is achieved, and the problems of real-time and tracking stability in the prior art are solved.

CN120147608AActive Publication Date: 2025-06-13BEIJING TOPMOO TECH

Patent Information

Application Number
CN202510201017.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

Existing single-object detection or tracking algorithms are difficult to achieve real-time and efficient tracking in video streams, especially when target appearance changes, occlusions or rapid movements are prone to losing targets.

Method used

Combining the YOLOv8 detection algorithm and the KCF tracking algorithm, the KCF tracker is initialized by using YOLOv8 to detect the first frame image, and then the KCF tracker processes the subsequent frame image. If the tracking failure occurs in the t-th frame, YOLOv8 is re-detected to try to retrieve the target. If successful, the target detection image is updated and the KCF tracker is re-initialized; if the target is not detected, the target detection image of the t-th frame is predicted based on the previous t-1 frame information.

Benefits of technology

It effectively ensures the continuity and accuracy of target vehicle object tracking, and solves the problems of real-time and tracking stability in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147608A_ABST
    Figure CN120147608A_ABST
Patent Text Reader

Abstract

The invention relates to the field of target detection and tracking, and provides a target detection and tracking method combining a YOLOv8 detection algorithm and a KCF tracking algorithm, and the method comprises the steps: firstly obtaining a video stream time sequence set of a tracking image of a target vehicle object, carrying out the target detection of a first frame through employing YOLOv8, so as to initialize a KCF tracker, then processing a subsequent frame image through the KCF tracker, and carrying out the target tracking, and if tracking fails in the tth frame, re-starting the YOLOv8 to detect the frame so as to try to retrieve the target, if the target is successfully detected, updating and generating a new target vehicle object detection image and re-initializing the KCF tracker, and if the target is not detected, predicting and generating a target vehicle object detection image of the tth frame based on the information of the first (t-1) frames. In this way, the continuity and accuracy of target vehicle object tracking can be effectively ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of object detection and tracking, and more specifically, to an object detection and tracking method that combines the YOLOv8 detection algorithm with the KCF tracking algorithm. Background Art

[0002] In fields such as intelligent transportation systems, autonomous vehicles, and video surveillance, accurately and efficiently detecting and tracking moving objects (such as vehicles, pedestrians, etc.) is crucial. Precise object detection and tracking can not only improve the overall performance of the system but also ensure traffic safety, optimize traffic flow, and promptly detect abnormal events, providing reliable guarantees for intelligent monitoring and autonomous driving.

[0003] However, existing single object detection or tracking algorithms have obvious limitations. Specifically, traditional object detection algorithms (such as the YOLO series) can identify target vehicles in static images, but in a video stream, due to the need to comprehensively detect each frame of the image, the computational complexity is huge, making it difficult to meet the real-time requirements. At the same time, when using a single tracking algorithm (such as KCF), although the computational efficiency is relatively high during the tracking process, once the target vehicle undergoes appearance changes, occlusion, or rapid movement, etc., the tracker is easily lost and the tracking fails. Moreover, KCF depends on the high-quality detection results of the initial frame. If the initialization is inaccurate, the subsequent tracking effect will be affected.

[0004] Therefore, an optimized object detection and tracking solution is desired. Summary of the Invention

[0005] This application aims at the shortcomings in the prior art and provides an object detection and tracking method that combines the YOLOv8 detection algorithm with the KCF tracking algorithm.

[0006] According to one aspect of this application, there is provided an object detection and tracking method that combines the YOLOv8 detection algorithm with the KCF tracking algorithm, which includes:

[0007] Obtain a video stream time series set of tracking images of a target vehicle object;

[0008] Use the YOLOv8 detection algorithm to perform object detection on the first frame of the tracking image in the video stream time series set of the tracking images to obtain a target vehicle object detection image;

[0009] Initialize the KCF tracker based on the target vehicle object detection image, and use the initialized KCF tracker to perform object tracking on other frames of the tracking images in the video stream time series set;

[0010] In response to the KCF tracker reporting that target tracking fails in the t-th frame tracking image in the video stream timing set of the tracking images, the YOLOv8 detection algorithm is used to re-detect the target in the t-th frame tracking image to obtain a target detection result;

[0011] In response to the target detection result being that the target vehicle object is re-detected, a t-th frame target vehicle object detection image is generated, and the KCF tracker is re-initialized based on the t-th frame target vehicle object detection image;

[0012] In response to the target detection result being that the target vehicle object is not detected, a target vehicle object prediction is performed based on the t - 1 tracking images before the t-th frame tracking image to generate a t-th frame target vehicle object detection image.

[0013] Due to the adoption of the above technical solutions, the present application has significant technical effects:

[0014] The target detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm provided by the present application first obtains a video stream timing set of tracking images of a target vehicle object, and uses YOLOv8 to perform target detection on the first frame to initialize the KCF tracker. Subsequently, the KCF tracker processes subsequent frame images for target tracking. If tracking fails in the t-th frame, YOLOv8 is re-enabled to detect this frame to try to retrieve the target. If the target is successfully detected, a new target vehicle object detection image is updated and generated, and the KCF tracker is re-initialized. If the target fails to be detected, a t-th frame target vehicle object detection image is predicted based on the information of the previous t - 1 frames. In this way, the continuity and accuracy of target vehicle object tracking can be effectively ensured. Description of the Drawings

[0015] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0016] Figure 1 It is a flowchart of the target detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to the embodiment of the present application.

[0017] Figure 2 It is a flowchart of step S6 in the target detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to the embodiment of the present application.

[0018] Figure 3 It is a flowchart of step S61 in the object detection and tracking method that combines the YOLOv8 detection algorithm and the KCF tracking algorithm according to an embodiment of the present application.

[0019] Figure 4 It is a flowchart of step S64 in the object detection and tracking method that combines the YOLOv8 detection algorithm and the KCF tracking algorithm according to an embodiment of the present application.

[0020] Figure 5 It is a flowchart of step S642 in the object detection and tracking method that combines the YOLOv8 detection algorithm and the KCF tracking algorithm according to an embodiment of the present application. Detailed implementation manners

[0021] Next, exemplary embodiments of the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0022] In the fields of intelligent transportation systems, autonomous vehicles, and video surveillance, accurately and efficiently detecting and tracking moving objects such as vehicles and pedestrians is extremely crucial. This can not only improve the system performance but also ensure traffic safety, optimize traffic flow, and promptly detect abnormal situations, providing a solid guarantee for intelligent monitoring and autonomous driving. However, existing single algorithms have limitations: Traditional object detection algorithms (such as the YOLO series) can identify objects in static images, but when processing video streams, due to the need to comprehensively detect each frame, the computational load is large, making it difficult to achieve real-time performance; while separate tracking algorithms (such as KCF), although the tracking efficiency is relatively high, are prone to losing the target in the face of target appearance changes, occlusion, or rapid movement, and their performance highly depends on the high-quality detection results of the initial frame. Inaccurate initialization will significantly affect the subsequent tracking effect.

[0023] It should be understood that the YOLOv8 detection algorithm is a deep learning-based object detection algorithm, belonging to the single-stage object detection algorithm. It continues the core idea of the YOLO series of algorithms, that is, predicting the class and bounding box position of the object through a neural network, and can complete the object detection task in one forward propagation process. YOLOv8 has been optimized and improved in architecture design, adopting a more efficient backbone network, neck network and detection head structure, with faster detection speed and higher detection accuracy. Since YOLOv8 can complete the detection of objects in the image in an extremely short time, it is suitable for fast object localization in real-time video streams. The KCF tracking algorithm is an object tracking algorithm based on correlation filters. It realizes efficient single-object tracking by using circulant matrices and the fast Fourier transform (FFT). And this algorithm uses kernel methods to handle non-linear features, making it have good robustness in complex scenarios. YOLOv8 and KCF can be combined and applied to the object detection and tracking algorithm. Specifically, YOLOv8 is responsible for quickly and accurately detecting objects, while KCF uses the detection results for efficient and stable tracking.

[0024] Based on this, the present application proposes an object detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm. By combining advanced technologies, the two algorithms are combined to solve the limitations of the prior art and achieve precise, stable and real-time detection and tracking of the target vehicle. Figure 1 The flowchart of the object detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to the embodiments of the present application. As Figure 1As shown, the object detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to the embodiments of the present application includes: S1, obtaining a video stream time series set of tracking images of a target vehicle object; S2, using the YOLOv8 detection algorithm to perform object detection on the first frame of tracking image in the video stream time series set of the tracking images to obtain a target vehicle object detection image; S3, initializing the KCF tracker based on the target vehicle object detection image, and using the initialized KCF tracker to perform object tracking on other frames of tracking images in the video stream time series set of the tracking images; S4, in response to the KCF tracker reporting that target tracking fails in the t-th frame of tracking image in the video stream time series set of the tracking images, using the YOLOv8 detection algorithm to re-perform object detection on the t-th frame of tracking image to obtain a target detection result; S5, in response to the target detection result being that the target vehicle object is redetected, generating a t-th frame target vehicle object detection image, and re-initializing the KCF tracker based on the t-th frame target vehicle object detection image; S6, in response to the target detection result being that the target vehicle object is not detected, performing target vehicle object prediction based on the t-1 tracking images before the t-th frame of tracking image to generate a t-th frame target vehicle object detection image.

[0025] In step S1, a video stream time series set of tracking images of a target vehicle object is obtained. It should be understood that the video stream time series set of tracking images of the target vehicle object can provide necessary input data for the subsequent object detection and tracking processes. Specifically, the video stream time series set of tracking images of the target vehicle object contains the target vehicle that needs to be analyzed, detected, and tracked. By obtaining a continuous video stream, dynamic monitoring of the target vehicle can be performed, not limited to the information in static images, but capable of tracking the changes of the target vehicle object in both the time and space dimensions.

[0026] In step S2, the YOLOv8 detection algorithm is used to perform object detection on the first frame of tracking image in the video stream time series set of the tracking images to obtain a target vehicle object detection image. Correspondingly, considering that YOLOv8 is an advanced object detection algorithm, it has high accuracy and fast detection speed in object detection tasks. In the first frame image, the target vehicle may be in a complex scene, and there may be other interference factors around it, such as other vehicles, pedestrians, obstacles, etc. YOLOv8 can use its powerful feature extraction and classification capabilities to accurately identify the target vehicle object and exclude the influence of other interference factors, thus providing reliable target information for subsequent tracking.

[0027] The following is a detailed elaboration of a specific implementation process for "performing object detection on the first frame of the tracking image in the video stream time series of the tracking image using the YOLOv8 detection algorithm to obtain the object vehicle object detection image":

[0028] First, a specific model needs to be selected. There are multiple versions of YOLOv8, such as YOLOv8n, YOLOv8s, YOLOv8m, YOLOv8 l, and YOLOv8x, and each version has differences in model size, speed, and accuracy. It is necessary to select a suitable version according to the specific application scenario and performance requirements. If real-time performance is emphasized, the lightweight YOLOv8n can be selected, and if high accuracy is required, the larger YOLOv8x can be selected. Usually, a model pre-trained on a large-scale dataset such as COCO is loaded. The pre-trained model contains rich general object feature information, which can shorten the training time and improve the detection effect.

[0029] Next, image preprocessing operations are performed. The first frame of the tracking image is extracted from the video stream time series and read into a processable data format using an image processing library. To adapt to the input requirements of the YOLOv8 model, the image size is adjusted to a specified size. Unifying the size can improve the processing efficiency and detection stability. At the same time, the pixel values of the image are normalized and scaled to the range of [0,1] or [-1,1] to reduce the influence of factors such as illumination changes.

[0030] After that, it enters the object detection inference stage. The preprocessed first-frame image is input into the loaded YOLOv8 model. Its backbone network will extract features from the image and convert the input image into feature maps of different scales and abstraction levels through operations such as convolution and pooling. These feature maps contain feature information such as the edges, textures, and shapes of the objects. The feature maps are then processed by the neck network, which will fuse and enhance the feature maps of different scales. Through the PANet structure, high-level semantic features and low-level detail features are effectively fused to improve the model's detection ability for objects of different sizes. Finally, the processed feature maps enter the detection head, which will parse the feature maps, predict the categories, bounding box positions, and confidence levels of the possible objects in the image, predict multiple bounding boxes at each grid point of the feature map, and assign class probabilities and confidence scores to them.

[0031] Post-processing is required after the inference is completed. Since the detection head will predict a large number of prediction boxes, many of which may overlap and point to the same object, non-maximum suppression is used to remove redundant prediction boxes. First, sort by confidence level, select the prediction box with the highest confidence level as the prediction box, calculate the intersection over union of other prediction boxes with the prediction box, and filter out the prediction boxes with an overlap degree exceeding the threshold. At the same time, a confidence level threshold is set to filter out the prediction boxes with a confidence level lower than this threshold to remove possible false detections. Then, according to the detected object category, the prediction boxes belonging to the vehicle category are selected.

[0032] Finally, the target vehicle object detection image is generated. Since the target vehicle refers to a single vehicle in this application, among the selected vehicle bounding boxes, the prediction box with the highest confidence is chosen as the detection result of the target vehicle object. According to the prediction box information of the selected target vehicle object, a bounding box is drawn on the first frame tracking image to visually display the position of the target vehicle. The class label of the target vehicle (such as "vehicle") and the confidence score are added near the bounding box to facilitate the user's understanding of the accuracy of the detection result. Then, the first frame tracking image with the drawn bounding box and label is saved as the target vehicle object detection image.

[0033] In step S3, the KCF tracker is initialized based on the target vehicle object detection image, and the initialized KCF tracker is used to perform target tracking on other frame tracking images in the video stream time series of the tracking images. It should be understood that KCF is an efficient tracking algorithm based on kernel correlation filtering, which can significantly reduce the demand for computing resources while maintaining a relatively high tracking accuracy. By using the initial target position information provided by the YOLOv8 detection algorithm, the KCF tracker can be initialized quickly and accurately. In a video stream, the target vehicle usually appears continuously in multiple consecutive frames. Initializing the KCF tracker based on the target detection result of the first frame can ensure continuous and stable tracking of the target in subsequent frames, and can effectively handle situations where the target undergoes slight deformation or occlusion.

[0034] The following is a detailed elaboration of a specific implementation process of "initializing the KCF tracker based on the target vehicle object detection image":

[0035] When initializing the KCF tracker based on the target vehicle object detection image, the target information needs to be accurately extracted from the target vehicle object detection image first. The target vehicle in the target vehicle object detection image is marked in the form of a bounding box. The upper left coordinates (x, y), width w, and height h of the bounding box are extracted, which precisely define the position and range of the target vehicle in the image. At the same time, features are extracted from the target area enclosed by the bounding box. Gray or color features are commonly used in the KCF tracker. If gray features are used, the color image will be converted into a grayscale image to highlight the shape information and reduce the amount of data. When using color features, the color characteristics of the target are described according to different color spaces such as RGB and HSV, providing rich information for subsequent tracking.

[0036] Next, the extracted target features are used to construct a target template. The KCF tracker is based on kernel correlation filtering technology and needs to construct a template that can represent the target vehicle. In the initialization stage, the extracted features are used to construct an initial kernel correlation filtering model. This model learns the target appearance pattern by calculating the correlation between target features. Specifically, the target features are mapped to a high-dimensional space, and a kernel function such as a Gaussian kernel is used to calculate the similarity, obtaining the correlation filtering response of the target. Subsequently, the extracted target features are used to train the correlation filter. During the training process, the filter parameters are continuously adjusted to minimize the error between the predicted output and the expected output. The expected output is generally a two-dimensional Gaussian distribution with a peak around the target center, which helps the filter learn the position information of the target in the image.

[0037] After that, the initial state of the KCF tracker is set. The initial position of the target vehicle in the first frame image is determined according to the target bounding box information, which is the starting point for subsequent tracking. At the same time, key parameters are set for the KCF tracker. The learning rate determines the speed at which the tracker updates the target template. A higher learning rate enables the tracker to quickly adapt to changes in the target appearance but is vulnerable to noise; a lower learning rate makes the tracker more stable but may not be able to keep up with rapid changes in the target. The size of the search area is also crucial, which determines the range within which the tracker searches for the target in subsequent frames. An appropriate size can improve the tracking efficiency and accuracy.

[0038] Finally, the initialization result is verified. After completion of initialization, the extracted target features and the constructed target template are carefully checked. The initialization effect can be judged by visualizing the target features or observing the correlation filter response. If the target features are blurred or the filter response is abnormal, re-initialization is required. Before formal tracking, the initialized tracker is used to perform simulated tracking on the current frame image to check whether the target vehicle can be accurately located. If the result of the simulated tracking is not satisfactory, all aspects of the initialization should be comprehensively checked to identify problems and make adjustments.

[0039] In step S4, in response to the KCF tracker reporting that target tracking fails in the t-th frame tracking image in the video stream timing set of the tracking image, the YOLOv8 detection algorithm is used to re-detect the target in the t-th frame tracking image to obtain the target detection result. It should be understood that although the KCF tracking algorithm has high tracking efficiency, it also has some limitations. For example, when the target vehicle is severely occluded, moving rapidly, undergoing drastic appearance changes (such as large appearance differences caused by changes in the viewing angle when the vehicle turns), or when the lighting conditions change drastically, the KCF tracker may lose the target, resulting in tracking failure. Once this situation occurs, relying solely on KCF itself cannot continue to accurately track the target, and a more powerful target detection mechanism needs to be introduced. At this time, reusing the powerful YOLOv8 target detection algorithm can help the system re-locate and recover the lost target. That is, YOLOv8 performs comprehensive feature extraction and target classification on the entire image during each detection, without relying on the tracking information of the previous frame. Therefore, when the KCF tracking fails, using YOLOv8 to re-detect the current frame can re-locate the target vehicle from a global perspective, unaffected by the previous tracking failure.

[0040] In step S5, in response to the target detection result being that the target vehicle object is re-detected, a t-th frame target vehicle object detection image is generated, and the KCF tracker is re-initialized based on the t-th frame target vehicle object detection image. Correspondingly, considering that during re-detection, the target model and position information stored internally can no longer accurately reflect the current state of the target vehicle. This may be due to reasons such as the target being occluded, moving rapidly, undergoing large appearance changes, or changes in lighting conditions, making the target features and position information learned by the KCF tracker based on the previous frame no longer applicable. Therefore, it is necessary to re-initialize it with accurate target information.

[0041] In step S6, in response to the target detection result being that the target vehicle object is not detected, target vehicle object prediction is performed based on the t - 1 tracking images before the t-th frame tracking image to generate a t-th frame target vehicle object detection image. It should be understood that in the target detection and tracking task, continuously tracking the target vehicle is a key requirement. The previous t - 1 tracking images contain rich historical information such as the shape and color of the target vehicle. This information can reflect the movement laws and trends of the target vehicle. Using this historical information for prediction can, to a certain extent, make up for the information loss caused by the detection failure in the current frame. For example, based on the shape and color of the target vehicle in the past few frames, its possible position in the t-th frame can be roughly inferred.

[0042] Based on this, the technical concept of this application is to use image analysis and feature extraction techniques based on machine vision to perform HOG feature extraction and color histogram calculation on each of the previous t - 1 tracking images. Then, semantic capture and multi - dimensional semantic fusion are performed on the extracted histogram of oriented gradients of each tracking image and the color histogram of each tracking image. Based on this, the dynamic causal context walk representation between the multi - dimensional semantic fusion characterization features of each tracking image after fusion is used to automatically predict and obtain the object detection image of the target vehicle in the t - th frame. In this way, through HOG feature extraction and color histogram calculation, the shape and color information of the target vehicle can be comprehensively described from different angles. At the same time, through the dynamic causal context walk representation, the dynamic changes of the target vehicle over time and its interaction with the environment are captured, thereby improving the understanding and prediction accuracy of the behavior of the target vehicle.

[0043] Specifically, Figure 2 It is a flowchart of step S6 in the object detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to an embodiment of the present application. As Figure 2 shown, the step S6 includes: S61, respectively performing HOG feature extraction and color histogram calculation on each of the t - 1 tracking images to obtain a sequence of histograms of oriented gradients of the tracking images and a sequence of color histograms of the tracking images; S62, performing histogram feature extraction on the sequence of histograms of oriented gradients of the tracking images and the sequence of color histograms of the tracking images to obtain a sequence of semantic feature vectors of the histograms of oriented gradients of the tracking images and a sequence of semantic feature vectors of the color histograms of the tracking images; S63, concatenating each corresponding semantic feature vector of the histogram of oriented gradients of the tracking image and the semantic feature vector of the color histogram of the tracking image in the sequence of semantic feature vectors of the histogram of oriented gradients of the tracking images and the sequence of semantic feature vectors of the color histograms of the tracking images to obtain a sequence of multi - dimensional semantic fusion representation vectors of the tracking images; S64, performing dynamic walk of the tracking image with causal modeling on the sequence of multi - dimensional semantic fusion representation vectors of the tracking images to obtain a semantic encoding vector of the dynamic walk of the tracking image context; S65, based on the semantic encoding vector of the dynamic walk of the tracking image context, obtaining the object detection image of the target vehicle in the t - th frame.

[0044] In step S61, HOG feature extraction and color histogram calculation are respectively performed on each of the t - 1 tracking images to obtain a sequence of histograms of oriented gradients of the tracking images and a sequence of color histograms of the tracking images. Specifically, Figure 3 It is a flowchart of step S61 in the object detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to an embodiment of the present application. As Figure 3As shown, step S61 includes: S611, performing HOG feature extraction on each of the t - 1 tracking images to obtain a sequence of histograms of oriented gradients of the tracking images; S612, calculating the color histogram of each of the t - 1 tracking images to obtain a sequence of color histograms of the tracking images.

[0045] In step S611, HOG feature extraction is performed on each of the t - 1 tracking images to obtain a sequence of histograms of oriented gradients of the tracking images. Correspondingly, considering that the shape of the target vehicle is one of its important features, in different frames, although the target vehicle may undergo position movement, pose change, or be affected by factors such as illumination, the shape feature is relatively stable. And considering that HOG features can effectively capture the shape and structure information of objects in the image. Based on this, in the technical solution of this application, HOG feature extraction is performed on each of the t - 1 tracking images to obtain a sequence of histograms of oriented gradients of the tracking images. That is, by performing HOG feature extraction on each frame of the image, the shape description of the target vehicle at different times can be obtained, providing a basis for subsequent analysis. For example, features such as the body contour and wheel shape of the vehicle can be well reflected in the HOG features. Even when the angle of the vehicle in the image changes, its basic shape features can still be reflected through the HOG features. And each histogram of oriented gradients of the tracking image records the shape features of the target vehicle at that moment, and the sequence composed of these histograms can reflect the change of the shape features of the target vehicle over time. For example, when the vehicle turns, the projection of its shape in different frames will change, and this change trend can be observed through the HOG feature sequence, thereby providing a basis for predicting the state of the target vehicle in the t-th frame.

[0046] The following is a detailed elaboration of a specific implementation process of "performing HOG feature extraction on each of the t - 1 tracking images to obtain a sequence of histograms of oriented gradients of the tracking images":

[0047] First, image preprocessing operations need to be performed. Since HOG features mainly focus on the gray - level changes in the image to highlight the object shape features, each tracking image needs to be converted from color to grayscale to reduce the data volume. To eliminate the influence of different illumination intensities between images on feature extraction, the grayscale image also needs to be normalized, mapping the image gray - level values to a range such as [0, 1] or [-1, 1]. In addition, the noise present in the image may interfere with feature extraction, so a Gaussian filter is used to perform a convolution operation on the image for Gaussian smoothing, making the image smoother and reducing the influence of noise.

[0048] After the image preprocessing is completed, the calculation of the image gradient begins. Calculating the image gradient is to obtain the change of gray values in the image, which is calculated separately in the horizontal and vertical directions. Usually, the Sobel operator is used to achieve this. Through its convolution operation with the image, the gradient Gx in the horizontal direction and the gradient Gy in the vertical direction of the image are obtained. Based on these two gradient components, the gradient magnitude and the gradient direction are further calculated.

[0049] Next, the image is divided into multiple small cells. The cell is the basic unit for HOG feature calculation, and its size is usually fixed. Commonly, there are 8x8 pixels or 16x16 pixels. Each cell independently performs the subsequent calculation of the histogram of oriented gradients. Through this division method, the gradient information can be carefully counted within a local range, and then the local shape features of the object can be captured more accurately.

[0050] After the cell division is completed, the calculation of the histogram of oriented gradients for each cell begins. The specific approach is to divide the gradient direction into several intervals. For example, the range from 0 - 180° is divided into 9 intervals, with each interval being 20°. Within each cell, based on the gradient direction and magnitude information of the pixels, a histogram of oriented gradients is constructed, and the sum of the gradient magnitudes within each interval is counted. When counting, according to the angle of the gradient direction, determine the interval it belongs to, and then accumulate the corresponding gradient magnitude into the statistical value of that interval. In this way, each cell can be represented by a vector containing the statistical values of multiple intervals, that is, the histogram of oriented gradients.

[0051] Finally, for each of the t - 1 tracking images, the HOG feature extraction is strictly performed according to the above steps from image preprocessing to calculating the histogram of oriented gradients of cells, so as to obtain the histogram of oriented gradients of each image. Arranging these histograms of oriented gradients in the chronological order of the images forms a sequence of histograms of oriented gradients of the tracking images.

[0052] In step S612, the color histograms of the respective tracking images among the t - 1 tracking images are calculated separately to obtain a sequence of the tracking image color histograms. Correspondingly, considering that color is a very prominent and distinguishable visual feature of the target vehicle. Vehicles of different brands and types often have unique colors. For example, fire trucks are usually red, ambulances are mostly white with special blue or red markings, etc. Even in a complex background environment, color information can help quickly identify and locate the target vehicle. Based on this, in this application, the color histograms of the respective tracking images among the t - 1 tracking images are calculated separately to obtain a sequence of the tracking image color histograms. In this way, by calculating the color histogram for each frame of the image, a sequence of color histograms is formed. This sequence records the color distribution characteristics of the target vehicle at different time points, reflects the change of color features over time, and provides a basis for subsequent prediction.

[0053] In step S62, histogram feature extraction is performed on the sequence of tracking image directional gradient histograms and the sequence of tracking image color histograms to obtain a sequence of tracking image directional gradient histogram semantic feature vectors and a sequence of tracking image color histogram semantic feature vectors. Specifically, in an embodiment of the present application, step S62 includes: passing the sequence of tracking image directional gradient histograms and the sequence of tracking image color histograms through a histogram feature extractor based on the FCN spatial model to obtain a sequence of tracking image directional gradient histogram semantic feature vectors and a sequence of tracking image color histogram semantic feature vectors. Accordingly, considering that the tracking image directional gradient histogram and the color histogram contain basic information such as the shape and color distribution of the target vehicle, this information is a relatively shallow and intuitive statistical result. Therefore, in order to further refine the deep semantic information contained in each histogram, the present application passes the sequence of the tracking image directional gradient histogram and the sequence of the tracking image color histogram through a histogram feature extractor based on the FCN spatial model to obtain a sequence of tracking image directional gradient histogram semantic feature vectors and a sequence of tracking image color histogram semantic feature vectors. It should be understood that the FCN spatial model has powerful convolution operations and feature mapping capabilities. It can gradually abstract and extract features from the input histogram through multiple layers of convolutional layers, and convert the original histogram data into a more representative and semantic feature representation. For example, it can learn the association between the shape and color features of the target vehicle and the target category, state, etc., thereby digging out the semantic information hidden in the histogram.

[0054] In step S63, each group of corresponding tracking image directional gradient histogram semantic feature vectors and tracking image color histogram semantic feature vectors in the sequence of tracking image directional gradient histogram semantic feature vectors and the sequence of tracking image color histogram semantic feature vectors are cascaded to obtain a sequence of tracking image multi-dimensional semantic fusion representation vectors. It should be understood that the tracking image directional gradient histogram semantic feature vector mainly reflects the shape and structure information of the target vehicle, while the tracking image color histogram semantic feature vector focuses on reflecting the color distribution characteristics of the target vehicle. These two features are highly complementary, and using one of the features alone may not be able to fully describe the characteristics of the target vehicle. Based on this, in the technical solution of the present application, each group of corresponding tracking image directional gradient histogram semantic feature vectors and tracking image color histogram semantic feature vectors in the sequence of tracking image directional gradient histogram semantic feature vectors and the sequence of tracking image color histogram semantic feature vectors are cascaded to integrate the different modal information they carry together to form a more comprehensive and richer feature representation, and obtain a sequence of tracking image multi-dimensional semantic fusion representation vectors.

[0055] In step S64, perform causal modeling of the sequence of the multi-dimensional semantic fusion representation vectors of the tracking image, namely, dynamic walk of the tracking image, to obtain the tracking image context dynamic walk semantic coding vector. Specifically, Figure 4 FIG. is a flowchart of step S64 in the object detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to an embodiment of the present application. As Figure 4 shown, the step S64 includes: S641, perform implicit feature mining on each multi-dimensional semantic fusion representation vector in the sequence of the multi-dimensional semantic fusion representation vectors of the tracking image to obtain a sequence of multi-dimensional semantic depth implicit feature coding vectors of the tracking image; S642, calculate the causal association topological features of the sequence of the multi-dimensional semantic depth implicit feature coding vectors of the tracking image to obtain a multi-dimensional semantic causal association topological feature matrix of the tracking image; S643, based on the multi-dimensional semantic causal association topological feature matrix of the tracking image, perform surface layer and hidden layer dynamic walk coding on the sequence of the multi-dimensional semantic fusion representation vectors of the tracking image and the sequence of the multi-dimensional semantic depth implicit feature coding vectors of the tracking image respectively to obtain a multi-dimensional semantic surface layer context dynamic walk coding vector and a multi-dimensional semantic hidden layer context dynamic walk coding vector of the tracking image; S644, fuse the multi-dimensional semantic surface layer context dynamic walk coding vector and the multi-dimensional semantic hidden layer context dynamic walk coding vector of the tracking image to obtain the tracking image context dynamic walk semantic coding vector.

[0056] It should be understood that there are dynamic relationships among the multi-dimensional semantic fusion representation vectors of each tracking image. For example, how the features such as the shape and color of the target vehicle at different times affect and change each other over time. During the driving process of the target vehicle, its attitude change may cause the shape feature to change, and at the same time, the change of the illumination condition may also affect the color feature. In order to capture these complex dynamic connections, the present application performs causal modeling of the sequence of the multi-dimensional semantic fusion representation vectors of the tracking image, namely, dynamic walk of the tracking image, to obtain the tracking image context dynamic walk semantic coding vector. In this way, the internal relationships between complex information such as target features in the image and key feature points can be more effectively mined, so that not only the local features of the target are focused on, but also the hidden global dependencies are fully captured, and at the same time, the interference of noise is reduced, thereby helping to better understand the state and changes of the target.

[0057] Specifically, first perform implicit feature mining on each multi-dimensional semantic fusion representation vector in the sequence of the multi-dimensional semantic fusion representation vectors of the tracking image to obtain a sequence of multi-dimensional semantic depth implicit feature coding vectors of the tracking image. The above process can be expressed by the formula:

[0058] O = {x 1 ,x2 ,..., x i ,..., x n}

[0059] v i = Sigmoid[Conv 1×1 (x i )]

[0060] D = {v 1 , v 2 ,..., v i ,..., v n}

[0061] Among them, O is the sequence of the multi-dimensional semantic fusion representation vectors of the tracking image, x 1 , x 2 , x i and x n are respectively the 1st, 2nd, ith, and nth multi-dimensional semantic fusion representation vectors of the tracking image in the sequence of the multi-dimensional semantic fusion representation vectors of the tracking image, Conv 1×1 is point convolution encoding, Sigmoid is the activation function of convolution encoding, v 1 , v 2 , v i , v j and v n are respectively the 1st, 2nd, ith, jth, and nth multi-dimensional semantic depth implicit feature encoding vectors of the tracking image in the sequence of the multi-dimensional semantic depth implicit feature encoding vectors of the tracking image, and D is the sequence of the multi-dimensional semantic depth implicit feature encoding vectors of the tracking image.

[0062] It should be understood that in the actual target vehicle detection and tracking scenario, the collected image data will inevitably be affected by various noises, such as the errors of the shooting equipment, the instability of the environmental illumination, the signal interference during the image transmission, etc. These noises will be reflected in the multi-dimensional semantic fusion representation vectors of the tracking image, affecting the accurate analysis of the target features in the subsequent stage. Through implicit feature mining, the influence of these noises can be ignored to a certain extent, making the generated multi-dimensional semantic depth implicit feature encoding vectors of the tracking image more pure and reliable. Through implicit feature mining, deep semantic features can also be extracted from the sequence of the input multi-dimensional semantic fusion representation vectors of the tracking image. For example, taking the posture of the target vehicle as an example, the features in the original sequence may only describe the general outline of the vehicle, while implicit feature mining can more accurately capture the details and dynamic changes of the vehicle posture by analyzing the change relationship of the features of each part of the vehicle under different postures, and thus can provide richer and more accurate feature information for the detection and tracking of the target vehicle.

[0063] Specifically, Figure 5It is a flowchart of step S642 in the object detection and tracking method that combines the YOLOv8 detection algorithm and the KCF tracking algorithm according to an embodiment of the present application. As Figure 5 shown, the step S642 includes: S6421, calculating the semantic causal association factor between any two tracking image multi-dimensional semantic depth implicit feature encoding vectors in the sequence of the tracking image multi-dimensional semantic depth implicit feature encoding vectors to obtain a tracking image multi-dimensional semantic causal association topology matrix composed of multiple tracking image multi-dimensional semantic causal association factors; S6422, triggering a causal gating activation function for the tracking image multi-dimensional semantic causal association topology matrix to obtain the tracking image multi-dimensional semantic causal association topology feature matrix.

[0064] More specifically, in an embodiment of the present application, the S6421 includes: calculating the association matrix between any two tracking image multi-dimensional semantic depth implicit feature encoding vectors in the sequence of the tracking image multi-dimensional semantic depth implicit feature encoding vectors to obtain a sequence of tracking image multi-dimensional semantic association matrices; calculating the semantic causal association factor of each tracking image multi-dimensional semantic association matrix in the sequence of the tracking image multi-dimensional semantic association matrices to obtain the tracking image multi-dimensional semantic causal association topology matrix composed of multiple tracking image multi-dimensional semantic causal association factors, and the tracking image multi-dimensional semantic causal association factor is related to the mean, variance, maximum value and tracking image causal association bias value of its corresponding tracking image multi-dimensional semantic association matrix; wherein, in response to the variance of the tracking image multi-dimensional semantic association matrix being greater than or equal to a predetermined threshold, the weighted average of the distances between any two tracking image multi-dimensional semantic depth implicit feature encoding vectors in the sequence of the tracking image multi-dimensional semantic depth implicit feature encoding vectors is used as the tracking image causal association bias value; in response to the variance of the tracking image multi-dimensional semantic association matrix being less than the predetermined threshold, the weighted mean of the tracking image multi-dimensional semantic association matrix is used as the tracking image causal association bias value. The above process can be expressed by the formula:

[0065]

[0066] where, v i , v j are respectively the i-th and j-th tracking image multi-dimensional semantic depth implicit feature encoding vectors in the sequence of the tracking image multi-dimensional semantic depth implicit feature encoding vectors, is matrix multiplication, v j T is the transposed vector of v j , M i-j is the tracking image multi-dimensional semantic association matrix between v i and v j , σ 2 (M i-j) is M i-j is the variance of, max(M i-j ) is to take the maximum value in M i-j , μ(M i-j ) is the mean of M i-j , λ is the tracking image causal association bias value, t i-j is the tracking image multi-dimensional semantic causal association factor corresponding to M i-j , d(v i , v j ) is the distance between v i and v j , L is the number of vectors in D, ε is a predetermined threshold, α and β are weighted hyperparameters, t 1-1 , t 1-n , t n-1 and t n-n are respectively the tracking image multi-dimensional semantic causal association factors at each position in the tracking image multi-dimensional semantic causal association topology matrix, and T is the tracking image multi-dimensional semantic causal association topology matrix.

[0067] It should be understood that in the actual scenario of object detection and tracking, the tracking image multi-dimensional semantic depth implicit feature encoding vector contains rich feature information of the target vehicle, and the relationships of these information are intricate. Calculating the semantic causal association factors between any two encoding vectors can uncover the hidden causal connections from the implicit feature sequence. For example, the color feature of a vehicle may change due to changes in lighting conditions, and the lighting change may be related to the environmental position where the vehicle is located. This deep causal association is revealed through this step, which helps to more comprehensively and deeply understand the change mechanism of the target vehicle features. Here, the tracking image multi-dimensional semantic causal association topology matrix composed of multiple tracking image multi-dimensional semantic causal association factors can present the global causal association structure among all the features of the target vehicle. By analyzing this matrix, complex patterns and rules in the feature data can be discovered, which helps to grasp the feature change law of the target vehicle as a whole.

[0068] In particular, by treating the low-level causal associations in a complex system as molecular-level relationships inferred based on statistical correlations, it is possible to further perform intervention prediction on the causal association energy of the tracking image based on the global fine-grained statistical association representation, so as to study the fine-grained structure and its dynamic regulation of the tracking image causal association based on the high-dimensional and heterogeneous representation of tracking image causal relationship omics. Among them, when the aggregative distribution representation of the causal graph is greater than a predetermined threshold, a bias occurs during the integration of source data based on the matrix-based graph node effect representation of the semantic causal association factor, while when the aggregative distribution representation of the causal graph is less than the predetermined threshold, a condensed structure modeling can be directly performed through feature pattern integration compression. In this way, not only can the causal association energy of the tracking image in the system be encoded and described, but also the implicit causal intervention prediction results can be condensed, thereby obtaining a more efficient revelation of the key tracking image causal associations.

[0069] Next, the causal gating activation function is triggered for the multi-dimensional semantic causal association topology matrix of the tracking image to obtain the multi-dimensional semantic causal association topology feature matrix of the tracking image. The above process can be expressed by the formula:

[0070]

[0071] where t i-j is the multi-dimensional semantic causal association factor corresponding to M i-j , T is the multi-dimensional semantic causal association topology matrix of the tracking image, softmax is a non-linear activation function, τ is a normalization threshold, and f trigger (T) represents the gating activation process for T, and M is the multi-dimensional semantic causal association topology feature matrix of the tracking image.

[0072] It should be understood that by triggering the causal gating activation function for the multi-dimensional semantic causal association topology matrix of the tracking image, more detailed feature extraction and modeling can be performed on the causal topology relationship in the original matrix. Specifically, its core operations are the dynamic gating mechanism and the non-linear activation function. The gating mechanism allows the model to identify key causal paths in a dynamic context, thereby strengthening important associations and weakening noise interference; while the activation function enhances the model's expressive power through introducing non-linearity and captures the high-order regularities hidden in complex causal structures. For example, as time goes by, the input tracking image data changes continuously, and the features and states of the target vehicle may change. The process of triggering the causal gating activation function can respond in a timely manner to these dynamic changes in the data and dynamically adjust the causal associations. For instance, when the target vehicle is occluded, accelerating, or decelerating, etc., the model can re-evaluate and adjust the causal relationships between features through the gating mechanism and the activation function to ensure accurate tracking and detection of the target vehicle and adapt to the dynamic changes in the data.

[0073] Specifically, in the embodiment of the present application, the step S643 includes: performing graph convolutional feature sequence dynamic walk encoding on the tracking image multi-dimensional semantic causal association topological feature matrix and the sequence of the tracking image multi-dimensional semantic fusion representation vectors to obtain the tracking image multi-dimensional semantic surface context dynamic walk encoding vector, and this process can be expressed by the formula:

[0074]

[0075] where x i is the i-th tracking image multi-dimensional semantic fusion representation vector in the sequence of the tracking image multi-dimensional semantic fusion representation vectors, M is the tracking image multi-dimensional semantic causal association topological feature matrix, GCN is graph convolutional processing, and H surface is the tracking image multi-dimensional semantic surface context dynamic walk encoding vector;

[0076] Performing the graph convolutional feature sequence dynamic walk encoding on the tracking image multi-dimensional semantic causal association topological feature matrix and the sequence of the tracking image multi-dimensional semantic depth implicit feature encoding vectors to obtain the tracking image multi-dimensional semantic hidden layer context dynamic walk encoding vector, and this process can be expressed by the formula:

[0077]

[0078] where v i are respectively the i-th tracking image multi-dimensional semantic depth implicit feature encoding vector in the sequence of the tracking image multi-dimensional semantic depth implicit feature encoding vectors, M is the tracking image multi-dimensional semantic causal association topological feature matrix, GCM is graph convolutional processing, GCN is graph convolutional processing, and H hidden is the tracking image multi-dimensional semantic hidden layer context dynamic walk encoding vector.

[0079] It should be understood that the tracking image multi-dimensional semantic causal association topological feature matrix shows the complex causal association structure among the target vehicle features. To generate the surface context semantic representation, the dynamic random walk mechanism simulates the propagation process of features in this topological structure. Just like information is transmitted in a network, features start from local nodes and gradually spread globally. For example, for the appearance features of the target vehicle, it starts from the features of a local component of the vehicle (such as the headlight), and through the association relationships in the topological structure, gradually spreads to the features of other components related to the headlight (such as the body color, overall shape, etc.), achieving feature aggregation from local to global. That is to say, during the feature propagation process, the dynamic random walk mechanism recursively aggregates the explicit semantics between nodes. This means it continuously integrates the explicit semantic information carried by different feature nodes. In this way, the explicit semantic information scattered in different nodes can be effectively aggregated to form a more comprehensive and representative surface context semantic representation, namely the tracking image multi-dimensional semantic surface context dynamic random walk encoding vector. This vector can reflect some obvious features of the target vehicle in the current state and their relationships as a whole, providing a basis for subsequent analysis. For the hidden layer context dynamic semantic encoding, the dynamic random walk mechanism conducts a deeper semantic exploration on the tracking image multi-dimensional semantic deep implicit feature encoding vector. During the exploration process of the hidden layer semantics, due to the enabling of the latent embedding of features, attention needs to be paid to the multi-hop propagation of high-order information and the distributed decoupling of deep features. Multi-hop propagation means that feature information can not only be transmitted between adjacent nodes, but also be transmitted over a long distance through multiple intermediate nodes, thereby capturing more profound temporal dependence relationships. For example, the change in the driving trajectory of a vehicle over a period of time may involve the feature information at multiple time points, and these information can be effectively integrated through multi-hop propagation. At the same time, the distributed decoupling of deep features can avoid the over-smoothing phenomenon caused by the propagation of deep topological features, that is, prevent features from becoming too similar during the propagation process and losing their uniqueness. In this way, a more discriminative and representative hidden layer context dynamic semantic encoding can be obtained, namely the tracking image multi-dimensional semantic hidden layer context dynamic random walk encoding vector, which can provide greater generalization ability for feature expression.

[0080] Finally, fuse the tracking image multi-dimensional semantic surface context dynamic random walk encoding vector and the tracking image multi-dimensional semantic hidden layer context dynamic random walk encoding vector to obtain the tracking image context dynamic random walk semantic encoding vector. The above process can be expressed by the formula:

[0081] H final =γ·H surface +(1 - γ)·H hidden

[0082] where, H surfaceis the dynamic random-walking encoded vector of the multi-dimensional semantic surface context of the tracking image, H hidden is the dynamic random-walking encoded vector of the multi-dimensional semantic hidden context of the tracking image, and γ is the fusion weighting parameter, H final is the dynamic random-walking semantic encoded vector of the context of the tracking image.

[0083] It should be understood that the dynamic random-walking encoded vector of the multi-dimensional semantic surface context of the tracking image captures the explicit semantic information of the target vehicle, such as the appearance features of the vehicle (color, shape, etc.); while the dynamic random-walking encoded vector of the multi-dimensional semantic hidden context of the tracking image mines the implicit semantic information, such as the driving intention of the vehicle, potential behavior patterns, and more abstract feature associations, etc. These two types of semantic information are complementary, and by fusing these two vectors, the explicit and implicit semantic information can be organically combined. That is, the dynamic random-walking semantic encoded vector of the context of the tracking image obtained by fusion integrates the semantic information of the surface layer and the hidden layer, providing a more comprehensive feature description of the target vehicle. It is no longer limited to a single explicit or implicit feature, but covers multiple aspects of the target. This comprehensive feature description helps to more accurately identify and track the target vehicle. In the specific implementation of this application, a weighted summation strategy is adopted to fuse these two vectors, which enables the dynamic adjustment of the importance distribution of the surface layer and hidden layer features in the fusion process. This ability to dynamically adjust the feature importance can make the fused feature vector better adapt to different situations and task requirements.

[0084] In step S65, based on the dynamic random-walking semantic encoded vector of the context of the tracking image, the object detection image of the target vehicle in the t-th frame is obtained. Specifically, in the embodiment of this application, step S65 includes: passing the dynamic random-walking semantic encoded vector of the context of the tracking image through an object prediction generator of the target vehicle based on AIGC to obtain the object detection image of the target vehicle in the t-th frame. That is, using the sequence of the multi-dimensional semantic fusion representation vectors of the tracking image to perform causal dynamic random-walking to obtain the dynamic random-walking semantic encoded vector of the context of the tracking image for generation processing, so as to automatically predict and obtain the object detection image of the target vehicle in the t-th frame. It is worth mentioning that the object prediction generator of the target vehicle based on AI GC (Artificial Intelligence Generated Content) technology has powerful generation capabilities. It can generate image content that conforms to logic and expectations based on the input dynamic random-walking semantic encoded vector of the context of the tracking image. In the object detection and tracking task, when the target vehicle is not detected in the t-th frame, this generation ability can be used to fill the information gap. By understanding and analyzing the information of the previous frames, it can generate a possible object detection image of the target vehicle in the t-th frame. In this way, by predicting the image to replace the actual detection image, the continuity of target tracking can be maintained.

[0085] The following is a detailed elaboration of a specific implementation process for "generating the object detection image of the target vehicle in the t-th frame by passing the dynamic roaming semantic encoding vector of the tracking image context through an AI GC-based target vehicle object predictor":

[0086] First, in the preparation stage of the predictor, it is necessary to select a suitable AI GC model according to the requirements of the actual task and the characteristics of the data at hand. Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), diffusion models, etc. are all common choices. GANs have the advantage of generating high-resolution and realistic images. Through the adversarial training between the generator and the discriminator, the generator can continuously optimize the generated images to make them closer to real vehicle images. VAEs are good at learning the latent distribution of data. By mapping the encoding vector to the latent space and sampling, and then using the decoder to generate images, they can produce diverse image samples. Diffusion models generate images by gradually removing noise and are outstanding in generating complex and high-quality images. After selecting the model, it is necessary to pre-train it with a large-scale and diverse vehicle image dataset. These datasets should contain vehicle images at different angles and poses in various scenarios. By learning these data, the model can deeply understand the features, appearance patterns, and structural information of vehicles, laying a solid foundation for generating accurate images based on specific encoding vectors in the future.

[0087] When the predictor is ready, input the dynamic roaming semantic encoding vector of the tracking image context into it. This encoding vector contains rich information about the target vehicle in the previous t-1 frame images and is the key basis for generating the t-th frame image. After receiving the encoding vector, the model will deeply analyze the vector using its internal neural network structure. For example, the convolutional layer is responsible for identifying local features in the vector, such as the contours and detailed components of the vehicle; the fully connected layer further integrates these local features, explores the hidden associations and potential patterns between them, and provides strong support for subsequent image generation.

[0088] Next, we enter the core part of image generation. Taking GAN as an example, the generator starts to attempt to generate the detection image of the target vehicle for the t-th frame based on the parsed feature information. In the initial stage, the generated image may have a large gap from the real vehicle image, but the discriminator will carefully distinguish between the generated image and the real image and feedback the judgment result to the generator. The generator continuously adjusts its own parameters according to these feedbacks to optimize the generated image. In the process of multiple iterations, the image generated by the generator will get closer and closer to the real vehicle image until it can deceive the discriminator. If the VAE model is used, it maps the encoded vector to the latent space, and each point in the latent space represents a unique combination of vehicle image features. The model samples in the latent space and uses the decoder to convert the sampled points into images. In this process, VAE can learn the distribution law of the data, so as to generate the t-th frame image that not only conforms to the current context information but also has a certain degree of diversity. The diffusion model starts from a random noise image and gradually predicts and removes the noise according to the information in the encoded vector. In each iteration, the model updates the image according to the previous image state and the encoded vector. After multiple iterations, the image gradually becomes clear and presents the accurate features of the target vehicle, and finally generates the detection image of the target vehicle for the t-th frame.

[0089] After the image is generated, optimization and post-processing are still needed. Since the generated image may have some defects or does not meet the requirements of practical applications in some aspects, it needs to be optimized. This can be achieved by adjusting the hyperparameters of the model, such as the learning rate, the number of iterations, etc., or by increasing the training data. For example, in GAN, reasonably adjusting the training ratio of the generator and the discriminator can make the generated image more realistic. At the same time, a series of post-processing operations are also required, including image filtering to remove possible noise, enhancing the contrast to make the image details clearer, adjusting the brightness to adapt to different visual needs, etc. In addition, according to the requirements of subsequent object detection algorithms, operations such as cropping and scaling the image may also be needed to ensure that the size and format of the image are compatible with the detection algorithm.

[0090] Finally, the generated detection image of the target vehicle for the t-th frame needs to be evaluated and the results need to be fed back. Evaluation metrics such as peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) are used to measure the similarity and quality difference between the generated image and the real image. If the quality of the generated image fails to meet the expected standard, the evaluation result needs to be fed back to the prediction generator to further train and adjust the model. By continuously optimizing the model, the quality of the generated image is improved to meet the accuracy and reliability requirements of the detection image of the target vehicle in practical applications.

[0091] In summary, step S6 is clearly described. It uses image analysis and feature extraction techniques based on machine vision to perform HOG feature extraction and color histogram calculation on each of the previous t-1 tracking images. Then, semantic capture and multi-dimensional semantic fusion are performed on the extracted histogram of oriented gradients of each tracking image and the color histogram of each tracking image, so as to automatically predict and obtain the target vehicle object detection image of the t-th frame according to the dynamic causal context walk representation between the multi-dimensional semantic fusion representation features of each tracking image after fusion. In this way, through HOG feature extraction and color histogram calculation, the shape and color information of the target vehicle can be comprehensively described from different angles. At the same time, through the dynamic causal context walk representation, the dynamic changes of the target vehicle over time and its interaction with the environment are captured, thereby improving the understanding and prediction accuracy of the target vehicle's behavior.

[0092] In summary, the object detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm based on the embodiments of the present application is elucidated. It first obtains the video stream time series set of the tracking images of the target vehicle object, and uses YOLOv8 to perform object detection on the first frame to initialize the KCF tracker. Subsequently, the KCF tracker processes the subsequent frame images for object tracking. If tracking fails at the t-th frame, YOLOv8 is re-enabled to detect this frame in an attempt to retrieve the target. If the target is successfully detected, a new target vehicle object detection image is updated and generated, and the KCF tracker is re-initialized. If the target cannot be detected, the target vehicle object detection image of the t-th frame is predicted and generated based on the information of the previous t-1 frames. In this way, the continuity and accuracy of the target vehicle object tracking can be effectively ensured.

Claims

1. A target detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm, characterized in that: include: Acquire a video stream time sequence set of tracking images of a target vehicle object; Using the YOLOv8 detection algorithm, target detection is performed on the first frame tracking image in the video stream time series set of the tracking image to obtain a target vehicle object detection image; Initializing a KCF tracker based on the target vehicle object detection image, and using the initialized KCF tracker to perform target tracking on other frame tracking images in the video stream time series set of the tracking image; In response to the KCF tracker reporting that target tracking fails in the t-th frame tracking image in the video stream time sequence set of the tracking image, re-performing target detection on the t-th frame tracking image using a YOLOv8 detection algorithm to obtain a target detection result; In response to the target detection result that the target vehicle object is re-detected, generating a t-th frame of target vehicle object detection image, and reinitializing the KCF tracker based on the t-th frame of target vehicle object detection image; In response to the target detection result that the target vehicle object is not detected, target vehicle object prediction is performed based on t-1 tracking images before the t-th frame tracking image to generate a t-th frame target vehicle object detection image.

2. The target detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to claim 1, characterized in that: In response to the target detection result indicating that the target vehicle object is not detected, performing target vehicle object prediction based on t-1 tracking images before the t-th frame tracking image to generate a t-th frame target vehicle object detection image, including: Performing HOG feature extraction and color histogram calculation on each tracking image in the t-1 tracking images to obtain a sequence of tracking image directional gradient histograms and a sequence of tracking image color histograms; Performing histogram feature extraction on the sequence of tracking image directional gradient histograms and the sequence of tracking image color histograms to obtain a sequence of tracking image directional gradient histogram semantic feature vectors and a sequence of tracking image color histogram semantic feature vectors; Cascading each group of corresponding tracking image directional gradient histogram semantic feature vectors and tracking image color histogram semantic feature vectors in the sequence of tracking image directional gradient histogram semantic feature vectors and the sequence of tracking image color histogram semantic feature vectors to obtain a sequence of tracking image multi-dimensional semantic fusion representation vectors; Performing causal modeling of the tracking image dynamic walk on the sequence of the tracking image multi-dimensional semantic fusion representation vector to obtain a tracking image context dynamic walk semantic encoding vector; Based on the tracking image context dynamic walk semantic coding vector, the target vehicle object detection image of the t-th frame is obtained.

3. The target detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to claim 2, characterized in that: Performing HOG feature extraction and color histogram calculation on each of the t-1 tracking images to obtain a sequence of tracking image directional gradient histograms and a sequence of tracking image color histograms, including: Performing HOG feature extraction on each tracking image in the t-1 tracking images to obtain a sequence of directional gradient histograms of the tracking images; The color histogram of each tracking image in the t-1 tracking images is calculated respectively to obtain a sequence of the tracking image color histograms.

4. The target detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to claim 3, characterized in that: Histogram feature extraction is performed on the sequence of tracking image directional gradient histograms and the sequence of tracking image color histograms to obtain a sequence of tracking image directional gradient histogram semantic feature vectors and a sequence of tracking image color histogram semantic feature vectors, including: passing the sequence of tracking image directional gradient histograms and the sequence of tracking image color histograms through a histogram feature extractor based on an FCN space model to obtain a sequence of tracking image directional gradient histogram semantic feature vectors and a sequence of tracking image color histogram semantic feature vectors.

5. The target detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to claim 4, characterized in that: The tracking image context dynamic walking semantic encoding vector of the tracking image is obtained by performing causal modeling on the sequence of the tracking image multi-dimensional semantic fusion representation vector, including: Performing implicit feature mining on each tracking image multi-dimensional semantic fusion representation vector in the sequence of tracking image multi-dimensional semantic fusion representation vectors to obtain a sequence of tracking image multi-dimensional semantic depth implicit feature encoding vectors; Calculating the causal correlation topological features of the sequence of multi-dimensional semantic depth implicit feature encoding vectors of the tracking image to obtain a multi-dimensional semantic causal correlation topological feature matrix of the tracking image; Based on the tracking image multi-dimensional semantic causal association topological feature matrix, respectively performing surface layer and hidden layer dynamic walking coding on the sequence of the tracking image multi-dimensional semantic fusion representation vector and the sequence of the tracking image multi-dimensional semantic depth implicit feature coding vector to obtain the tracking image multi-dimensional semantic surface layer context dynamic walking coding vector and the tracking image multi-dimensional semantic hidden layer context dynamic walking coding vector; The tracking image multi-dimensional semantic surface context dynamic wandering coding vector and the tracking image multi-dimensional semantic hidden context dynamic wandering coding vector are fused to obtain the tracking image context dynamic wandering semantic coding vector.

6. The target detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to claim 5, characterized in that: Calculating the causal association topological features of the sequence of multi-dimensional semantic depth implicit feature encoding vectors of the tracking image to obtain a multi-dimensional semantic causal association topological feature matrix of the tracking image includes: Calculating the semantic causal association factor between any two tracking image multi-dimensional semantic depth implicit feature coding vectors in the sequence of tracking image multi-dimensional semantic depth implicit feature coding vectors to obtain a tracking image multi-dimensional semantic causal association topological matrix composed of multiple tracking image multi-dimensional semantic causal association factors; A causal gating activation function is triggered on the multi-dimensional semantic causal association topological matrix of the tracking image to obtain the multi-dimensional semantic causal association topological feature matrix of the tracking image.

7. The target detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to claim 6, characterized in that: Calculating the semantic causal association factor between any two tracking image multi-dimensional semantic depth implicit feature coding vectors in the sequence of the tracking image multi-dimensional semantic depth implicit feature coding vectors to obtain a tracking image multi-dimensional semantic causal association topological matrix composed of a plurality of tracking image multi-dimensional semantic causal association factors, including: Calculating the association matrix between any two tracking image multi-dimensional semantic depth implicit feature coding vectors in the sequence of tracking image multi-dimensional semantic depth implicit feature coding vectors to obtain a sequence of tracking image multi-dimensional semantic association matrices; Calculating the semantic causal association factor of each tracking image multidimensional semantic association matrix in the sequence of the tracking image multidimensional semantic association matrix to obtain the tracking image multidimensional semantic causal association topological matrix composed of a plurality of tracking image multidimensional semantic causal association factors, wherein the tracking image multidimensional semantic causal association factor is related to the mean, variance, maximum value and tracking image causal association bias value of the corresponding tracking image multidimensional semantic association matrix; In which, in response to the variance of the tracking image multidimensional semantic association matrix being greater than or equal to a predetermined threshold, a weighted average of the distances between any two tracking image multidimensional semantic depth implicit feature coding vectors in the sequence of the tracking image multidimensional semantic depth implicit feature coding vectors is used as the tracking image causal association bias value; In response to the variance of the tracking image multidimensional semantic association matrix being smaller than the predetermined threshold, a weighted mean of the tracking image multidimensional semantic association matrix is ​​used as the tracking image causal association bias value.

8. The target detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to claim 7, characterized in that: Based on the tracking image multi-dimensional semantic causal association topological feature matrix, the sequence of the tracking image multi-dimensional semantic fusion representation vector and the sequence of the tracking image multi-dimensional semantic depth implicit feature coding vector are respectively subjected to surface layer and hidden layer dynamic walking coding to obtain the tracking image multi-dimensional semantic surface layer context dynamic walking coding vector and the tracking image multi-dimensional semantic hidden layer context dynamic walking coding vector, including: Performing graph convolution feature sequence dynamic walking coding on the tracking image multi-dimensional semantic causal association topological feature matrix and the tracking image multi-dimensional semantic fusion representation vector sequence to obtain the tracking image multi-dimensional semantic surface context dynamic walking coding vector; The graph convolution feature sequence dynamic walking coding is performed on the sequence of the tracking image multi-dimensional semantic causal association topological feature matrix and the tracking image multi-dimensional semantic depth implicit feature coding vector to obtain the tracking image multi-dimensional semantic hidden layer context dynamic walking coding vector.

9. The target detection and tracking method combining the YOLOv8 detection algorithm and the KCF tracking algorithm according to claim 8, characterized in that: Based on the tracking image context dynamic wandering semantic coding vector, the t-th frame target vehicle object detection image is obtained, including: passing the tracking image context dynamic wandering semantic coding vector through an AIGC-based target vehicle object prediction generator to obtain the t-th frame target vehicle object detection image.

Citation Information

Patent Citations

  • Video semi-automatic target labeling method integrating target detection and tracking

    CN110929560A

  • Image tracking method in audio and video control equipment

    CN113920168A

  • Power grid safety operation monitoring image processing method and intelligent video monitoring system

    CN116109975A

  • Target tracking method and device based on feature fusion and loss judgment mechanism

    CN116664628A

  • Dynamic link prediction method and system based on time sequence heterogeneous graph attention network

    CN118916786A

Cited By

  • Intelligent identification system and method for working stage of mining excavator

    CN120119697A