Object-oriented feature point tracking method
Through the object-oriented feature point tracking method, combined with YOLOv5 and SURF/KCF algorithms, the problem of poor results in complex scenarios and target deformation is solved, and the target is achieved quickly and accurately tracked and efficiently calculated.
Patent Information
- Application Number
- CN202411901960.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional goal tracking methods are not effective when dealing with complex scenarios and target deformation, and have high computational complexity and are difficult to meet real-time requirements.
The object-oriented feature point tracking method is adopted to locate feature points through the object detection algorithm, and the feature point tracking algorithm is used for stable tracking. Combined with the YOLOv5 object detection algorithm and the SURF/KCF feature point detection tracking algorithm, efficient and accurate tracking of the target is achieved.
Fast and accurate tracking of the target is achieved, the robustness in complex scenarios and target deformation is improved, the calculation complexity is reduced, and the real-time requirement is met.
Smart Images

Figure CN120013989A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of video processing and target tracking, in particular to an object-oriented feature point tracking method. Background Art
[0002] With the rapid development of computer vision and image processing technology, target tracking has become a key research focus in many fields such as video surveillance, human-computer interaction, and robot navigation. Its core task is to accurately identify and lock the position of the same target in consecutive image frames, and then depict its motion trajectory. However, factors such as noise interference in the image, changes in lighting conditions, and deformation of the target itself make target tracking a very challenging task.
[0003] Traditional target tracking methods mostly rely on filtering technology or machine learning algorithms, such as Kalman filtering, particle filtering, and support vector machines. These methods often need to extract a large amount of feature information from the image first, and then build a target model based on these features and match it. However, these methods have high computational complexity during the processing process and are difficult to meet real-time requirements. In addition, when dealing with complex scenes and target deformation, their tracking effects are often unsatisfactory.
[0004] In recent years, object-oriented feature point tracking methods have gradually emerged and received widespread attention in the industry. This method accurately locates and extracts the key feature points of the target and continuously tracks these feature points to achieve stable tracking of the entire target. Compared with traditional target tracking methods, object-oriented feature point tracking methods have significant advantages in computational efficiency, real-time performance, and robustness to target deformation and illumination changes. Summary of the invention
[0005] In view of the above-mentioned problems existing in the prior art, an embodiment of the present invention provides an object-oriented feature point tracking method, which makes full use of the positioning accuracy of the target detection algorithm and combines the stability of feature point tracking to achieve more efficient and accurate target tracking, and provide strong technical support for applications in related fields.
[0006] An embodiment of the present invention provides an object-oriented feature point tracking method, comprising:
[0007] Perform target detection on the video frame to obtain the target detection frame;
[0008] Extract feature points in the detection frame of target detection and track the feature points;
[0009] Calculate the target motion trajectory.
[0010] In some embodiments of the present invention, before performing target detection on a video frame, the method includes:
[0011] Perform video frame preparation and preprocessing, using the video processing library to split the video file into individual frames as input data for the subsequent object detection model;
[0012] In the preparation stage, the object categories and bounding boxes in the video frames are annotated to build an object detection dataset;
[0013] During preprocessing, each frame is resized to fit the model’s input requirements;
[0014] Normalization is performed to scale pixel values to a specific range to eliminate differences between images due to lighting and contrast factors;
[0015] Generate more training samples by using data augmentation techniques that at least include randomly flipping, rotating, cropping, and scaling the images.
[0016] In some embodiments of the present invention, the performing target detection on the video frame includes:
[0017] The real-time target detection algorithm YOLOv5 is used for model training. YOLOv5 is a target detection algorithm based on deep learning. YOLOv5 is used to detect targets in video frames to obtain the target detection frame.
[0018] During the training process, the preprocessed video frames are first loaded as input data; then, the input data is passed to the target detection network through forward propagation. The target detection network includes a feature extraction network, a feature pyramid network, and a detection network; among them, the feature extraction network is responsible for extracting features from the input image, the feature pyramid network is used to fuse feature information of different scales, and the detection network is responsible for predicting the position, size, and category probability of the target bounding box.
[0019] In some embodiments of the present invention, the steps of using the real-time target detection algorithm YOLOv5 for model training are as follows:
[0020] Forward propagation: Input the video frame with the target bounding box annotation into the target detection network, and obtain the predicted position of the target box through the feature extraction network, feature pyramid network, and detection network;
[0021] Loss calculation: It includes three parts: bounding box regression loss, target confidence loss and classification loss. The bounding box regression loss is used to measure the position difference between the predicted bounding box and the actual bounding box, the target confidence loss is used to measure the difference between the probability of the predicted target and the actual annotation, and the classification loss is used to measure the difference between the predicted category and the actual category. CIOU loss is used as the calculation formula for bounding box regression loss. CIOU loss considers at least the aspect ratio and center point distance of the bounding box on the basis of IOU loss. The confidence loss and classification loss are calculated using cross entropy loss respectively.
[0022] The bounding box regression loss uses CIOU loss, and the calculation formula is as follows:
[0023]
[0024] Among them, B represents the bounding box and gt represents the true value;
[0025] The confidence loss uses cross entropy loss, and the calculation formula is as follows:
[0026]
[0027] P i =Sigmoid(IoU)
[0028] Among them, P represents the confidence level;
[0029] The classification loss also uses cross entropy loss, and the calculation formula is as follows:
[0030]
[0031] C i =Sigmoid(c)
[0032] Among them, X represents the category score;
[0033] Back propagation and optimization: The gradient of the loss function to the model parameters is calculated through the back propagation algorithm, and the optimizer SGD is used to update the parameters to minimize the loss function;
[0034] Iterative training: Repeat the process of forward propagation, loss calculation, and backpropagation until the model reaches the preset number of training rounds.
[0035] In some embodiments of the present invention, after completing the target detection model training, the method deploys the model into the video frame processing flow, and the method further includes:
[0036] Use the non-maximum suppression algorithm to eliminate the redundant bounding boxes generated by YOLOv5 when training the model;
[0037] Filter out low-confidence detection results according to the confidence threshold, retain bounding boxes with confidence higher than the set threshold, draw the retained bounding boxes on the original frames, including the coordinates, size and category labels of the bounding boxes, and save the frames with detection results as image sequences or reassemble them into video files.
[0038] In some embodiments of the present invention, extracting feature points in a detection frame of target detection includes:
[0039] After obtaining the target bounding box of each frame, a feature point detection algorithm is used to extract the feature points inside the target bounding box; specifically,
[0040] The SURF algorithm is used to generate the descriptor of the feature points by calculating the gradient and direction information of the local area of the image;
[0041] For the SURF algorithm, the descriptor of the feature point is calculated based on the determinant value of the Hessian matrix. The specific formula is as follows:
[0042] Calculating Hessia n matrix:
[0043]
[0044] Among them, L xx , L xy and L yy is the Gaussian second-order derivative of the image at point x;
[0045] Compute the determinant of the Hessian matrix:
[0046] det(H)=L xx (x,σ)·L yy (x,σ)-L xy (x,σ) 2
[0047] When det(H) is greater than a certain threshold, the point is considered to be a feature point;
[0048] Assign a direction to each feature point to ensure the rotation invariance of the descriptor;
[0049] Select a neighborhood around the feature point, and calculate the gradient direction and gradient magnitude of the pixels in the neighborhood to generate the descriptor of the feature point;
[0050] In order to ensure the rotation invariance of the descriptor, the SURF algorithm assigns a main direction to each feature point. The main direction is obtained by counting the gradient direction and gradient magnitude of the pixels in the neighborhood of the feature point.
[0051] By rotating the pixels in the neighborhood of the feature point to the main direction, the descriptor is ensured to be robust to the rotation changes of the target;
[0052] The SURF algorithm selects a neighborhood around the feature point and calculates the gradient direction and gradient magnitude of the pixels in the neighborhood to generate a descriptor of the feature point. The descriptor is a vector used to represent the local feature information of the feature point. By comparing the descriptors of feature points in different frames, the correspondence between feature points is established to achieve target tracking.
[0053] In some embodiments of the present invention, tracking the feature points includes:
[0054] During the tracking process, feature points are extracted for each frame and their descriptors are calculated;
[0055] Using the target detector in the KCF algorithm, search for the position that best matches the target descriptor in the current frame; by comparing the similarity between descriptors, establish the correspondence between feature points and determine the position of the target in the current frame;
[0056] The KCF algorithm tracking process is expressed as the following formula:
[0057] Calculate the objective function:
[0058]
[0059] Among them, α i are the coefficients of the linear regression model, k(x,z i ) is the kernel function used to calculate the target position x and the training sample z in the current frame i The similarity between
[0060] The optimal solution is found by minimizing the objective function:
[0061]
[0062] Where K is the kernel matrix, λ is the regularization parameter, I is the identity matrix, and y is the label of the training sample;
[0063] Use the coefficients found To update the model and calculate the position of the target in the next frame;
[0064] In order to update the position and motion model of feature points in real time, an iterative method is adopted. In each frame, the position of the feature points is updated according to the prediction results of the KCF algorithm, and its descriptor is recalculated; then, the updated descriptor is used to track the next frame; the position and motion model of the feature points are gradually updated to adapt to the motion changes of the target between consecutive frames.
[0065] In some embodiments of the present invention, the calculating the target motion trajectory includes:
[0066] Calculate the overall motion trajectory of the target object based on the position changes of the tracked feature points;
[0067] Assuming that the tracked feature point positions are (x1, y1), (x2, y2), ..., (xn, yn), use the least squares method to fit a straight line as the target's motion trajectory; the specific formula is as follows,
[0068] Assume that the target's trajectory can be represented by a straight line, and its equation is:
[0069] y=mx+b
[0070] Among them, m is the slope of the line, and b is the intercept of the line. The goal is to find the best m and b so that the line can best fit the tracked feature point position.
[0071] The least squares method finds the best fitting parameters by minimizing the sum of squared errors; the error e i Defined as the perpendicular distance from each feature point to the line:
[0072] e i =y i -(mx i +b)
[0073] The total error sum of squares S is expressed as:
[0074]
[0075] In order to find the best m and b, we need to find the partial derivative of S and set it equal to zero:
[0076]
[0077] Solve this set of equations to obtain the optimal solution for m and b, thereby drawing the trajectory of the target.
[0078] In some embodiments of the present invention, the object-oriented feature point tracking method further comprises:
[0079] Perform trajectory visualization and performance evaluation, visualize the calculated motion trajectory, and overlay it with the original video frame to more intuitively observe the motion of the target object;
[0080] The accuracy and stability of trajectory calculation are quantitatively evaluated using indicators including at least trajectory length, trajectory smoothness, and deviation from the true trajectory;
[0081] Among them, the trajectory length reflects the total moving distance of the target in the video sequence; the trajectory smoothness describes the continuity and change rate of the trajectory; and the deviation from the true trajectory measures the degree of difference between the calculated trajectory and the actual trajectory.
[0082] Compared with the prior art, the beneficial effect of the object-oriented feature point tracking method provided by the embodiment of the present invention is that it realizes fast and accurate tracking of the target through the object-oriented feature point tracking method, effectively solving the shortcomings of the traditional target tracking method in complex scenes and target deformation and other problems; at the same time, by combining key steps such as target detection, feature point extraction, feature point tracking and trajectory calculation, the motion trajectory extraction of the target object in the video frame is realized, which can be used in video surveillance, autonomous driving, sports analysis and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 A flowchart of an object-oriented feature point tracking method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0084] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0085] Various aspects and features of the present application are described herein with reference to the accompanying drawings.
[0086] These and other characteristics of the present application will become apparent from the following description of a preferred form of embodiment given as a non-limiting example with reference to the accompanying drawings.
[0087] It should also be understood that, although the present application has been described with reference to some specific examples, those skilled in the art will be able to realize many other equivalent forms of the present application that have the features described in the claims and are therefore within the scope of protection defined thereby.
[0088] The above and other aspects, features and advantages of the present application will become more apparent in view of the following detailed description when taken in conjunction with the accompanying drawings.
[0089] Specific embodiments of the present application are described hereinafter with reference to the accompanying drawings; however, it should be understood that the embodiments applied for are merely examples of the present application, which may be implemented in a variety of ways. Well-known and / or repeated functions and structures are not described in detail to determine the true intent based on the user's historical operations and to avoid unnecessary or redundant details that make the present application unclear. Therefore, the specific structural and functional details applied for herein are not intended to be limiting, but are merely used as the basis and representative basis for the claims to teach those skilled in the art to use the present application in a variety of ways with substantially any suitable detailed structure.
[0090] This specification may use the phrases "in one embodiment," "in another embodiment," "in yet another embodiment," or "in other embodiments," all of which may refer to one or more of the same or different embodiments according to the present application.
[0091] The embodiment of the present invention provides an object-oriented feature point tracking method, such as Figure 1 As shown, the method includes:
[0092] Perform target detection on the video frame to obtain the target detection frame;
[0093] Extract feature points in the detection frame of target detection and track the feature points;
[0094] Calculate the target motion trajectory.
[0095] To facilitate understanding of the above technical solution, the following is an explanation with reference to the accompanying drawings and specific examples. The above technical solution specifically includes the following contents:
[0096] Video frame preparation and preprocessing
[0097] First, split the video file into individual frames using a video processing library such as OpenCV. These frames will serve as input data for the subsequent object detection model. In the preparation stage, the object categories and bounding boxes in the video frames need to be manually annotated to build an object detection dataset. The annotation work usually involves using specialized annotation tools to select the objects in the video frames and assign corresponding category labels.
[0098] The preprocessing step is crucial to improving model performance. First, since the input size of the model is usually fixed (such as 640x640), each frame needs to be resized to fit the input requirements of the model. This usually involves scaling or cropping operations to ensure that the size of the frame matches the model input size.
[0099] Next, normalization is performed. Normalization is to scale pixel values to a specific range (usually between 0 and 1) to eliminate differences between different images due to factors such as lighting and contrast. Normalization helps improve the stability and generalization ability of the model.
[0100] Finally, in order to increase the diversity of the dataset and improve the generalization performance of the model, data augmentation technology is used. Data augmentation includes randomly flipping, rotating, cropping, scaling, and other operations on the image to generate more training samples. This helps the model learn more different target poses and background changes, thereby improving detection accuracy.
[0101] Video frame object detection model training
[0102] Target detection model training is one of the key steps in the present invention. We use the real-time target detection algorithm YOLOv5 for model training. YOLOv5 is a target detection algorithm based on deep learning, which is efficient and accurate.
[0103] During the training process, we first load the preprocessed video frames as input data. Then, the input data is passed to the object detection network through forward propagation. The object detection network usually consists of three parts: feature extraction network, feature pyramid network and detection network. The feature extraction network is responsible for extracting features from the input image, the feature pyramid network is used to fuse feature information of different scales, and the detection network is responsible for predicting the location, size and category probability of the object bounding box.
[0104] The detailed steps are as follows:
[0105] (1) Forward propagation: Input a video frame with annotated target bounding boxes into the target detection network, and obtain the predicted position of the target box through the feature extraction network, feature pyramid network, and detection network.
[0106] (2) Loss calculation: It includes three parts: bounding box regression loss, target confidence loss and classification loss. The bounding box regression loss is used to measure the position difference between the predicted bounding box and the actual bounding box, the target confidence loss is used to measure the difference between the probability of the predicted target existence and the actual annotation, and the classification loss is used to measure the difference between the predicted category and the actual category. In the present invention, we use CIOU loss as the calculation formula for bounding box regression loss. CIOU loss takes into account factors such as the aspect ratio and center point distance of the bounding box on the basis of IOU loss, and can more accurately measure the position difference between bounding boxes. The confidence loss and classification loss are calculated using cross entropy loss respectively.
[0107] The bounding box regression loss uses CIOU loss, and the calculation formula is as follows:
[0108]
[0109] Where B represents the bounding box and gt represents the true value. The confidence loss uses the cross entropy loss, and the calculation formula is as follows:
[0110]
[0111] P i =Sigmoid(IoU)
[0112] Where P represents the confidence. The classification loss also uses the cross entropy loss, and the calculation formula is as follows:
[0113]
[0114] C i =Sigmoid(c)
[0115] Where C represents the category score.
[0116] (3) Backpropagation and optimization: The gradient of the loss function with respect to the model parameters is calculated through the backpropagation algorithm, and the optimizer SGD is used to update the parameters to minimize the loss function.
[0117] (4) Iterative training: Repeat the process of forward propagation, loss calculation, and back propagation until the model reaches the preset number of training rounds
[0118] 3. Video frame object detection model reasoning
[0119] After completing the object detection model training, deploy the model into the video frame processing pipeline. YOLOv5 generates a large number of bounding box predictions, many of which are redundant. In order to eliminate these redundant bounding boxes, the non-maximum suppression (NMS) algorithm is used. NMS retains the best bounding box based on the confidence and IoU (intersection over union) of the bounding box.
[0120] Filter out low-confidence detections based on a confidence threshold. Only bounding boxes with confidence above the set threshold are retained. Draw the retained bounding boxes on the original frame, including the coordinates, size, and category label of the bounding box. Save the frames with detection results as an image sequence or reassemble them into a video file.
[0121] 4. Feature point extraction within the detection frame
[0122] After obtaining the target bounding box of each frame, we can use the feature point detection algorithm to extract the feature points inside the target bounding box. The present invention adopts the SURF (Speeded Up Robust Features) algorithm to generate the descriptor of the feature points by calculating the gradient, direction and other information of the local area of the image.
[0123] For the SURF algorithm, the descriptor of the feature point is calculated based on the determinant value of the Hessian matrix. The specific formula is as follows:
[0124] (1) Calculate the Hessian matrix:
[0125]
[0126] Among them, L xx , L xy and L yy is the Gaussian second derivative of the image at point x.
[0127] (2) Calculate the determinant of the Hessian matrix:
[0128] det(H)=L xx (x, σ)·L yy (x, σ)-L xy (x, σ) 2
[0129] When det(H) is greater than a certain threshold, the point is considered to be a feature point.
[0130] (3) Assign a direction to each feature point to ensure the rotation invariance of the descriptor.
[0131] (4) Select a neighborhood around the feature point and calculate the gradient direction and gradient magnitude of the pixels in the neighborhood to generate a descriptor for the feature point.
[0132] To ensure the rotation invariance of the descriptor, the SURF algorithm assigns a main direction to each feature point. The main direction is obtained by counting the gradient direction and gradient magnitude of the pixels in the neighborhood of the feature point. By rotating the pixels in the neighborhood of the feature point to the main direction, we can ensure that the descriptor is robust to rotation changes of the target.
[0133] Finally, the SURF algorithm selects a neighborhood around the feature point and calculates the gradient direction and gradient magnitude of the pixels in the neighborhood to generate the descriptor of the feature point. The descriptor is a vector that represents the local feature information of the feature point. By comparing the descriptors of feature points in different frames, we can establish the correspondence between feature points and thus achieve target tracking.
[0134] 5. Feature point tracking
[0135] During the tracking process, feature points are extracted for each frame and their descriptors are calculated. Then, the object detector in the KCF algorithm is used to search for the position that best matches the object descriptor in the current frame. By comparing the similarity between descriptors, the correspondence between feature points can be established and the position of the object in the current frame can be determined.
[0136] The KCF algorithm tracking process is expressed as the following formula:
[0137] (1) Calculate the objective function:
[0138]
[0139] Among them, α i are the coefficients of the linear regression model, k(x, z i ) is the kernel function used to calculate the target position x and the training sample z in the current frame i The similarity between .
[0140] (2) Find the optimal solution by minimizing the objective function:
[0141]
[0142] Among them, K is the kernel matrix, λ is the regularization parameter, I is the identity matrix, and y is the label of the training sample.
[0143] (3) Use the coefficients obtained To update the model and calculate the position of the target in the next frame.
[0144] In order to update the position and motion model of feature points in real time, an iterative approach is adopted. In each frame, the position of the feature points is updated according to the prediction results of the KCF algorithm, and its descriptor is recalculated. Then, the updated descriptor is used to track the next frame. In this way, the position and motion model of feature points can be gradually updated to adapt to the motion changes of the target between consecutive frames.
[0145] 6. Calculate the target trajectory
[0146] In this step, the overall motion trajectory of the target object is calculated based on the position changes of the tracked feature points.
[0147] Assuming that the tracked feature point positions are (x1, y1), (x2, y2), ..., (xn, yn), use the least squares method to fit a straight line as the target's motion trajectory. The specific formula is as follows
[0148] Assume that the target's trajectory can be represented by a straight line, and its equation is:
[0149] y=mx+b
[0150] Among them, m is the slope of the line and b is the intercept of the line. Our goal is to find the best m and b so that the line can best fit the tracked feature point position.
[0151] The least squares method is a common mathematical optimization technique that finds the best fitting parameters by minimizing the sum of squared errors. In this problem, the error e i It can be defined as the perpendicular distance from each feature point to the line:
[0152] e i =y i -(mx i +b)
[0153] The total error sum of squares S can be expressed as:
[0154]
[0155] In order to find the best m and b, we need to find the partial derivative of S and set it equal to zero:
[0156]
[0157] By solving this set of equations, we can get the optimal solution for m and b, and thus draw the trajectory of the target.
[0158] 7. Trajectory Visualization and Performance Evaluation
[0159] Finally, the calculated motion trajectory is visualized and superimposed with the original video frame to more intuitively observe the motion of the target object. Trajectory visualization helps to better understand the dynamic behavior and motion pattern of the target.
[0160] In order to evaluate the accuracy and stability of trajectory calculation, a series of performance indicators are used for quantitative evaluation. These indicators include trajectory length, trajectory smoothness, deviation from the true trajectory, etc. The trajectory length reflects the total moving distance of the target in the video sequence; the trajectory smoothness describes the continuity and change rate of the trajectory; and the deviation from the true trajectory measures the degree of difference between the calculated trajectory and the actual trajectory.
[0161] In summary, the present invention realizes accurate detection and trajectory calculation of target objects in video frames through steps such as video frame preparation and preprocessing, target detection model training, target detection and bounding box screening, feature point extraction, feature point tracking and position update, target motion trajectory calculation, trajectory visualization and performance evaluation. This method is efficient, accurate and robust, and is of great significance for target tracking and trajectory analysis in practical applications.
[0162] The above embodiments are only exemplary embodiments of the present invention and are not intended to limit the present invention. The protection scope of the present invention is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present invention within the essence and protection scope of the present invention, and such modifications or equivalent substitutions shall also be deemed to fall within the protection scope of the present invention.
Claims
1. An object-oriented feature point tracking method, characterized in that: include: Perform target detection on the video frame to obtain the target detection frame; Extract feature points in the detection frame of target detection and track the feature points; Calculate the target motion trajectory.
2. The object-oriented feature point tracking method according to claim 1, characterized in that: Before performing target detection on a video frame, the method includes: Perform video frame preparation and preprocessing, using the video processing library to split the video file into individual frames as input data for the subsequent object detection model; In the preparation stage, the object categories and bounding boxes in the video frames are annotated to build an object detection dataset; During preprocessing, each frame is resized to fit the model’s input requirements; Normalization is performed to scale pixel values to a specific range to eliminate differences between images due to lighting and contrast factors; Generate more training samples by using data augmentation techniques that at least include randomly flipping, rotating, cropping, and scaling the images.
3. The object-oriented feature point tracking method according to claim 2, characterized in that: The performing target detection on the video frame comprises: The real-time target detection algorithm YOLOv5 is used for model training. YOLOv5 is a target detection algorithm based on deep learning. YOLOv5 is used to detect targets in video frames to obtain the target detection frame. During the training process, the preprocessed video frames are first loaded as input data; then, the input data is passed to the target detection network through forward propagation. The target detection network includes a feature extraction network, a feature pyramid network, and a detection network; among them, the feature extraction network is responsible for extracting features from the input image, the feature pyramid network is used to fuse feature information of different scales, and the detection network is responsible for predicting the position, size, and category probability of the target bounding box.
4. The object-oriented feature point tracking method according to claim 3, characterized in that: The steps of model training using the real-time target detection algorithm YOLOv5 are as follows: Forward propagation: Input the video frame with the target bounding box annotation into the target detection network, and obtain the predicted position of the target box through the feature extraction network, feature pyramid network, and detection network; Loss calculation: It includes three parts: bounding box regression loss, target confidence loss and classification loss. The bounding box regression loss is used to measure the position difference between the predicted bounding box and the actual bounding box, the target confidence loss is used to measure the difference between the probability of the predicted target and the actual annotation, and the classification loss is used to measure the difference between the predicted category and the actual category. CIOU loss is used as the calculation formula for bounding box regression loss. CIOU loss considers at least the aspect ratio and center point distance of the bounding box on the basis of IOU loss. The confidence loss and classification loss are calculated using cross entropy loss respectively. The bounding box regression loss uses CIOU loss, and the calculation formula is as follows: Among them, B represents the bounding box and gt represents the true value; The confidence loss uses cross entropy loss, and the calculation formula is as follows: P i =Sigmoid(IoU) Among them, P represents the confidence level; The classification loss also uses cross entropy loss, and the calculation formula is as follows: C i =Sigmσid(c) Where C represents the category score; Back propagation and optimization: The gradient of the loss function to the model parameters is calculated through the back propagation algorithm, and the optimizer SGD is used to update the parameters to minimize the loss function; Iterative training: Repeat the process of forward propagation, loss calculation, and backpropagation until the model reaches the preset number of training rounds.
5. The object-oriented feature point tracking method according to claim 4, characterized in that: After completing the target detection model training, the method deploys the model into the video frame processing flow, and the method further includes: Use the non-maximum suppression algorithm to eliminate the redundant bounding boxes generated by YOLOv5 when training the model; Filter out low-confidence detection results according to the confidence threshold, retain bounding boxes with confidence higher than the set threshold, draw the retained bounding boxes on the original frames, including the coordinates, size and category labels of the bounding boxes, and save the frames with detection results as image sequences or reassemble them into video files.
6. The object-oriented feature point tracking method according to claim 5, characterized in that: The step of extracting feature points in the detection frame of the target detection includes: After obtaining the target bounding box of each frame, a feature point detection algorithm is used to extract the feature points inside the target bounding box; specifically, The SURF algorithm is used to generate the descriptor of the feature points by calculating the gradient and direction information of the local area of the image; For the SURF algorithm, the descriptor of the feature point is calculated based on the determinant value of the Hessian matrix. The specific formula is as follows: Compute the Hessian matrix: Among them, L xx ,Lx y and L yy is the Gaussian second-order derivative of the image at point x; Compute the determinant of the Hessian matrix: det(H)=L xx (x,σ)·L yy (x,σ)-L xy (x,σ) 2 When det(H) is greater than a certain threshold, the point is considered to be a feature point; Assign a direction to each feature point to ensure the rotation invariance of the descriptor; Select a neighborhood around the feature point, and calculate the gradient direction and gradient magnitude of the pixels in the neighborhood to generate the descriptor of the feature point; In order to ensure the rotation invariance of the descriptor, the SURF algorithm assigns a main direction to each feature point. The main direction is obtained by counting the gradient direction and gradient magnitude of the pixels in the neighborhood of the feature point. By rotating the pixels in the neighborhood of the feature point to the main direction, the descriptor is ensured to be robust to the rotation changes of the target; The SURF algorithm selects a neighborhood around the feature point and calculates the gradient direction and gradient magnitude of the pixels in the neighborhood to generate a descriptor of the feature point. The descriptor is a vector used to represent the local feature information of the feature point. By comparing the descriptors of feature points in different frames, the correspondence between feature points is established to achieve target tracking.
7. The object-oriented feature point tracking method according to claim 6, characterized in that: The tracking of the feature points includes: During the tracking process, feature points are extracted for each frame and their descriptors are calculated; Using the target detector in the KCF algorithm, search for the position that best matches the target descriptor in the current frame; by comparing the similarity between descriptors, establish the correspondence between feature points and determine the position of the target in the current frame; The KCF algorithm tracking process is expressed as the following formula: Calculate the objective function: Among them, α i are the coefficients of the linear regression model, k(x,z i ) is the kernel function used to calculate the target position x and the training sample z in the current frame i The similarity between The optimal solution is found by minimizing the objective function: Where K is the kernel matrix, λ is the regularization parameter, I is the identity matrix, and y is the label of the training sample; Use the coefficients found To update the model and calculate the position of the target in the next frame; In order to update the position and motion model of feature points in real time, an iterative method is adopted. In each frame, the position of the feature points is updated according to the prediction results of the KCF algorithm, and its descriptor is recalculated; then, the updated descriptor is used to track the next frame; the position and motion model of the feature points are gradually updated to adapt to the motion changes of the target between consecutive frames.
8. The object-oriented feature point tracking method according to claim 7, characterized in that: The calculating target motion trajectory comprises: Calculate the overall motion trajectory of the target object based on the position changes of the tracked feature points; Assuming that the tracked feature point positions are (x1, y1), (x2, y2), …, (xn, yn), use the least squares method to fit a straight line as the target's motion trajectory; the specific formula is as follows, Assume that the target's trajectory can be represented by a straight line, and its equation is: y=mx+b Among them, m is the slope of the line, and b is the intercept of the line. The goal is to find the best m and b so that the line can best fit the tracked feature point position. The least squares method finds the best fitting parameters by minimizing the sum of squared errors; the error e i Defined as the perpendicular distance from each feature point to the line: e i =y i -(mx i +b) The total error sum of squares S is expressed as: In order to find the best m and b, we need to find the partial derivative of S and set it equal to zero: Solve this set of equations to obtain the optimal solution for m and b, thereby drawing the target's motion trajectory.
9. The object-oriented feature point tracking method according to claim 8, characterized in that: The method further comprises: Perform trajectory visualization and performance evaluation, visualize the calculated motion trajectory, and overlay it with the original video frame to more intuitively observe the motion of the target object; The accuracy and stability of trajectory calculation are quantitatively evaluated using indicators including at least trajectory length, trajectory smoothness, and deviation from the true trajectory; Among them, the trajectory length reflects the total moving distance of the target in the video sequence; the trajectory smoothness describes the continuity and change rate of the trajectory; and the deviation from the true trajectory measures the degree of difference between the calculated trajectory and the actual trajectory.