Vehicle auxiliary driving moving target real-time detection and tracking method
By introducing convolutional filtering methods of scale evolution factors and fHOG dynamic features, the real-time and robustness of the vehicle-mounted tracking system in complex environments is solved, and real-time object detection and tracking in vehicle-assisted driving is realized.
Patent Information
- Application Number
- CN202411974485.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing vehicle-mounted tracking systems are difficult to meet the requirements of real-time, reliability and accuracy in complex and diverse environments. Traditional visual tracking algorithms are slow and cannot effectively deal with problems such as lighting changes, scale changes and target occlusion in vehicle-assisted driving.
The minimum output square deviation and squared and weak deviation filtering method based on convolution filtering is adopted to introduce scale evolution factors, and fHOG dynamic features are used instead of grayscale features. By updating the model appearance, center position and scale evolution factors of the target online, the target position and scale changes are quickly estimated.
It has achieved real-time and robustness improvements in vehicle-mounted systems, and can effectively deal with lighting changes, scale evolution and target occlusion in complex environments, meeting the tracking needs of vehicle-assisted driving.
Smart Images

Figure CN120375321A_ABST
Abstract
Description
Technical Field
[0001] This application relates to a method for detecting and tracking moving targets of an automobile, and particularly to a method for real-time detection and tracking of moving targets for vehicle assisted driving, belonging to the technical field of vehicle assisted intelligent driving. Background Art
[0002] The detection and tracking of moving targets are hotspots in computer vision, which are divided into three steps: extraction of target features, target detection, and target tracking, so as to obtain moving target information, including the speed, position, and moving trajectory of the moving target, etc., which can further interpret the moving target and facilitate decision-makers to make decisions. In practical applications, due to many adverse interference factors, higher requirements are put forward for the detection and tracking algorithms of moving targets. The detection of moving targets has important applications in many fields: (1) Video surveillance system: monitoring all changes occurring in the monitored scene, analyzing suspicious persons, and the security of the monitored scene, etc.; (2) Human-computer interaction system: recognizing gestures made by the human body and interpreting possible behaviors, and recognizing faces, etc.; (3) Video indexing: annotating multimedia data in a multimedia database and extracting videos of interest to people; (4) Monitoring traffic; (5) Navigation system: navigating and planning the route of a vehicle or aircraft during travel, avoiding obstacles, and ensuring smooth and safe travel of the vehicle or aircraft. The detection and tracking of moving targets have high frontiers and practicality.
[0003] The popularization of automobiles has brought convenience and speed to people's lives, but also brought certain potential safety hazards. For the safety of themselves and others, people are more eager to make their automobiles intelligent. In recent years, it has become increasingly common to install a micro camera or in-vehicle monitoring on an automobile. As a direction in the field of computer vision, the ultimate goal of vehicle assisted driving is to enable the vehicle to drive smoothly and safely without a driver, using an in-vehicle camera to detect basic information on the road, realizing real-time detection and tracking of moving vehicles or pedestrians, and ensuring the environmental safety of the driving vehicle. The research on the detection and tracking algorithm of moving targets based on an in-vehicle camera has very important application value in a vehicle assisted driving system.
[0004] The problems to be solved in the registration of multi-source vector-raster remote sensing images in the prior art and the key technical difficulties of this application include:
[0005] (1) To extract the features of a moving target, certain information about the target, such as its size and appearance, must be known in advance. Therefore, different feature extraction operators are selected according to the characteristics of the target, and different target models are selected according to different scenarios. Moving target tracking means specifying the tracking target in the first frame, calculating the initial state of the target in the first frame, and then combining time and space to compare the subsequent state of the target with the initial state to find the alternative target with the highest similarity. For a good tracking system, it needs to meet reliability, real-time performance, and accuracy at the same time. However, it is very difficult to achieve these three points simultaneously. For systems under different applications, the emphasis requirements will be different. For the moving object detection and tracking system of in-vehicle cameras, real-time performance is the most important for drivers, followed by accuracy. All in all, the complex and diverse environment poses a great challenge to the tracking system. The algorithm analysis and processing capabilities of existing tracking systems cannot well handle the complex and changing environment and cannot meet the requirements of in-vehicle tracking systems for the real-time performance, reliability, and robustness of tracking algorithms.
[0006] (2) The complex and diverse vehicle assisted driving environment poses a great challenge to the tracking system. Traditional visual tracking algorithms are slow and can no longer meet the requirements of in-vehicle tracking systems. There is an urgent need for a novel, fast, and robust tracking algorithm. However, existing technologies lack a tracking algorithm based on correlation filtering to establish the minimum output square deviation and the sum of squared weak deviations, with a very high computational complexity and unable to meet the real-time requirements of in-vehicle tracking systems in terms of speed. It cannot resist scale changes. When the target scale changes too much, the learned target appearance shape will be very different from the actual shape, resulting in an unsatisfactory tracking effect and ultimately leading to the failure of tracking the moving target. It does not consider the real-time scale change of the target, does not introduce a scale evolution factor into the target model, lacks fHOG dynamic features to replace gray features and cannot represent the texture information and edge information of the target, lacks a fast estimation of the position of the target in the new frame, cannot calculate the scale evolution factor of the target in the new frame, lacks online updating of the model appearance, center position, and scale evolution factor of the target, cannot resist the change of the model scale, cannot balance real-time performance and reliability, is not suitable for use in in-vehicle tracking systems, and cannot provide good support for vehicle assisted driving.
[0007] (3) In the actual target tracking applications of the prior art, the following conditions are defaulted: 1) The movement of the target is continuous and smooth without sudden changes, such as sudden emergency braking of a vehicle, sudden U-turn, sudden fall or sudden disappearance of a pedestrian, etc.; 2) The target moves at a constant speed or in a steady motion; 3) The prior knowledge of the target is known, such as the number, size, contour and appearance of the moving target, etc. In practice, it is very difficult to have such ideal conditions. Due to the variability of road conditions, it is very difficult for a vehicle to move at a constant speed during driving. During the tracking process, the tracked vehicle disappears from the field of view, an unknown vehicle enters the field of view, and the number of vehicles in the field of view constantly changes, which are all common situations. There are still many unsolved problems and difficulties in the field of moving target tracking, including how to solve the interference caused by noise in the image, the irregular motion mode of the tracked target, the non-rigid transformation of the target, the local or global occlusion of other objects, the illumination change in the scene, etc. The prior art cannot well solve the above problems, and the robustness, reliability and real-time performance of the real-time detection and tracking of moving targets in vehicle-assisted driving are poor. Summary of the Invention
[0008] This application establishes a minimum output square deviation and sum of squared weak deviation filtering method based on convolutional filtering, which fully meets the real-time requirements of the vehicle-mounted tracking system in terms of speed. This algorithm has a certain robustness in the tracking effect of problems such as illumination change, scale evolution, rotation change, and target occlusion, meets the real-time requirements of the vehicle-mounted tracking system, and its tracking performance is better than other tracking algorithms. During the tracking process, the scale of the vehicle is constantly changing, and it is difficult for this algorithm to accurately track the target. Based on the sum of squared weak deviation algorithm, the change of the real-time scale of the target is taken into account, and a scale evolution factor is added to the target model to resist the change of the model scale. The improved algorithm performs well in both real-time performance and robustness, and can also well handle the scale evolution of moving targets, meeting the requirements of the complex and diverse environment of vehicle-assisted driving for the tracking system. The algorithm analysis and processing ability of the tracking system can well handle the complex and changeable environment and can meet the requirements of the vehicle-mounted tracking system for the real-time performance, reliability and robustness of the tracking algorithm.
[0009] To achieve the above technical effects, the technical solutions adopted in this application are as follows:
[0010] A real-time detection and tracking method for moving targets in vehicle assisted driving. Based on convolutional filtering, a minimum output sum of squared deviations and sum of squared weak deviations filtering method is established. On the basis of the sum of squared weak deviations algorithm, the change of the real-time scale of the target is considered. A scale evolution factor is introduced into the target model, and the fHOG motion feature is used to replace the gray feature in the sum of squared weak deviations to characterize the texture information and edge information of the target. The specific method is to first use the improved sum of squared weak deviations function to quickly estimate the position of the target in the new frame, then calculate the scale evolution factor of the target in the new frame, and finally uniformly calculate and update the filter and the scale evolution factor. By online updating the model appearance, center position and scale evolution factor of the target, the change of the model scale is resisted;
[0011] The following improvement strategies are adopted for the sum of squared weak deviations algorithm with low feature dimension and no consideration of the change of the real-time scale of the target:
[0012] 1) Obtain high-dimensional features based on the fHOG transform: The fHOG is obtained by optimizing the HOG speed, maintaining optical and geometric invariance, and the fHOG is used as the characterization of the target feature;
[0013] 2) Introduce a scale evolution factor into the sum of squared weak deviations algorithm: First, use the initial sum of squared weak deviations algorithm to track and obtain the center point of the target in the new frame, calculate the scale evolution factor of the target in the new frame, and jointly update the filter and the scale evolution factor to ensure the algorithm speed while solving the influence of scale evolution on target tracking;
[0014] The improved algorithm based on the real-time performance of target tracking is divided into two steps: First, use the improved sum of squared weak deviations function to quickly estimate the position of the target in the new frame to achieve the tracking of the moving target position, then calculate the scale evolution factor of the target in the new frame, calculate the fHOG motion feature, and finally jointly update the filter and the scale evolution factor to achieve the real-time scale estimation of the target.
[0015] Preferably, for the moving target position tracking: The error function of the sum of squared weak deviations filtering tracking algorithm is as follows:
[0016]
[0017] d is the feature dimension, represents the circular convolution operation, and the feature map is f l (l = 1, 2, 3, …, d), h l (l = 1, 2, 3, …, d) represents the correlation filter under the corresponding feature dimension, obtained by solving through the error function formula 9, g is the desired output, f l 、h l and g have the same dimension and the same size, but simply relying on formula 9 to solve for h requires a division operation, and formula 9 needs to be improved:
[0018]
[0019] By adding the λ term, the situation of division by zero in the solution process is excluded. In addition, the variation range of the filter parameters is also controlled. The smaller the λ, the larger the variation range of the filter parameters. Whether to add the λ term has little impact on the initial frame perturbation. Finally, the λ term is not added.
[0020] Equation 10 is a real convex function, and there is only one optimal solution, which is the point where the derivative is 0. Perform Fourier transform on Equation 10, take the derivative, and set the derivative to 0. After the transformation, it is as follows:
[0021]
[0022] In Equation 11, all are capital letters, indicating that all functions in the spatial domain have been transformed into the frequency domain. It represents the conjugate complex number of the frequency domain matrix G obtained after the Fourier transform of the expected output matrix g. It represents the fHOG motion feature f of the k-th dimension of the target region. k The conjugate complex number of the k-th dimension feature F in the frequency domain obtained after the Fourier transform. In the actual tracking process, the following formula is used to update the model online: k The conjugate complex number of, and in the actual tracking process, the following formula is used to update the model online:
[0023]
[0024] n is the learning rate, and its value range is [0, 1]. The larger the learning rate, the more the model can learn the new appearance of the moving target. If it is 1, it means the model represents the target with the latest appearance. If it is 0, it means the model is not updated from beginning to end. In this application, the learning rate of 0.015 is the most appropriate.
[0025] After updating the filter, the following formula is used to predict the center position of the target in the new frame:
[0026]
[0027] Z l It is the l-th dimension feature map of the fHOG motion feature calculated for the target candidate region of the latest frame. The position of the maximum value of y is determined as the center point position of the target region in the (t + 1)-th frame.
[0028] Preferably, calculate the fHOG motion feature: calculate the histogram of oriented gradients of the local area of the vehicle-mounted camera image, and then calculate the obtained histogram of oriented gradients as a feature operator representing the moving target, that is, first divide the target image into small cell units, and then calculate the histogram of the gradient magnitude and gradient direction of each pixel point in each cell unit. The calculation of the histogram of oriented gradients of the vehicle-mounted moving image is the final fHOG motion feature.
[0029] The specific calculation steps of the fHOG motion feature operator are as follows:
[0030] Step 1: Select the vehicle assisted driving moving window or target area to be detected;
[0031] Step 2: If the input vehicle-mounted image is a color image, first grayscale the image, and then normalize the color space of the grayscale image using the Gamma correction method to adjust the contrast of the image, reduce the influence brought by local shadows and lighting changes in the image, and at the same time resist noise interference;
[0032] Step 3: Calculate the gradient magnitude and direction of each pixel point in the vehicle assisted driving moving target area;
[0033] Step 4: Divide the image into small cell units cel1, and there are 5*5 pixels in one cell unit;
[0034] Step 5: Calculate the gradient histogram of each cell unit;
[0035] Step 6: A certain number of cell units form a block, 4*4 cell units form a block, and the target is composed of fHOG motion feature descriptors of multiple blocks. The feature descriptor of the block is calculated from the feature descriptors of all cell units within the block;
[0036] Step 7: The fHOG motion feature of the target is obtained by combining the fHOG motion feature descriptors of all blocks in the vehicle-mounted image.
[0037] Preferably, for real-time scale estimation of the target: First, obtain the latest position, take the scale evolution factor into account, first obtain the center point of the target, then calculate the scale evolution factor of the target in the new frame, and finally uniformly calculate and update the filter and the scale evolution factor, with a smaller time complexity and more guaranteed real-time performance. The algorithm steps include: ① Input, ② Position estimation, ③ Scale estimation, ④ Update model.
[0038] Preferably, ① Input:
[0039] The image I of the t-th frame t ;
[0040] The position p of the previous frame t-1 and the scale evolution factor s t-1 ;
[0041] The center position model filter
[0042] Preferably, ② Position estimation:
[0043] The first step: In the image I of the t-th frame t according to the position p where the center of the target was located in the previous framet-1 and the target real-time scale evolution factor s of the previous frame t-1 Obtain the candidate region of the target, and obtain the fHOG motion feature map z of the target through calculation t trans ;
[0044] Step 2: The feature map z within the region where the estimated target position is located t trans is subjected to correlation convolution calculation with the position filter to obtain the candidate matrix y of the latest position t trans ;
[0045] Step 3: Set the point with the maximum value of y t trans as the latest center position p of the moving target in the current frame t .
[0046] Preferably, ③ Estimate the scale:
[0047] Step 4: In the t-th frame image I t , according to the latest position p t and the scale evolution factor s of the previous frame t-1 delimit the region where the target is located, obtain S regions through scaling, and extract the feature maps z of the S regions t scale ;
[0048] Step 5: The S region feature maps z within the region where the estimated target position is located t scale are subjected to correlation calculation with the filter to obtain y t scale ;
[0049] Step 6: Set the maximum value of y t scale as the target real-time scale evolution factor s of the current frame t .
[0050] Preferably, ④ Update the model:
[0051] Step 7: In the t-th frame image I t , according to the target center position p estimated in steps ② and ③ t and the scale evolution factor s t delimit the region where the target is located, extract the features of the target, and obtain the training sample f after Fourier transform t trans ;
[0052] Step 8: Update the position filter
[0053] 1) The movement of the target is smooth in a very short period of time, and the approximate center position of the target in the current frame is obtained through the state of the previous frame; 2) The scale evolution of the target changes little in a very short period of time, and the region of the target in the current frame is obtained through the scale evolution factor of the previous frame and the latest center position of the target.
[0054] Preferably, the sum of squares and weak deviation filtering tracking: The convolution calculation in the spatial domain of the digital image is equivalent to the element-by-element multiplication between two matrices after the image is transformed into the frequency domain, saving the calculation time of the convolution calculation in the spatial domain. Assume that the training samples are a series of digital images f i which becomes F after Fourier transform i and the output result G is given artificially i to obtain the filter H i * is:
[0055]
[0056] The final filter H is optimized by minimizing the output mean square sum deviation:
[0057]
[0058] In the formula, ⊙ represents element-by-element multiplication between matrices. The elements of the filter H are calculated at the corresponding positions in the frequency domain space, and any element in the filter H is solved separately. Assume that H wv is an independent element in H, and the filter is solved by separately solving H wv to obtain the filter, and Equation 2 is converted into the following form:
[0059]
[0060] The above function is a real function and a convex function, with only one optimal point. To find the optimal point, only the derivative of the function needs to be taken, and the point where the derivative function is 0 is the optimal point of the function, that is:
[0061]
[0062] to obtain the expression of H wv :
[0063]
[0064] After separately obtaining the elements in the filter H, all the elements in the filter H are obtained, and the filter H is solved.
[0065] Preferably, multi-filter online update for change tracking: Moving objects continuously change in the actual scene image, including changes in scene illumination, scale evolution, rotation, pose, and object position. The real-time multi-filter is updated online to adapt to the above changes. The following method is used to update the model:
[0066]
[0067]
[0068]
[0069]
[0070]
[0071] where t represents the frame number, and F t * represents the matrix obtained after Fourier transform of the image within the target area of the t-th frame, and G t represents the desired output matrix G t * is the conjugate complex number of G, and F t represents the conjugate complex number of F t * . η represents the learning rate, and the learning rate is taken as 0.125 to make the update speed of the filter keep up with the appearance change of the target;
[0072] The filter H of the t-th frame is obtained by solving. For the (t + 1)-th frame, its feature map Z is calculated, and the correlation convolution of H t and Z is performed. In the finally output result, the point where the peak value is located is the center position of the tracked target; t In the finally output result, the position of the maximum value of y is regarded as the center point of the target in the (t + 1)-th frame image;
[0073]
[0074] Based on the tracking algorithm of convolutional filtering, the peak-to-sidelobe ratio PSR is used to represent the tracking situation of the target, and the definition is as follows:
[0075] The steepness of the two-dimensional Gaussian distribution surface affects the tracking accuracy. The tracking situation of the target is characterized by the above formula, and μ and σ represent the mean and variance of the pixels excluding the peak point within the 11 * 11 window with the peak as the center point.
[0076]
[0077] Compared with the prior art, the innovation points and advantages of this application are as follows:
[0078] Compared with the prior art, the innovation points and advantages of this application are:
[0079] (1) This application establishes a minimum output square deviation and sum of squared weak deviation tracking algorithm based on correlation filtering. The sum of squared weak deviation algorithm has a very low computational complexity, and its speed can reach 669fps, fully meeting the real-time requirements of in-vehicle tracking systems in terms of speed. On the basis of the sum of squared weak deviation algorithm, the real-time scale change of the target is taken into account, a scale evolution factor is introduced into the target model, and the fHOG motion feature is used to replace the gray feature in the sum of squared weak deviation to represent the texture information and edge information of the target. An improved sum of squared weak deviation function is used to quickly estimate the position of the target in the new frame, then the scale evolution factor of the target in the new frame is calculated, and finally the filter and scale evolution factor are updated uniformly. By online updating the model appearance, center position and scale evolution factor of the target, the change of the model scale is resisted. The improved algorithm in this application can well resist the change of the model scale. Through a large number of dataset tests and comparative experiments with measured data, the results show that the real-time performance and robustness of the improved algorithm are both very good. The improved algorithm takes into account both real-time performance and reliability, indicating that this algorithm is suitable for use in in-vehicle tracking systems and provides support for vehicle assisted driving.
[0080] (2) This application establishes a minimum output square deviation and sum of squared weak deviation filtering method based on convolutional filtering, which fully meets the real-time requirements of in-vehicle tracking systems in terms of speed. This algorithm has a certain robustness in tracking effects for problems such as illumination changes, scale evolution, rotation changes, and target occlusion, meeting the real-time requirements of in-vehicle tracking systems and having better tracking performance than other tracking algorithms. In view of the fact that during the tracking process, the scale of the vehicle is constantly changing and it is difficult for this algorithm to accurately track the target. On the basis of the sum of squared weak deviation algorithm, the real-time scale change of the target is taken into account, and a scale evolution factor is added to the target model to resist the change of the model scale. The improved algorithm shows very good real-time performance and robustness, and can also well handle the scale evolution of moving targets, meeting the requirements of the complex and diverse environment of vehicle assisted driving for the tracking system. The algorithm analysis and processing ability of the tracking system can well handle the complex and changeable environment and can meet the requirements of in-vehicle tracking systems for the real-time performance, reliability and robustness of tracking algorithms.
[0081] (3) Based on the drawbacks of the minimum output deviation sum of squares and weak deviation filtering algorithm, the filter contains the appearance shape information of the target and models the target appearance. Once the scale evolution is too large, the learned target appearance will be quite different from the actual target, ultimately resulting in the failure of target tracking. To address this defect, the position factor is solved based on the sum of squares weak deviation algorithm, and the scale evolution factor is introduced into the algorithm. This can not only solve the problem of scale evolution but also ensure that the robustness of the improved algorithm is not affected. This application demonstrates the real-time performance, reliability, and robustness of the improved algorithm from multiple aspects. The improvement idea is to first obtain the center point of the target, then calculate the scale evolution factor of the target in the new frame, and finally update the filter and the scale evolution factor uniformly. The computational complexity of the improved algorithm is only a linear multiple of the time of the sum of squares weak deviation algorithm, so the real-time performance can be guaranteed. The position factor and the scale evolution factor are solved separately because the movement of the target is smooth and the scale change is relatively small within a short period. Therefore, the region of the target in the current frame can be obtained relatively reliably through the state of the previous frame (center point position and scale evolution factor). With the alternative region, all the basic conditions of the sum of squares weak deviation algorithm are satisfied. After testing with relevant specific datasets, such as when the target in the scene is occluded, the illumination intensity changes, rotates, or undergoes scale evolution, the experimental results show that the improved algorithm can well solve the above problems encountered in the dataset, leading to the conclusion that the improved algorithm has good robustness, reliability, and real-time performance. Using on-vehicle cameras to detect basic information on the road, realizing real-time detection and tracking of moving vehicles or pedestrians, and ensuring the environmental safety of driving vehicles have important application values in vehicle assisted driving systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 It is a schematic diagram for estimating the scale evolution factor.
[0083] Figure 2 It is a schematic diagram of actual samples of different sizes shown by filters of different scales.
[0084] Figure 3 It is a comparison chart of the tracking results when the target in the hiding dataset is completely occluded.
[0085] Figure 4 It is a schematic diagram of the tracking results when the illumination changes within the scene of the David test dataset.
[0086] Figure 5 It is a schematic diagram of the test data results when the target in the Sylv test dataset undergoes rotational changes.
[0087] Figure 6 It is a comparison chart of Test Experiment 1 of different tracking algorithms for the same dataset.
[0088] Figure 7 It is a comparison chart of the test experiments 2 of two tracking algorithms for the same data set. The left figure shows the effect of the minimum output square deviation and sum of squares weak deviation filtering method, and the right figure shows the effect of the improved method. Specific implementation manners
[0089] The following further describes the technical solutions of the vehicle assisted driving moving target real-time detection and tracking method provided by the present application with reference to the accompanying drawings, so that those skilled in the art can better understand the present application and be able to implement it.
[0090] The popularity of automobiles has brought convenience and speed to people's lives, but at the same time, it has also brought certain potential safety hazards, intensifying the demand for vehicle intelligence. Target detection and tracking in specific environments are particularly important. Mobile target detection and tracking based on in-vehicle cameras have become one of the important topics in the current fields of computer vision and intelligent vehicles due to their important application value in vehicle assisted driving systems. A good tracking system needs to meet reliability, real-time performance, and accuracy at the same time. Real-time performance is the most important for drivers, followed by reliability and accuracy. The complex and diverse environment poses a great challenge to the tracking system. Although the traditional visual tracking algorithms meet the accuracy requirements, they are slow. The visual tracking algorithms of the existing technologies can no longer meet the requirements of the in-vehicle tracking system, and there is an urgent need for a novel, fast, and robust tracking algorithm.
[0091] A minimum output square deviation and sum of squares weak deviation tracking algorithm is established based on correlation filtering, and good tracking effects have been achieved. The sum of squares weak deviation algorithm has a very small computational complexity, and the speed can reach 669fps, which fully meets the real-time requirements of the in-vehicle tracking system in terms of speed. However, the tracking algorithm based on sum of squares weak deviation filtering can only resist scale changes within a certain range. In the actual in-vehicle camera environment, the real-time scale of the tracked target is constantly changing. Therefore, if the real-time scale change of the target is too large, it may cause a too large difference between the learned target appearance shape and the actual shape, resulting in an unsatisfactory tracking effect and ultimately leading to the failure of tracking the moving target. Based on the sum of squares weak deviation algorithm, the present application takes into account the change of the real-time scale of the target, introduces a scale evolution factor into the target model, and uses the fHOG motion feature to replace the gray feature in the sum of squares weak deviation to characterize the texture information and edge information of the target. The specific method is to first use the improved sum of squares weak deviation function to quickly estimate the position of the target in the new frame, then calculate the scale evolution factor of the target in the new frame, and finally uniformly calculate and update the filter and the scale evolution factor. By online updating the model appearance, center position, and scale evolution factor of the target, the change of the model scale can be resisted.
[0092] The improved algorithm of this application can well resist the changes in the model scale. Through the comparative tests of a large number of data sets and actual measured data, the results show that the improved algorithm performs very well in real-time performance and robustness. The improved algorithm takes into account both real-time performance and reliability, indicating that this algorithm is suitable for use in vehicle tracking systems and provides support for vehicle assisted driving.
[0093] I. Square Sum Weak Deviation Filter Tracking
[0094] The convolution calculation in the spatial domain of digital images is equivalent to the element-by-element multiplication between two matrices after the image is transformed into the frequency domain, saving the calculation time of convolution calculation in the spatial domain. Assume that the training samples are a series of digital images f i , which becomes F after Fourier transform i , and the output result G is given artificially i , and the filter H is obtained i * as follows:
[0095]
[0096] The final filter H is optimized by minimizing the output mean square sum deviation:
[0097]
[0098] In the formula, ⊙ represents element-by-element multiplication between matrices. The elements of the filter H are calculated at the corresponding positions in the frequency domain space, and any one element in the filter H is solved separately. Assume that H wv is an independent element in H, and the filter is solved by separately solving H wv , and formula 2 is converted into the following form:
[0099]
[0100] The above function is a real function and a convex function, with only one optimal point. To find the optimal point, only the derivative of the function needs to be calculated, and the point where the derivative function is 0 is the optimal point of the function, that is:
[0101]
[0102] Obtain the expression of H wv :
[0103]
[0104] After separately obtaining the elements in the filter H, all the elements in the filter H are obtained, and the filter H is solved.
[0105] II. Multi-Filter Online Update Change Tracking
[0106] The appearance of the moving target remains unchanged in the actual scene, but it continuously changes in the image, including changes in scene illumination, scale evolution, rotation, pose, and target position. The real-time multi-filter is updated online to adapt to the above changes. The model is updated using the following method:
[0107]
[0108]
[0109]
[0110]
[0111]
[0112] In the formula, t represents the number of frames, and F t * represents the matrix obtained after the Fourier transform of the image within the target area of the t-th frame. G t represents the expected output matrix G t * 's conjugate complex number. F t represents F t * 's conjugate complex number. η represents the learning rate, and the learning rate is taken as 0.125 to make the filter update speed keep up with the appearance change of the target.
[0113] The filter H of the t-th frame is obtained by solving. For the (t + 1)-th frame, its feature map Z is calculated. The correlation convolution of H t and Z is performed. In the finally output result, the point where the peak value is located is the center position of the tracked target; t In the finally output result, the position of the maximum value of y is recognized as the center point of the target in the (t + 1)-th frame image;
[0114]
[0115] The position of the maximum value of y is recognized as the center point of the target in the (t + 1)-th frame image;
[0116] Based on the tracking algorithm of convolutional filtering, the peak-to-sidelobe ratio PSR is used to represent the tracking situation of the target, and the definition is as follows:
[0117]
[0118] The steepness of the two-dimensional Gaussian distribution surface affects the tracking accuracy. The tracking situation of the target is characterized by the above formula. μ and σ represent the mean and variance of the pixels excluding the peak point within the 11 * 11 window with the peak as the center point.
[0119] III. Improved algorithm of sum of squares and weak deviation based on scale evolution
[0120] Based on the Sum of Squares Weak Deviation Filter Tracking Algorithm, the relevant convolution filter updates the filter in an online manner. The update of filter H only simply updates and learns the target appearance, and only uses grayscale as the feature of the target during tracking. Since the dimension reflecting the model feature is too low, it is difficult to well represent the target. If there is a deviation in one of the filters H, it may cause a deviation in the subsequent learning of the target model and result in the failure of target tracking. If the feature dimension is high, this problem does not need to be worried about at all. And based on the Sum of Squares Weak Deviation Filter Tracking Algorithm, only the translational movement of the center point of the target area between each frame of images is estimated. If the real-time scale change of the target exceeds a certain degree, it will cause a deviation in the filter, and in severe cases, it will result in the failure of target tracking.
[0121] The defects and deficiencies of the Sum of Squares Weak Deviation Algorithm are due to the low feature dimension and the lack of consideration of the real-time scale change of the target. To address these two problems, the present application adopts the following improvement strategies:
[0122] 1) Obtain high-dimensional features based on the fHOG transformation: fHOG is obtained by optimizing the HOG speed, with small computational complexity and maintaining optical and geometric invariance. fHOG is used as the representation of the target feature;
[0123] 2) Introduce a scale evolution factor in the Sum of Squares Weak Deviation Algorithm: First, use the initial Sum of Squares Weak Deviation Algorithm to track and obtain the center point of the target on the new frame, then calculate the scale evolution factor of the target in the new frame, and finally jointly update the filter and the scale evolution factor to ensure the algorithm speed while solving the impact of scale evolution on target tracking.
[0124] Based on the real-time requirement of the target tracking algorithm for tracking, constructing the Sum of Squares Weak Deviation Algorithm based on scale evolution is divided into two steps: First, use the improved Sum of Squares Weak Deviation function to quickly estimate the position of the target on the new frame to achieve the tracking of the moving target position, then calculate the scale evolution factor of the target in the new frame, calculate the dynamic fHOG feature, and finally calculate and update the filter and the scale evolution factor uniformly to achieve the real-time scale estimation of the target.
[0125] (1) Tracking of the moving target position
[0126] The error function of the Sum of Squares Weak Deviation Filter Tracking Algorithm is as follows:
[0127]
[0128] d is the feature dimension, represents the circular convolution operation, and the feature map is f l (l = 1, 2, 3, …, d), h l(l = 1, 2, 3, …, d) represents the correlation filter under the corresponding feature dimension, which is obtained by solving through the error function formula 9, where g is the expected output, and f l , h l has the same dimension as g and the same size. However, simply relying on formula 9 to solve for h requires a division operation, so formula 9 needs to be improved:
[0129]
[0130] By adding the λ term, the situation of division by zero during the solution process is excluded. Additionally, the variation range of the filter parameters is also controlled. The smaller λ is, the larger the variation range of the filter parameters. Whether to add the λ term has little impact on the initial frame perturbation, and finally, it is chosen not to add the λ term;
[0131] Formula 10 is a real convex function, and there is only one optimal solution, which is the point where the derivative is 0. Perform Fourier transform on formula 10, take the derivative, and set the derivative to 0. After the transformation, it is as follows:
[0132]
[0133] In formula 11, all are capital letters, indicating that all functions in the spatial domain have been transferred to the frequency domain, represents the conjugate complex number of the frequency domain matrix G obtained after Fourier transform of the expected output matrix g, represents the k - dimensional fHOG dynamic feature f of the target region, k and the k - dimensional feature F in the frequency domain obtained after Fourier transform of f k is the conjugate complex number. In the actual tracking process, the following formula is used to update the model online:
[0134]
[0135] n is the learning rate, and its value range is [0, 1]. The larger the learning rate, the more the model can learn the new appearance of the moving target. If it is 1, it means the model represents the target with the latest appearance. If it is 0, it means the model is not updated from beginning to end. The most suitable learning rate in this application is 0.015, which not only enables the filter to update fast enough to keep up with the appearance change of the target but also ensures the robustness of the tracking performance;
[0136] After updating the filter, the following formula is used to predict the center position of the target in the new frame:
[0137]
[0138] Z l is the l - dimensional feature map of the fHOG dynamic feature calculated for the target candidate region in the latest frame, and the position of the maximum value of y is determined as the center point position of the target region in the (t + 1) - th frame.
[0139] (2) Calculate the fHOG motion feature
[0140] Calculate the histogram of oriented gradients (HOG) for the local region of the in-vehicle camera image, and then use the calculated HOG as a feature operator to represent the moving target. That is, first divide the target image into small cell units, and then calculate the histogram of the gradient magnitude and gradient direction for each pixel in each cell unit. The calculation of the HOG for the in-vehicle moving image is the final fHOG motion feature;
[0141] The specific calculation steps of the fHOG motion feature operator are as follows:
[0142] Step 1: Select the moving window or target area of vehicle assisted driving to be detected;
[0143] Step 2: If the input in-vehicle image is a color image, first grayscale the image, and then normalize the color space of the grayscale image using the Gamma correction method to adjust the contrast of the image, reduce the influence of local shadows and lighting changes in the image, and at the same time resist noise interference;
[0144] Step 3: Calculate the gradient magnitude and direction of each pixel in the moving target area of vehicle assisted driving;
[0145] Step 4: Divide the image into small cell units cel1, and there are 5*5 pixels in one cell unit;
[0146] Step 5: Calculate the histogram of oriented gradients for each cell unit;
[0147] Step 6: A certain number of cell units form a block, 4*4 cell units form a block, and the target is composed of the fHOG motion feature descriptors of multiple blocks. The feature descriptor of a block is calculated from the feature descriptors of all the cell units within the block;
[0148] Step 7: The fHOG motion feature of the target is obtained by combining the fHOG motion feature descriptors of all the blocks in the in-vehicle image;
[0149] The fHOG motion feature describes the edge structure feature of the dynamic target. Normalizing the histogram of oriented gradients in the local area of the target makes the fHOG motion feature have anti-interference performance against changes in the illumination intensity in the scene. The fHOG motion feature is invariant to geometric and optical changes and is suitable for human detection. The fHOG is obtained by solving in the densely sampled in-vehicle image blocks. The calculated fHOG motion feature implies the spatial position relationship between the block and the detection window. The dimension of the fHOG motion feature is set to 31, so the computational amount is much smaller compared to the SIFT (128-dimensional) operator.
[0150] (3) Real-time scale estimation of the target
[0151] The above series of calculations do not consider scale evolution for the time being. In this application, the latest position is first obtained. If the scale evolution factor is considered, two ideas are proposed: The first is to solve the latest position and the scale evolution factor simultaneously. The algorithm complexity is: O(dMNS*log(MNS)), where d represents the dimension of the feature, M*N is the number of pixels in the tracking window, and S is the real-time number of available scale coefficients. Solving the position and the scale evolution factor simultaneously greatly increases the time complexity, resulting in the speed not meeting the real-time requirement; The second idea is to first obtain the center point of the target in the first step, then calculate the scale evolution factor of the target in the new frame, and finally calculate and update the filter and the scale evolution factor uniformly. The time complexity of this method becomes O(dMN*log(MN)+dMNS*log(S)). By comparison, it can be found that the time complexity of the second method is significantly smaller and more guaranteed for real-time performance.
[0152] Based on the above discussion, the second method is finally adopted to implement the tracking algorithm for scale evolution. The algorithm steps are as follows:
[0153] ① Input:
[0154] The t-th frame image I t ;
[0155] The position p of the previous frame t-1 and the scale evolution factor s t-1 ;
[0156] The center position model filter
[0157] ② Position estimation:
[0158] The first step: In the t-th frame image I t , obtain the candidate region of the target according to the position p t-1 where the center of the target was located in the previous frame and the real-time scale evolution factor s t-1 of the target in the previous frame. Calculate the fHOG motion feature map z t trans of the target;
[0159] The second step: Correlate and convolve the feature map z t trans in the region where the estimated target position is located with the position filter to obtain the candidate matrix y t trans of the latest position;
[0160] The third step: Set the point with the maximum value of y t trans as the latest center position p of the moving target in the current framet ;
[0161] ③ Estimation scale:
[0162] Step 4: In the t-th frame image I t locate the region where the target is located according to the latest position p t and the scale evolution factor s of the previous frame t-1 define the region where the target is located, obtain S regions through scaling, and extract the feature maps z of the S regions t scale ;
[0163] Step 5: Calculate the correlation between the S region feature maps z t scale within the region where the estimated target position is located and the filter to obtain y t scale ;
[0164] Step 6: Set the maximum value of y t scale as the real-time scale evolution factor s of the target for the current frame t ;
[0165] ④ Update the model:
[0166] Step 7: In the t-th frame image I t locate the region where the target is located according to the estimated target center position p t and the scale evolution factor s in Steps ② and ③ t extract the features of the target to obtain the training sample f after Fourier transform t trans ;
[0167] Step 8: Update the position filter
[0168] 1) The movement of the target is smooth in a very short time. Obtain the approximate center position of the target in the current frame through the state of the previous frame;
[0169] 2) The change in the scale evolution of the target is small in a very short time. Obtain the region of the target in the current frame through the scale evolution factor of the previous frame and the latest center position of the target;
[0170] Compared with the sum of squares and weak deviation algorithm process, the estimation of the scale evolution factor is added. When initializing in the first frame, the scale of the target is defined as 1, as Figure 1As shown, the latest center position of the target obtained by quickly solving the sum of squares and weak deviation energy equation in the current frame is used. Then, the candidate regions of the target in the current frame are obtained by using the center position and 31 scale evolution factors of the fHOG motion feature. All candidate regions are scaled to a specified size, which is the same as the target size in the first frame. Then, through the fHOG motion feature extraction, Fourier transform, and calculation with the same filter, a series of responses are obtained. The maximum values corresponding to these scale evolution factors are found respectively, and the scale evolution factor corresponding to the largest maximum value is the required scale evolution factor;
[0171] The selection method for scale evolution is as follows:
[0172]
[0173] Where P and R represent the width and height of the target in the previous frame respectively, a = 1.02 is the scale evolution factor, S = 31 is the number of scales. The above scales are not in a linear relationship, but a process of estimation from fine to coarse, which is manifested as detection and tracking in the direction from inside to outside on the image. Figure 2 They are actual samples of different sizes shown by different scales in the image.
[0174] IV. Experimental Results and Analysis
[0175] The test datasets mainly come from the VOT dataset (http: / / www.votchallenge.net) and the OTB dataset (http: / / cvlab.hanyang.ac.kr / tracker_benchmark / datasets.html). These two datasets are the most authoritative evaluation platforms in the current field of visual tracking. These two platforms not only provide special test datasets, but also calibrate the datasets. During the test, not only can the tracking effect be observed visually, but the test results can be compared with the groundtruth in the VOT dataset and the OTB dataset to evaluate the tracking performance of the tracker, and the tracking performance comparison and analysis with other tracking algorithms can also be carried out. Selecting these two platforms to test the improved algorithm can not only evaluate the performance of the tracking system visually, but also compare it with other tracking algorithms to evaluate the advantages and disadvantages of this algorithm.
[0176] (I) Occlusion
[0177] In the hiding test data in the VOT dataset, there are situations where the target is partially occluded and completely occluded by other objects during the movement. Judging from the experimental results, the tracking effect is still relatively ideal during partial occlusion. Through Figure 3It can be seen that before the small figure a6, the tracking effect is relatively ideal. However, after the small figure a6, the tracking fails. The actually tracked target is no longer the target because the position of the target in the frame of the small figure a6 is the position where the target is completely occluded. Although the center point of the target can still be detected here, the appearance of the target has changed greatly. Therefore, the appearance of the learned model is very different from the actual appearance shape, ultimately resulting in the failure of target tracking.
[0178] In the VOT crossing dataset, the target is partially occluded by other objects during the movement. Whether during the occlusion process or after the occlusion, the tracking effect is very ideal. Although there is partial occlusion, the appearance model of the target can still well represent the target, and the real-time learning of the target appearance model also enables the continuous tracking of the target. One reason for the good robustness of this algorithm to partial occlusion is that the feature operator fHOG motion feature has local invariance. Although the target is partially occluded, the fHOG is used to extract the feature points of the unoccluded part, and the extracted features can still well represent the target. Although there are differences between the learned target appearance model and the actual target appearance model, the differences between the two are not large enough to affect the target tracking.
[0179] From the results of the above tests, it can be seen that the algorithm of the present application can well adapt to the situation of partial occlusion of the target and has good robustness.
[0180] (2) Scene illumination change
[0181] During the movement of the moving target in the David test dataset, the illumination intensity in the scene changes. To be precise, the target to be tracked moves from a darker place to a place with higher light intensity. The final test results are as Figure 4 shown. From the experimental results, the tracking effect is not affected by the change in illumination intensity. This shows that the improved algorithm is robust to illumination changes. Another small detail is that whether David wears glasses or not does not affect the final tracking result.
[0182] During the movement of the target, the illumination of the above dataset changed significantly. At the beginning, the illumination intensity was very low, and only the outline of the target could be barely seen. By the end, all corners of the indoor area could be clearly seen. In such cases, the target could be successfully tracked, and not even a single frame of data was lost. This indicates that the proposed algorithm can handle most of the illumination changes that occur during the actual tracking of moving targets. Moreover, when the appearance of the target changed during movement, such as taking off and putting on glasses, the tracking was not affected at all, demonstrating that the algorithm of this application is highly robust to changes in the target appearance.
[0183] (III) Rotation and Appearance Changes
[0184] The appearance of a moving target changes more or less during movement, which poses a challenge to tracking algorithms based on target appearance shape modeling. This application is an improvement on the sum of squares and weak deviation algorithm, while retaining the tracking idea of modeling the target appearance in the sum of squares and weak deviation algorithm. The Sylv test data undergoes rotational changes during movement, and thus the appearance shape of the target changes accordingly in the image. Figure 5 The key frames in the tracking results of the test dataset are shown. The tracking effect is not affected by the appearance changes brought about by the rotational changes of the target, indicating that this application is robust to target rotation and appearance changes.
[0185] (IV) Scale Evolution
[0186] During vehicle driving, due to the different speeds of different vehicles, the distance between vehicles changes continuously, which is manifested as the continuous change of the vehicle size in the in-vehicle camera video. Therefore, the greatest challenge of this application is the real-time scale change of the target. By testing the Dogs dataset, it can be seen that the improved algorithm can handle scale evolution well, indicating that the algorithm idea is correct and it has good robustness to scale evolution.
[0187] Through the test results of the above typical datasets, it shows that the improvement proposed in this application based on the sum of squares and weak deviation algorithm has strong robustness to illumination changes, scale evolution, rotational changes, and occlusion. However, the data of the above experiments are all in a relatively ideal environment, with relatively single influencing factors and insufficient persuasiveness. Therefore, it is necessary to use measured data to further test and verify the tracking algorithm of this application.
[0188] (V) Experiment with Measured Data
[0189] The actual measured data here were captured using different devices in different scenarios. For the actual measured data 1, the scenario was that the photographer used a mobile phone to photograph an oncoming vehicle in a deep alley. The actual measured data 2 was that the photographer used an in-vehicle camera in a taxi to photograph the driving process of the vehicle ahead. The actual measured data 3 was of a sanitation worker cleaning the road on the far right during a turning process. The actual measured data 4 was of a private car passing the vehicle the photographer was in. The above actual measured data cover the situations that a driver may encounter during the driving process, with the main targets being people and vehicles.
[0190] From the test results of the above data, it can be seen that the algorithm has strong robustness to the scale evolution during the movement of the target. In actual tests, for different tracking targets, the tracking speeds are also different, and some speeds can even reach over 120fps. From this, it can be seen that the algorithm of this application can achieve real-time tracking of moving targets.
[0191] (V) Comparative experiments
[0192] Through the above series of qualitative experiments, it can be concluded that the algorithm of this application has good robustness in terms of scale evolution, illumination change, and occlusion problems. The following experiment is to compare with the sum of squares and weak deviation algorithm in terms of reliability and speed to analyze the advantages and disadvantages of the improved algorithm compared with the sum of squares and weak deviation algorithm in tracking. Figure 6 The figure shows the test comparison of different tracking algorithms for the same data set. On the left is the test result of the sum of squares and weak deviation, and on the right is the test result of the algorithm of this application.
[0193] Through Figure 6 It can be seen that during the movement of the vehicle being tracked, the appearance of the vehicle is very blurred, and its scale continuously changes in the image plane. During the tracking process, the sum of squares and weak deviation has an obvious drift, while the algorithm of this application does not have the situation of tracking drift. Through comparison, it can be seen that the tracking effect of the improved algorithm of this application is significantly better than that of the sum of squares and weak deviation algorithm, and the tracking result is also very stable and has strong real-time performance.
[0194] Figure 7 The comparison results shown in the figure are even more obvious. The scale evolution of the vehicle in the data is large. Although the sum of squares and weak deviation can track the vehicle, as the scale becomes larger, the sum of squares and weak deviation algorithm only tracks a part of the vehicle, and the visual effect is obviously not as good as that of the algorithm of this application. It can be hypothesized that if this part is suddenly blocked by an object, it will cause the sum of squares and weak deviation tracking to fail, while the tracking algorithm of this application will not fail.
Claims
1. A real-time detection and tracking method for moving targets in vehicle assisted driving, characterized in that, Based on convolution filtering, a minimum output square deviation and sum of squares weak deviation filtering method is established. On the basis of the sum of squares weak deviation algorithm, the change of the real-time scale of the target is taken into account. A scale evolution factor is introduced into the target model, and the fHOG motion feature is used to replace the gray feature in the sum of squares weak deviation to characterize the texture information and edge information of the target. The specific method is to first use the improved sum of squares weak deviation function to quickly estimate the position of the target in the new frame, then calculate the scale evolution factor of the target in the new frame, and finally calculate and update the filter and the scale evolution factor uniformly. By updating the model appearance, center position and scale evolution factor of the target online, the change of the model scale is resisted; The following improvement strategies are adopted for the sum of squares weak deviation algorithm with low feature dimension and no consideration of the change of the real-time scale of the target: 1) Obtain high-dimensional features based on the fHOG transform: fHOG is obtained by optimizing the HOG speed, maintaining optical and geometric invariance, and fHOG is used as the representation of the target feature; 2) Introduce a scale evolution factor into the sum of squares weak deviation algorithm: First, use the initial sum of squares weak deviation algorithm to track the center point of the target in the new frame, calculate the scale evolution factor of the target in the new frame, and jointly update the filter and the scale evolution factor to ensure the algorithm speed while solving the influence of scale evolution on target tracking; The improved algorithm based on the real-time nature of target tracking is divided into two steps: First, use the improved sum of squares weak deviation function to quickly estimate the position of the target in the new frame to achieve the tracking of the moving target position, then calculate the scale evolution factor of the target in the new frame, calculate the fHOG motion feature, and finally jointly update the filter and the scale evolution factor to achieve the real-time scale estimation of the target.
2. The real-time detection and tracking method of a moving target for vehicle assisted driving according to claim 1, characterized in that Tracking of the moving target position: The error function of the sum of squares weak deviation filtering tracking algorithm is as follows: d is the feature dimension, represents the circular convolution operation, and the feature map is f l (l = 1, 2, 3, …, d), h l (l = 1, 2, 3, …, d) represents the correlation filter under the corresponding feature dimension, which is obtained by solving through the error function formula 9. g is the expected output, f 1 , h 1 and g have the same dimension and the same size, but simply relying on formula 9 to solve h requires a division operation, and formula 9 needs to be improved: By adding the λ term, the case of division by zero in the solution process is excluded. In addition, the change range of the filter parameters is also controlled. The smaller the λ, the larger the change range of the filter parameters. Whether to add the λ term has little influence on the initial frame perturbation. Finally, the λ term is not added; Equation 10 is a real convex function, and there is only one optimal solution, which is the point where the derivative is 0. Perform Fourier transform on Equation 10, take the derivative, and set the derivative to 0. After transformation, it is as follows: All are capital letters in Equation 11, indicating that all functions in the spatial domain have been transformed into the frequency domain. Indicates the conjugate complex number of the frequency domain matrix G obtained after the Fourier transform of the expected output matrix g. Indicates the fHOG dynamic feature f of the k-th dimension of the target region. k The conjugate complex number of the k-th dimensional feature F in the frequency domain obtained after the Fourier transform. In the actual tracking process, the following equation is used to update the model online: k In the actual tracking process, the following equation is used to update the model online: n is the learning rate, and its value range is [0,1]. The larger the learning rate, the more the model can learn the new appearance of the moving target. If it is 1, it means that the model represents the target with the latest appearance. If it is 0, it means that the model is not updated from beginning to end. The learning rate of this application is most suitable to take 0.015; After updating the filter, the following formula is used to predict the center position of the target in the new frame: Z l It is the feature map of the l-th dimension of the fHOG motion feature calculated for the target candidate region of the latest frame, and the position of the maximum value of y is determined to be the position of the center point of the target region in the (t + 1)-th frame.
3. The vehicle assisted driving moving target real-time detection and tracking method according to claim 1, characterized in that Calculating the fHOG motion feature: Calculate the gradient histogram of the local area of the vehicle-mounted camera image, and then calculate the calculated gradient histogram as a feature operator representing the moving target, that is, first divide the target image into small cell units, and then calculate the gradient magnitude and gradient direction histogram of each pixel point in each cell unit. The calculation of the vehicle-mounted moving image gradient histogram is the final fHOG motion feature; The specific calculation steps of the fHOG motion feature operator are as follows: Step 1: Select the vehicle assisted driving moving window or target area to be detected; Step 2: If the input vehicle-mounted image is a color image, first grayscale the image, and then normalize the color space of the grayscale image using the Gamma correction method to adjust the contrast of the image, reduce the impact of local shadows and lighting changes in the image, and at the same time resist noise interference; Step 3: Calculate the gradient magnitude and direction of each pixel point in the vehicle assisted driving moving target area; Step 4: Divide the image into small cell units cel1, with 5*5 pixels in one cell unit; Step 5: Calculate the gradient histogram of each cell unit; Step 6: A certain number of cell units form a block. 4*4 cell units form a block. The target consists of fHOG motion feature descriptors of multiple blocks. The feature descriptor of a block is calculated from the feature descriptors of all cell units within the block; Step 7: The target fHOG motion feature obtained by combining the fHOG motion feature descriptors of all blocks in the vehicle-mounted image.
4. The real-time detection and tracking method of a moving target for vehicle assisted driving according to claim 1, characterized in that, Target real-time scale estimation: First obtain the latest position, take the scale evolution factor into account, first obtain the center point of the target, then calculate the scale evolution factor of the target in the new frame, and finally uniformly calculate and update the filter and the scale evolution factor. The time complexity is smaller and the real-time performance is more guaranteed. The algorithm steps include: ① Input, ② Position estimation, ③ Scale estimation, ④ Update model.
5. The method for real-time detection and tracking of moving targets for vehicle assisted driving according to claim 4, wherein ① Input: The t-th frame image l t ; Previous frame position p t-1 and scale evolution factor s t-1 ; Central position model filter 6. The real-time detection and tracking method for moving targets in vehicle assisted driving according to claim 5, characterized in that ② Position estimation: The first step: in the t-th frame image I t obtain the candidate regions of the target according to the position p t-1 where the target center was located in the previous frame and the real-time scale evolution factor s t-1 of the target in the previous frame, and obtain the fHOG motion feature map z of the target through calculation t trans ; Step 2: Feature map z within the area where the estimated target position is located t trans is convolved with the position filter to obtain the alternative matrix y of the latest position t trans ; Step 3: Set the point with the maximum value of y t trans as the latest center position p of the moving target in the current frame t .
7. The method for real-time detection and tracking of moving targets for vehicle assisted driving according to claim 6, wherein ③ Scale estimation: Step 4: In the t-th frame image I t according to the latest position p t and the scale evolution factor s of the previous frame t-1 delimit the region where the target is located, obtain S regions through scaling, and extract the feature maps z of the S regions t scale ; Step 5: Perform correlation calculation on the S regional feature maps z within the area where the estimated target position is located t scale and the filter to obtain y t scale ; Step 6: y t scale Set the maximum value to the target real-time scale evolution factor s of the current frame t .
8. The real-time detection and tracking method of a moving target for vehicle assisted driving according to claim 7, characterized in that, ④ Update model: Step 7: In the t-th frame image I t , according to the target center position p t estimated in Steps ② and ③ and the scale evolution factor s t , delimit the region where the target is located, extract the features of the target, and obtain the training sample f after Fourier transform t trans ; Step 8: Update the position filter 1) The motion of the target is smooth in a very short time. Obtain the approximate center position of the target in the current frame through the state of the previous frame; 2) The scale evolution of the target changes little in a very short time. Obtain the area of the target in the current frame through the scale evolution factor of the previous frame and the latest center position of the target.
9. The method for real-time detection and tracking of moving targets for vehicle assisted driving according to claim 1, wherein Sum of Squares Weak Deviation Filter Tracking: The convolution calculation in the spatial domain of a digital image is equivalent to the element-by-element multiplication between two matrices after the image is transformed into the frequency domain, saving the computational time of the convolution calculation in the spatial domain. Assume that the training samples are a series of digital images f i , which becomes F after Fourier transform i . Manually given the output result G i , the filter H is obtained i * as follows: The final filter H is optimized by minimizing the output mean square sum deviation: where ⊙ represents element-by-element multiplication between matrices, the elements of the filter H are calculated at the corresponding positions in the frequency domain space, and any one element in the filter H is solved separately. Assume H wv is an independent element in H, and the filter is obtained by solving H wv separately. Equation 2 is converted into the following form: The above function is a real function and a convex function, with only one optimal point. To find the optimal point, only need to take the derivative of the function. The point where the derivative function is 0 is the optimal point of the function, that is: Obtain H wv Expression of: After separately obtaining the elements in the filter H, obtain all the elements in the filter H, and solve to obtain the filter H.
10. The method for real-time detection and tracking of moving targets for vehicle assisted driving according to claim 1, wherein Multi-filter online update change tracking: The moving target continuously changes in the actual scene image, including scene illumination changes, scale evolution, rotation changes, pose changes, and target position changes. Real-time multi-filter online update is used to adapt to the above changes. The following method is used to update the model: where \(t\) represents the number of frames, \(F\) t * represents the matrix obtained after the Fourier transform of the image within the target area of the \(t\)-th frame, \(G\) t represents the desired output matrix \(G\) t * is the conjugate complex number of \(G\), \(F\) t represents \(F\) t * is the conjugate complex number of \(F\), \(\eta\) represents the learning rate, and the learning rate is taken as 0.125 to make the updating speed of the filter keep up with the appearance change of the target; The filter H for the t-th frame is obtained by solving t , for the (t + 1)-th frame, calculate its feature map Z, and perform a correlation convolution on H t and Z. In the finally output result, the point where the peak value is located is the center position of the target to be tracked; The position of the maximum value of y is determined to be the center point of the target in the (t + 1)-th frame image; Based on the convolutional filter tracking algorithm, the peak-to-sidelobe ratio PSR is used to represent the tracking situation of the target, and the definition is as follows: The steepness of the two-dimensional Gaussian distribution surface affects the tracking accuracy. The tracking situation of the target is characterized by the above formula. μ and σ represent the mean and variance of the pixels excluding the peak point within the 11*11 window centered on the peak.