Video image target identification and tracking method and system, storage medium and equipment
By extracting and fusing fast and slow time-domain features from video image sequences, and combining feature enhancement and motion parameter prediction, the problem of target feature degradation in video target tracking is solved, achieving higher tracking accuracy and stability.
Patent Information
- Application Number
- CN202511287529.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-12-30
AI Technical Summary
Existing video image target tracking technologies are prone to feature degradation when the target object moves and deforms, leading to inaccurate tracking, a problem that existing technologies cannot effectively solve.
This method involves extracting fast and slow temporal features from video image sequences, combining feature fusion and enhancement processing, and using feature enhancement parameters and motion parameters to predict the target's search region.
It improves the accuracy of video target tracking and avoids the problems of target tracking divergence or loss.
Smart Images

Figure CN121236710A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a method, system, storage medium, and device for identifying and tracking targets in video images. Background Technology
[0002] With the rapid development of computer vision technology, video image target tracking technology has been widely used in fields such as intelligent surveillance, human-computer interaction, and autonomous driving. The core of video image target tracking technology is to locate and track the position of a specific target object in real time within a continuous video sequence.
[0003] Currently, the common technical approach is feature extraction and target matching. Specifically, feature information of the target object is first extracted from the video sequence, and then the region that best matches that feature is searched in subsequent frames to achieve continuous target tracking. However, in practical applications, due to interference factors such as the target object's own deformation, occlusion, and complex backgrounds—for example, when the target object moves rapidly or undergoes drastic deformation—the extracted features are prone to degradation, leading to a decrease in the target feature representation ability. This can easily cause target tracking divergence or loss, severely affecting the accuracy of video target tracking. Summary of the Invention
[0004] This application provides a method, system, storage medium, and device for identifying and tracking video image targets, which can improve the accuracy of video target tracking.
[0005] In a first aspect, this application provides a method for identifying and tracking targets in video images, the method comprising: A video image sequence is acquired, and features are extracted from the video image sequence in both fast and slow time domains. The extracted features are then analyzed to obtain a feature sequence corresponding to at least one target object. The feature enhancement parameters are determined based on the distribution parameters of the target object in the feature sequence; The target object is enhanced based on the feature enhancement parameters to obtain the enhanced target sequence. The motion parameters of the target object in the current tracking period of the target sequence for multiple consecutive frames are obtained, and the search area for the next tracking period is determined based on the motion parameters. The target object is matched for features in the search area, and the target area with the highest matching degree is taken as the latest position of the target object in the next tracking cycle, so as to realize continuous tracking of the target object.
[0006] By adopting the above technical solution, firstly, by extracting features in both the fast and slow time domains of the video image sequence and performing correlation analysis, the instantaneous change features and stable structural features of the target object can be obtained simultaneously, thereby improving the completeness of feature representation. Then, based on the distribution parameters of the target object, feature enhancement parameters are determined, and the target object is enhanced, which can effectively improve the expressive power of the target features and reduce the impact of feature degradation when the target moves rapidly or undergoes drastic deformation. Finally, by obtaining the motion parameters of the target object in the current tracking cycle, the search area for the next tracking cycle is determined, and feature matching is performed within this area, which can accurately predict the target's motion trend, improve the accuracy of target tracking, and effectively avoid the problem of target tracking divergence or loss.
[0007] A second aspect of this application provides a video image target recognition and tracking system, the system comprising: The feature sequence determination module is used to acquire video image sequences, extract features from the video image sequences in both fast and slow time domains, and perform correlation analysis on the extracted features to obtain a feature sequence corresponding to at least one target object; the enhancement parameter determination module is used to determine feature enhancement parameters based on the distribution parameters of the target objects in the feature sequences. The target sequence determination module is used to enhance the target object based on the feature enhancement parameters to obtain the enhanced target sequence; The search area determination module is used to obtain the motion parameters of the target object in multiple consecutive frames within the current tracking period of the target sequence, and determine the search area for the next tracking period based on the motion parameters. The target tracking and matching module is used to perform feature matching on the target object in the search area, and take the target area with the highest matching degree as the latest position of the target object in the next tracking cycle, so as to realize continuous tracking of the target object.
[0008] A third aspect of this application provides a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method steps described above.
[0009] A fourth aspect of this application provides an electronic device, comprising: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the above-described method steps.
[0010] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: First, by extracting features from video image sequences in both the fast and slow time domains and performing correlation analysis, this application can simultaneously obtain the instantaneous change features and stable structural features of the target object, thereby improving the completeness of feature representation; then, based on the distribution parameters of the target object, feature enhancement parameters are determined, and the target object is enhanced, which can effectively improve the expressive power of the target features and reduce the impact of feature degradation when the target moves rapidly or undergoes drastic deformation; finally, by obtaining the motion parameters of the target object in the current tracking cycle to determine the search area for the next tracking cycle and performing feature matching within that area, the motion trend of the target can be accurately predicted, improving the accuracy of target tracking and effectively avoiding the problem of target tracking divergence or loss. Attached Figure Description
[0011] Figure 1 This is a flowchart illustrating a method for identifying and tracking a target in a video image, as provided in an embodiment of this application. Figure 2 This is a schematic diagram illustrating the recognition effect of a target object according to an embodiment of this application; Figure 3 This is a schematic diagram of target tracking under occlusion conditions provided in an embodiment of this application; Figure 4 This is a schematic diagram of a video image target recognition and tracking system provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0012] Explanation of reference numerals in the attached drawings: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Implementation
[0013] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0014] In the description of the embodiments in this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Specifically, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.
[0015] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0016] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0017] The video image target recognition and tracking method provided in this application can be applied to a variety of practical scenarios. For example, in intelligent monitoring systems, it can be used to track the movement trajectory of suspicious persons or vehicles within the monitored area; in intelligent transportation systems, it can be used to track the movement status of specific vehicles in traffic flow; and in UAV visual navigation, it can be used to track specific targets to achieve autonomous following, etc.
[0018] As an optional scenario, this method can be applied to video target tracking scenarios for UAV electro-optical pods. Specifically, the video acquisition equipment of the UAV electro-optical pod needs to acquire images for target recognition and tracking of the identified targets during mission execution. In this application scenario, due to the flight motion of the UAV itself and changes in the motion state of the target object, the target object in the acquired video image may exhibit complex situations such as rapid movement and drastic deformation. For example, when the UAV tracks a moving vehicle on the ground, changes in the UAV's heading and the vehicle's turning acceleration will cause significant changes in the target features in the video image. In this specific application scenario, the video image target recognition and tracking method provided in this application embodiment, through the combination of fast and slow time-domain features, feature enhancement, and adaptive search region techniques, can effectively cope with various complex situations encountered by the UAV electro-optical pod during target tracking, ensuring the accuracy and stability of tracking. Subsequent embodiments will use this scenario as an example to explain in detail the specific implementation methods of each technical feature of this application.
[0019] Please refer to Figure 1This paper presents a flowchart illustrating a method for identifying and tracking targets in video images. This method can be implemented using a computer program, a microcontroller, or run on a video image target identification and tracking system. The computer program can be integrated into a computer device or run as a standalone application. Specifically, the method includes steps 10 to 50, as follows: Step 10: Obtain the video image sequence, perform fast time domain and slow time domain feature extraction on the video image sequence respectively, and perform correlation analysis on the extracted features to obtain the feature sequence corresponding to at least one target object.
[0020] In this embodiment of the application, a video image sequence refers to a set of images consisting of multiple consecutive frames arranged in chronological order, with each frame recording the visual information of the scene at a specific moment.
[0021] A feature sequence is a sequence of feature information extracted from a target object in multiple consecutive frames of images, arranged in chronological order. It includes a combination of fast temporal features (such as instantaneous motion features) and slow temporal features (such as stable structural features) of the target object. This feature sequence reflects the feature change process of the target object in the video sequence.
[0022] Please see Figure 2 This is a schematic diagram illustrating the recognition effect of a target object according to an embodiment of this application; as shown below. Figure 2 As shown, this image is any frame in a video image sequence and can include multiple target objects, which can be various vehicles.
[0023] Specifically, in the video target tracking scenario of the UAV electro-optical pod, the first step is to acquire a sequence of video images captured by the video acquisition device of the electro-optical pod. Since the target object exhibits rapid movement and deformation characteristics in the video during UAV flight, while also possessing relatively stable structural features, it is necessary to extract features from both the fast and slow time domains to comprehensively describe the target object's characteristic information.
[0024] Furthermore, foreground segmentation is first performed on the video image sequence to extract at least one candidate target region containing the target object. For each candidate target region, to obtain the fast motion features of the target object, difference images of adjacent frames in the video image sequence are acquired, and the instantaneous motion features of the target object are extracted as fast temporal features based on the gradient changes in the difference images. These fast temporal features can promptly reflect the dynamic changes of the target object caused by the drone's movement or its own movement. Simultaneously, to obtain the stable structural features of the target object, multi-directional edge detection is performed on the target object within a preset time window to obtain the edge response spectrum of the target object. The stable structural features of the target object are extracted as slow temporal features based on the temporal cumulative value of the edge response spectrum. These slow temporal features can characterize the relatively stable structural information of the target object during rapid movement.
[0025] Because fast time-domain features are highly transient while slow time-domain features are more stable, a correlation analysis between these two types of features is necessary. Specifically, feature fusion weights are determined based on the transient nature of fast time-domain features and the stability of slow time-domain features. The fast and slow time-domain features are then weighted and combined according to their respective feature fusion weights to obtain the feature sequences corresponding to each target object. This feature extraction and fusion method can capture the dynamic feature changes of the target object caused by the drone and its own movement in a timely manner, while also including the relatively stable structural information of the target, thus providing a more reliable feature foundation for subsequent target tracking.
[0026] Based on the above embodiments, as an optional embodiment, the step of performing fast time-domain and slow time-domain feature extraction on the video image sequence, and performing correlation analysis on the extracted features to obtain a feature sequence corresponding to at least one target object, further includes the following steps: Step 101: Perform foreground segmentation on the video image sequence and extract at least one candidate target region containing the target object.
[0027] Specifically, in the video image sequence acquired by the UAV's electro-optical pod, to accurately locate the target object region, foreground segmentation is first required. For each frame, a mixed background model containing 3-5 Gaussian distributions is established, each with three parameters: mean, variance, and weight. For each pixel in the image, its matching degree with each Gaussian distribution is calculated (the difference between the pixel value and the Gaussian distribution mean does not exceed 2.5 times the standard deviation), and the corresponding model parameters are updated. Then, the Gaussian distributions are sorted according to the ratio of weight to standard deviation, and the top few distributions with a cumulative weight exceeding 0.7 are selected as the background model. By comparing the current frame image with the background model, pixels that do not match the background model are marked as foreground. The obtained foreground region is then subjected to erosion using a 3×3 structuring element to remove noise, followed by dilation using a 5×5 structuring element to fill holes. Finally, an 8-connected component labeling algorithm is used for connected component analysis to segment the foreground region into multiple independent candidate target regions. Finally, based on preset target feature requirements (e.g., area between 100-1000 pixels, aspect ratio between 0.5-2), at least one candidate target region containing the target object is obtained.
[0028] Step 102: For each candidate target region, obtain the difference image of adjacent frames in the video image sequence, and extract the instantaneous motion features of the target object as fast temporal features based on the gradient changes in the difference image.
[0029] Specifically, for each candidate target region, a difference image is first obtained by calculating the difference between corresponding pixels in two adjacent frames. Then, a 3×3 Sobel operator is used to calculate the gradient values in the horizontal and vertical directions, and the gradient magnitude and orientation angle of each pixel are calculated based on these gradient values. The orientation angle is uniformly divided into 8 intervals (each interval is 45 degrees), and a gradient histogram is constructed by statistically analyzing the gradient orientation distribution and then normalized. Features such as the main motion direction (the orientation angle corresponding to the peak of the histogram), motion intensity (peak size), and motion complexity (the entropy value of the histogram) are extracted from the normalized histogram. At the same time, 9 feature points are uniformly sampled within the target region, and the optical flow vectors of these points are calculated to describe the local motion features. These motion-related features together constitute the instantaneous motion features of the target object, serving as fast temporal features that can effectively describe the instantaneous motion state of the target.
[0030] Step 103: Perform multi-directional edge detection on the target object within a preset time window to obtain the edge response spectrum of the target object, and extract the stable structural features of the target object as slow time domain features based on the temporal cumulative value of the edge response spectrum.
[0031] Specifically, to extract stable structural features of the target object, a preset time window (e.g., 15 consecutive frames) is first set. Within this time window, multi-directional edge detection is performed on the target object. A Gabor filter bank is used to filter the target region, with the filter directions including 0°, 45°, 90°, 135°, etc. For each direction, the amplitude of the filtered response is calculated to obtain the edge response in that direction. The edge responses of all directions are combined into an edge response spectrum. Then, the edge response spectrum is accumulated within the time window, and the accumulated response value in each direction is calculated. Finally, based on the distribution characteristics of the accumulated response values, the directions with response values greater than the preset response value and stable are extracted as the stable structural features of the target object, i.e., slow temporal features.
[0032] Step 104: Determine the feature fusion weights based on the instantaneous nature of the fast time-domain features and the stability of the slow time-domain features, and then combine the fast time-domain features and the slow time-domain features according to their respective feature fusion weights to obtain the feature sequences corresponding to each target object.
[0033] Specifically, to effectively fuse fast and slow temporal features, the rate of change of the fast temporal features between adjacent frames is first calculated (obtained by normalizing the Euclidean distance of the features), and the stability of the slow temporal features within the time window is calculated (represented by the reciprocal of the variance of the feature sequence). Then, the Softmax function is used to calculate the fusion weights, where the weight of the fast temporal feature is inversely proportional to its rate of change, and the weight of the slow temporal feature is directly proportional to its stability; the sum of the two weights is 1. By adjusting the balancing parameter in the weight calculation (usually the fast temporal feature parameter is set to 1, and the slow temporal feature parameter to 2), the relative importance of the two features can be controlled. Finally, the fast and slow temporal features are weighted and summed according to the calculated weights to obtain the fused features, and the fused features of consecutive frames are arranged in chronological order to form a feature sequence. This adaptive weighted fusion method based on the dynamic properties of features can automatically adjust the weight ratio of fast and slow temporal features according to the target's motion state, preserving both the instantaneous features of the target motion and containing stable structural information, thus improving the robustness of feature representation.
[0034] Step 20: Determine the feature enhancement parameters based on the distribution parameters of the target object in the feature sequence.
[0035] In this embodiment, the feature enhancement parameter refers to a parameter obtained by combining a contrast adjustment coefficient and a texture enhancement coefficient. The contrast adjustment coefficient is used to adjust the degree of contrast enhancement of the target feature, and the texture enhancement coefficient is used to adjust the degree of texture detail enhancement of the target feature. The combination of these two coefficients constitutes the feature enhancement parameter, which is used for subsequent enhancement processing of the target feature.
[0036] Specifically, in the video target tracking scenario of UAV optoelectronic pods, since the contrast and texture features of the target object in the image will change with the flight state of the UAV and environmental conditions, it is necessary to dynamically determine the contrast adjustment coefficient and texture enhancement coefficient based on the distribution parameters of the target object in the feature sequence, and combine them to form feature enhancement parameters.
[0037] First, statistical analysis is performed on the target object in the feature sequence, calculating the gray-level distribution histogram of the target region to obtain statistical characteristics such as the average gray-level value, standard deviation, and gray-level distribution skewness. Based on these statistical characteristics, the contrast difference between the target region and the background region is calculated, and the contrast adjustment coefficient is determined based on this difference value. Simultaneously, the texture complexity of the target region is analyzed, including calculating features such as local gradient changes and edge density, and the texture enhancement coefficient is determined based on these features. Finally, the contrast adjustment coefficient and the texture enhancement coefficient are combined to obtain the feature enhancement parameters. The feature enhancement parameters determined in this way consider both the enhancement requirements of target contrast and the degree of enhancement of target texture details, providing more comprehensive guidance for subsequent feature enhancement processing and improving the quality of target features.
[0038] Based on the above embodiments, as an optional embodiment, the step of determining the feature enhancement parameters according to the distribution parameters of the target object in the feature sequence further includes the following steps: Step 201: Calculate the contrast distribution and texture complexity distribution of the target object in the feature sequence.
[0039] Specifically, in the video images acquired by the UAV's electro-optical pod, the contrast distribution and texture complexity distribution of the target object in the feature sequence are calculated. For the target region R and its surrounding background region B in each frame of the image, the target-background contrast value Ci is calculated as follows: Where, μ R and μ B These are the average gray values of the target area and the background area, respectively. and Let these be the variances. Statistical analysis of the contrast values of N consecutive frames yields the contrast distribution characteristics: Mean contrast: Contrast variance: Contrast skewness: Where N represents the number of frames and Ci represents the contrast value of the i-th frame, these statistical features together describe the distribution characteristics of the target contrast.
[0040] Simultaneously, to evaluate the texture features of the target region, a Gabor filter is applied to the target region of each frame image. The Gabor filter is as follows: Perform multi-scale, multi-directional decomposition to obtain the texture complexity value of the i-th frame image. Where: (x,y) represents the pixel coordinate position in image space, R(x,y) represents the Gabor filter response value at pixel position (x,y), x'=xcosθ+ysinθ represents the rotated X coordinate, y'=-xsinθ+ycosθ represents the rotated Y coordinate; λ represents the wavelength of the filter, θ represents the direction angle of the filter, ψ represents the phase shift, σ represents the standard deviation of the Gaussian envelope, used to control the spatial range of the filter, and γ represents the spatial aspect ratio, used to control the ellipticity of the filter.
[0041] Step 202: Determine the contrast adjustment coefficient based on the contrast distribution, and determine the texture enhancement coefficient based on the texture complexity distribution.
[0042] Specifically, the contrast adjustment coefficient α is determined based on the contrast distribution. Considering the three statistical characteristics of contrast—mean, variance, and skewness—a composite adaptive mapping method is adopted: in, The desired contrast ratio is set to 1.0; k1 is the basic adjustment rate parameter, set to 2.0; λ1 is the variance weighting coefficient, typically set to 0.5; λ2 is the skewness weighting coefficient, set to 0.3. The first term in this formula... The baseline adjustment coefficient is determined based on the difference between the mean contrast value and the expected value; (Second item) Reduce the adjustment range when contrast fluctuates significantly; third item Compensation adjustment is performed based on the distribution skewness. By setting parameters k1, λ1, and λ2, the value of α is ensured to be within the range of [0.5, 2.0].
[0043] Furthermore, the texture enhancement coefficient β is determined based on the texture complexity value: Where T0 is the desired texture complexity value, set according to the specific application scenario; k2 is the texture adjustment rate parameter, set to 1.5; and Ti is the texture complexity value of the i-th frame image. By setting the parameter k2, the value of β is ensured to be between [0.3, 1.5]. When the actual texture complexity T is lower than the desired value T0, the value of β is increased to enhance texture details; conversely, the enhancement level is reduced.
[0044] Step 203: Combine the contrast adjustment coefficient and the texture enhancement coefficient to obtain the feature enhancement parameter, which is used to enhance the contrast and texture features of the target object.
[0045] Specifically, the contrast adjustment coefficient α and the texture enhancement coefficient β are used as two components of the feature enhancement parameters to obtain the feature enhancement parameters. The contrast adjustment coefficient, with a value range of [0.5, 2.0], is used to adjust the low-frequency components of the target image; the texture enhancement coefficient, with a value range of [0.3, 1.5], is used to adjust the high-frequency components of the target image.
[0046] Step 30: Enhance the target object based on the feature enhancement parameters to obtain the enhanced target sequence.
[0047] In this embodiment, the target sequence refers to a continuous image sequence obtained after feature enhancement processing of target objects in the feature sequence. Specifically, each target object, after contrast adjustment and texture enhancement processing, forms an enhanced target image region, which contains the enhanced image content of the target object. For example, for N target objects in the feature sequence, the target sequence contains these N feature-enhanced target image regions, which are arranged in chronological order to form a dynamically changing sequence.
[0048] Specifically, the process first involves acquiring target objects from the feature sequence. Each target object contains feature information obtained through a weighted combination of fast and slow temporal features. Then, the low-frequency components of the target objects are mapped using the contrast adjustment coefficient α in the feature enhancement parameters. When α is greater than 1, the mapping range of the low-frequency components is expanded to enhance the contrast between the target and the background; when α is less than 1, the mapping range of the low-frequency components is compressed to reduce excessive contrast. Simultaneously, the high-frequency components of the target objects are gain-adjusted using the texture enhancement coefficient β in the feature enhancement parameters. When β is greater than 1, the gain parameter of the high-frequency components is increased to enhance the texture details and edge features of the target; when β is less than 1, the gain parameter of the high-frequency components is decreased to suppress noise and pseudo-textures. This frequency-division enhancement method processes each target object in the feature sequence, ultimately yielding the enhanced target sequence. This enhancement method maintains the basic structure of the target features while improving the discriminative power of the features through the synergistic enhancement of contrast and texture, resulting in better feature representation of the target objects in the target sequence.
[0049] Based on the above embodiments, as an optional embodiment, the step of enhancing the target object based on feature enhancement parameters to obtain the enhanced target sequence further includes the following steps: Step 301: Decompose the target object into low-frequency components that characterize the overall features of the target and high-frequency components that characterize the detailed features of the target.
[0050] Specifically, in order to separately enhance the features of the target object at different scales, it is necessary to first decompose the target object into low-frequency components and high-frequency components. Specifically, the wavelet transform method is used to perform three-layer wavelet decomposition on the target object. In wavelet decomposition, the db4 wavelet basis function is selected, which has good orthogonality and compact support, and is suitable for multi-resolution analysis of images. Through wavelet decomposition, low-frequency approximation coefficients (LL3) and high-frequency detail coefficients (LH3, HL3, HH3) can be obtained. Among them, the low-frequency approximation coefficient LL3 reflects the contour and overall brightness distribution characteristics of the target, and serves as the low-frequency component representing the overall characteristics of the target; the high-frequency detail coefficients in the three directions of LH3, HL3, and HH3 are combined as the high-frequency component representing the detail characteristics of the target. This decomposition method based on wavelet transform can effectively separate the target features in different frequency domains, laying a foundation for subsequent frequency-divided enhancement processing.
[0051] Step 302: Determine the gray-scale mapping interval of the low-frequency component according to the contrast adjustment coefficient, and perform non-linear contrast enhancement processing on the low-frequency component.
[0052] Specifically, in order to enhance the contrast between the target and the background, non-linear contrast enhancement processing is performed on the low-frequency component. First, determine the gray-scale mapping interval [L min , L max according to the contrast adjustment coefficient α, and its calculation method is: L min = μ L - α·σ L ; L max = μ L + α·σ L ; where μ L and σ L are the mean and standard deviation of the low-frequency component respectively. When α > 1, the mapping interval is expanded to enhance the contrast; when α < 1, the mapping interval is compressed to avoid over-enhancement. Then, use the S-shaped non-linear mapping function to enhance the low-frequency component: for any pixel value X in the low-frequency component, when x < Lmin, L'(x) = Lmin; when Lmin ≤ x ≤ Lmax, L'(x) = Lmin + (Lmax - Lmin) / (1 + e^(-k(x - μL))); when x > Lmax, L'(x) = Lmax; where k is the curve slope parameter, which is dynamically adjusted by α, k = α·4 / (Lmax - Lmin). This non-linear mapping method can adaptively enhance the contrast of the target area while maintaining the continuity of the image gray scale, avoiding the problem of detail loss that may be caused by traditional linear mapping. When the brightness of the image area is close to the mean, the enhancement effect is relatively mild; when the brightness difference is large, the enhancement effect is more obvious, thus achieving adaptive enhancement of the contrast.
[0053] Step 303: Determine the gain parameters of the high-frequency components based on the texture enhancement coefficient, and perform gain processing on the high-frequency components.
[0054] Specifically, adaptive gain processing is applied to the high-frequency components to enhance the texture and edge features of the target. First, the gain parameter g = β·(1+exp(-σH)) is determined based on the texture enhancement coefficient β, where σH is the standard deviation of the high-frequency components. When β > 1, the gain parameter is increased to enhance texture details; when β < 1, the gain parameter is decreased to suppress noise. Then, gain processing is applied to each directional component of the high-frequency components separately: H' = g·H·w(H), where H is the original high-frequency coefficient, and w(H) is the weighting function used to suppress noise. The weighting function is defined as: w(H) = exp(-|H| / T), where T is the adaptive threshold, determined based on the local variance of the high-frequency components. This gain processing method can adaptively adjust the enhancement level according to the local characteristics of the texture, effectively improving the detail representation of the target.
[0055] Step 304: Calculate the fusion weights of the processed low-frequency components and high-frequency components respectively, and combine the processed low-frequency components and high-frequency components according to the fusion weights to obtain the enhanced target sequence.
[0056] Specifically, to properly fuse the processed low-frequency and high-frequency components, it is necessary to calculate and combine the fusion weights of each component. First, the initial weights are calculated based on the local standard deviations of the low-frequency component L' and the high-frequency component H': Where σ L′ and σ H′ The local standard deviations for the low-frequency and high-frequency components are given, respectively. Then, considering the local brightness variations in the target region, a brightness adjustment factor is introduced to correct the initial weights: w L =w L0 ·(1+exp(-|μ L′ -μ0|)); w H =w H0 ·(1+exp(-|μ H′ |));where μ L′ The local mean of the low-frequency components is μ0, and the desired brightness value is μ0. H′ This represents the local mean of the high-frequency components. Based on the corrected fusion weights, a weighted combination method is used to obtain the enhanced target object: I′=w L ·L′+w H•H′. The above processing is performed on each target object in the feature sequence. The processed target objects are then reorganized in chronological order to obtain the enhanced target sequence. This fusion method can effectively balance the contributions of low-frequency and high-frequency information, so that the enhanced target sequence maintains both contrast enhancement and rich texture details, thereby improving the feature representation ability of the target sequence.
[0057] Step 40: Obtain the motion parameters of the target object in multiple consecutive frames within the current tracking period of the target sequence, and determine the search area for the next tracking period based on the motion parameters.
[0058] In this embodiment, the current tracking cycle refers to the time period from acquiring a frame of image to completing target detection and tracking for that frame and predicting the target position for the next frame. For example, for a video sequence with a frame rate of 30fps, each tracking cycle is approximately 1 / 30 of a second.
[0059] The search region refers to the area within the current image frame centered on the predicted target location, determined based on motion parameters. This region is typically set to a size 2-3 times the target size and is used to search for and locate the target in the next image frame. Setting a suitable search region can effectively reduce the detection range and improve algorithm efficiency.
[0060] Specifically, within the current tracking cycle, the motion parameters of the target object in N consecutive frames (N≥3) are first acquired, including the target center coordinates, motion velocity, and scale change rate. Then, the Kalman filtering method is used to model and predict these motion parameters. A target motion state model can be constructed, using the target's position, velocity, and scale change rate as state variables. A state transition equation describes the target's motion law, and an observation equation reflects the relationship between actual measurements and state variables. By filtering and calculating the motion parameters of the N consecutive frames, the predicted state for the next frame is obtained, including predicted position, predicted velocity, and predicted scale. Based on the predicted state, the search area for the next tracking cycle is determined: centered on the predicted position, the size of the search area is adaptively adjusted according to the predicted velocity and predicted scale. When the predicted velocity is large, the search area is appropriately expanded to cope with rapid target movement; when the predicted scale changes significantly, the search area size is adjusted accordingly to adapt to target scale changes; when the target motion is stable, a smaller search range is maintained.
[0061] Based on the above embodiments, as an optional embodiment, the step of obtaining the motion parameters of the target object in multiple consecutive frames within the current tracking period of the target sequence, and determining the search area for the next tracking period based on the motion parameters, further includes the following steps: Step 401: Extract the motion trajectory of the target object in multiple consecutive frames, and calculate the displacement vector sequence between adjacent sampling points in the motion trajectory.
[0062] Specifically, to accurately analyze the motion characteristics of a target, it is necessary to extract the target's motion trajectory in consecutive image frames. First, the center position coordinates (xi, yi) of the target object are recorded in N consecutive frames (N≥5), where i represents the frame number. The displacement vector between two adjacent frames can be represented as vi = (xi+1-xi, yi+1-yi). By calculating N-1 displacement vectors, a displacement vector sequence V = {v1, v2, ..., vN-1} is obtained. This displacement vector-based representation method can intuitively reflect the target's motion characteristics, providing fundamental data for subsequent motion analysis.
[0063] Step 402: Calculate the tangent and normal directions of the motion trajectory at each sampling point based on the displacement vector sequence to obtain the motion direction sequence of the target object.
[0064] Specifically, to describe the change in the target's motion direction, it is necessary to calculate the tangent and normal directions of the trajectory at each sampling point. For each vector vi in the displacement vector sequence, its tangent direction θi can be calculated using the arctangent function: θi = arctan((yi+1-yi) / (xi+1-xi)). The normal direction φi is perpendicular to the tangent direction, i.e., φi = θi + π / 2. The tangent and normal directions of all sampling points are combined to form the motion direction sequence {(θ1,φ1),(θ2,φ2),...,(θN-1,φN-1)}. This direction representation method fully characterizes the target's motion trend and helps predict the target's turning behavior.
[0065] Step 403: Calculate the curvature value and angle change of the motion trajectory based on the motion direction sequence to obtain the steering characteristic parameters of the target object.
[0066] Specifically, to quantify the turning characteristics of a target, curvature values and angle changes need to be calculated based on the motion direction sequence. For curvature calculation, three consecutive sampling points on the trajectory are selected. First, the area of the triangle formed by these three points is calculated, then the pairwise distances between the three points are calculated. The curvature value at that point is obtained by dividing the triangle area by the product of its three side lengths and multiplying by a coefficient of 4, which effectively reflects the curvature of the trajectory. When calculating angle changes, the change in tangent direction at adjacent sampling points needs to be considered. For each sampling point, the angle between its tangent direction and the tangent direction of the next sampling point is calculated. To avoid periodic jumps in angle calculation, normalization is required when calculating the angle difference to ensure that the angle change always reaches its minimum value. For example, when the angles of two adjacent directions are 175 degrees and -175 degrees respectively, their actual angle difference should be 10 degrees instead of 350 degrees. The curvature values and angle changes are used as turning characteristic parameters of the target object.
[0067] Step 404: Calculate the shape parameters of the search area based on the steering feature parameters. The shape parameters include the major axis, minor axis, and center position of the search area.
[0068] Specifically, to construct a suitable search region, shape parameters need to be calculated based on the steering feature parameters. First, a baseline value is determined based on the current size of the target's bounding box. The baseline major axis ab is set as the diagonal length of the bounding box, and the baseline minor axis bb is the length of the shorter side of the bounding box. Then, the baseline values are adjusted according to the steering feature parameters. The specific calculation method is as follows: the major axis a of the search region is a = ab·(1+α·κm), where κm is the average curvature value, and α is the curvature adjustment coefficient, ranging from 1.2 to 1.5; the minor axis b of the search region is b = bb·(1+β·θt), where θt is the cumulative angle change, and β is the angle adjustment coefficient, ranging from 0.8 to 1.2. This adaptive size adjustment method ensures that the search region can dynamically change with the target's motion characteristics.
[0069] Furthermore, the center position of the search area is determined through motion prediction. Specifically, using the position data of the most recent N frames (N≥3), the coordinates (xi, yi) of the target center are recorded for each frame. Based on these historical coordinate data, the average velocities vx and vy in the x and y directions, and the average accelerations ax and ay, are calculated respectively. Then, the coordinates of the target center in the x and y directions in the next frame are predicted according to the kinematic formulas. For the x direction, the predicted coordinate is: the current x coordinate plus the product of the x-direction velocity and the frame interval, plus half of the product of the x-direction acceleration and the square of the frame interval; the predicted coordinate for the y direction is calculated similarly. The predicted x and y coordinates are used together to determine the center position of the search area.
[0070] Based on the above embodiments, as another optional embodiment, the step of calculating the shape parameters of the search region based on the steering feature parameters may further include the following steps: Step 4041: Calculate the ratio of the arc length to the chord length of the motion trajectory.
[0071] Specifically, to accurately describe the curvature characteristics of the target's trajectory, it is necessary to calculate the ratio of arc length to chord length. First, the arc length is calculated by accumulating the distances between adjacent sampling points on the trajectory. For two adjacent points Pi(xi,yi) and Pi+1(xi+1,yi+1), their straight-line distance is calculated as di, and the total arc length L is the sum of all di values. The chord length D is obtained by calculating the straight-line distance between the trajectory's starting point P1 and ending point Pn. Finally, the ratio of arc length to chord length, R = L / D, is calculated. This ratio reflects the curvature of the trajectory; the more curved the trajectory, the larger the ratio.
[0072] Step 4042: Determine the center position of the search area based on the ratio. The center position is located at the center of the arc of the motion trajectory.
[0073] Specifically, to ensure the search area better covers the target's potential movement range, the center position of the search area needs to be rationally determined. Based on the calculated arc length to chord length ratio R, a weighted average method is used to determine the arc center position. When R is close to 1, it indicates the trajectory is approximately a straight line, and the predicted target position is taken as the center of the search area. When R is significantly greater than 1, it indicates the trajectory has a large curvature, and the search area center needs to be shifted towards the trajectory arc center. The arc center position is calculated as follows: first, find the center of the arc formed by three adjacent points on the trajectory; then, calculate a weighted average of these center positions based on the R value, with the weight proportional to the local curvature. This arc center-based center position determination method can better adapt to the target's turning motion.
[0074] Step 4043: Calculate the major axis length of the search area based on the arc length, and calculate the minor axis length of the search area based on the chord length.
[0075] Specifically, to reasonably set the size of the search area, the lengths of the major and minor axes are determined using arc length and chord length, respectively. The length of the major axis of the search area should be related to the arc length of the trajectory to ensure that it can cover the entire range of motion of the target. Specifically, the length of the major axis is set to γ times the arc length, where γ is the major axis coefficient, with a value ranging from 1.2 to 1.5. The length of the minor axis is related to the chord length and is set to δ times the chord length, where δ is the minor axis coefficient, with a value ranging from 0.8 to 1.2. This size setting method based on motion trajectory features allows the shape of the search area to better match the motion characteristics of the target.
[0076] Step 4044: Perform scaling correction on the major axis length and minor axis length according to the current size of the target object, and use the corrected major axis, minor axis, and center position as shape parameters.
[0077] Specifically, to adapt to dynamic changes in target size, the calculated major and minor axis lengths need to be scaled. First, the current target size information is obtained, including the width W and height H of the target region. Then, the scale correction factor is calculated. Where D0 is the preset reference size. The major and minor axis lengths are multiplied by a correction factor s to obtain the final shape parameters. This target-size-based correction method ensures that the size of the search area can be dynamically adjusted as the target size changes, avoiding the problem of the search area being too large or too small. Simultaneously, the corrected major and minor axis lengths, along with the determined center position, are used as the shape parameters of the search area, providing a complete parameter basis for subsequent search area construction.
[0078] Step 405: Construct a search region based on the current position and shape parameters of the target object, which will serve as the search range for the next tracking cycle.
[0079] Specifically, an elliptical search region is constructed based on the calculated shape parameters. First, the direction of the ellipse is determined, using the average value θd of the motion directions from the most recent three frames as the tilt angle of the ellipse's major axis. Then, with the predicted center position (xc, yc) as the center, and the major axis a and minor axis b as parameters, an elliptical search region is constructed. For any point (x, y) within the search region, the ellipse equation must be satisfied: ((x-xc)·cosθd+(y-yc)·sinθd). 2 / a 2 +((y-yc)·cosθd-(x-xc)·sinθd) 2 / b 2 ≤1. This elliptical search region considers both the target's direction of motion and the turning characteristics reflected by the ratio of the major and minor axes, providing a reasonable search range for the next tracking cycle. Simultaneously, to handle prediction errors, an additional transition region is extended outside the elliptical boundary, with a width of 10% of the minor axis, thus improving the tracking's fault tolerance.
[0080] Step 50: Perform feature matching on the target object in the search area, and take the target area with the highest matching degree as the latest position of the target object in the next tracking cycle to achieve continuous tracking of the target object.
[0081] Specifically, to achieve accurate localization and continuous tracking of the target object, a correlation filtering method is used for feature matching within the search area. First, the target object is transformed to the frequency domain, and multi-channel HOG features are extracted as template features. Then, a cyclic matrix is constructed, and the coefficients of the correlation filter are calculated using a Fast Fourier Transform. The search area is then filtered to obtain a response map; the location with the largest response value corresponds to a candidate target region. The matching degree between the candidate region and the template features is calculated, and the region with the highest matching degree is selected as the latest position of the target object in the next tracking cycle. The filter coefficients are updated online to achieve continuous tracking of the target object.
[0082] Based on the above embodiments, as another optional embodiment, the step of performing feature matching on the target object in the search area and taking the target area with the highest matching degree as the latest position of the target object in the next tracking cycle to achieve continuous tracking of the target object may further include the following steps: Step 501: Extract the edge response spectrum of the target object as a reference feature, and calculate the cross-correlation coefficient between the edge response spectrum of each local region within the search area and the reference feature.
[0083] Specifically, to obtain the frequency domain feature representation of the target object, the edge response spectrum of the target object is first extracted. The Sobel operator is applied to the target region to calculate the gradients in the horizontal and vertical directions, obtaining a gradient magnitude map. A two-dimensional Fourier transform is performed on the gradient magnitude map to obtain the frequency domain representation, denoted as the reference feature F(u,v). Then, a sliding window approach is used within the search region to calculate the edge response spectrum G(u,v) for each local region. The cross-correlation coefficient ρ between the reference feature and the response spectra of each local region is calculated using the following formula: Where G* represents the conjugate complex number, F(u,v) is the edge response spectrum of the target object; G(u,v) is the edge response spectrum of the local region; G*(u,v) is the conjugate complex number of G(u,v), and (u,v) are the frequency domain coordinates. This frequency domain-based feature representation method has good robustness to target rotation and scale changes.
[0084] Step 502: Construct a frequency response diagram based on the cross-correlation coefficient, and extract multiple response points in the frequency response diagram whose response values exceed a preset threshold.
[0085] Specifically, to effectively locate candidate target regions, a frequency response map needs to be constructed and significant response points extracted. The cross-correlation coefficient ρ is mapped to the corresponding positions in the search region, forming a two-dimensional frequency response map R(x,y). To highlight significant responses, non-maximum suppression is applied to the response map to suppress local non-maximum response values. Then, a preset threshold is set: T = μ + 2σ, where μ and σ are the mean and standard deviation of the response map, respectively. Local maxima points with response values greater than the threshold T are extracted as the response point set {(xi,yi)}. This threshold-based response point extraction method can effectively filter weak response regions and retain the locations most likely to contain the target.
[0086] Step 503: Construct candidate matching regions centered on each response point.
[0087] Specifically, to determine candidate matching regions, a search sub-region is constructed centered on each response point. For each point (xi, yi) in the set of response points, a rectangular region slightly larger than the current size of the target is constructed centered on that point as a candidate matching region. For example, if the current size of the target is width W and height H, then the size of the candidate matching region is set to (1+α)W and (1+α)H, where α is the size expansion coefficient, typically between 0.2 and 0.3. This expansion size setting can accommodate possible scale changes and positional shifts of the target. If the response point is close to the boundary of the search region, the boundary of the candidate region needs to be adjusted to ensure that the candidate region is completely within the search range. This candidate region construction method based on response points ensures sufficient matching space while avoiding an excessively large search range, effectively improving the accuracy and efficiency of subsequent matching.
[0088] Step 504: Calculate the frequency feature similarity between each candidate matching region and the target object, and select the candidate matching region with the highest similarity as the latest position of the target object in the next tracking cycle.
[0089] Specifically, to select the best matching position from the candidate matching regions, the frequency feature similarity between each candidate region and the target object is calculated. For each candidate matching region, its frequency feature spectrum H(u,v) is extracted, and the similarity with the target reference feature F(u,v) is calculated using the phase correlation method. Specifically, the cross-spectrum is first calculated: C(u,v) is the cross spectrum; F(u,v) is the edge response spectrum of the target object; H(u,v) is the edge response spectrum of the candidate region; and H*(u,v) is the complex conjugate of H(u,v). Then, an inverse Fourier transform is performed on C(u,v) to obtain the correlation surface, and its peak value is taken as the similarity score. The candidate matching region with the highest similarity score is selected as the latest position of the target object in the next tracking cycle. This frequency-feature-based matching method has good robustness to illumination changes and partial occlusion, and can achieve stable target tracking.
[0090] Please see Figure 3 , Figure 3 This is a schematic diagram of target tracking under occlusion conditions provided in an embodiment of this application; as shown... Figure 3 As shown, Figure 3 The diagrams show the target object before occlusion, during occlusion, and when the target is re-locked. Before occlusion, the target object can be clearly identified. When the target object enters the occlusion area (such as a vehicle entering under a tree), the tracking of the target object disappears. According to the solution provided in this application, the latest position of the target object in the next tracking cycle can be predicted. That is, even when the target object is occluded, the approximate position of the target object can be predicted. This also facilitates the subsequent tracking of the target object after it leaves the occlusion area, thus achieving anti-occlusion tracking of the target object.
[0091] Please see Figure 4 This is a schematic diagram of a video image target recognition and tracking system provided in an embodiment of this application, wherein the system includes: The feature sequence determination module is used to acquire video image sequences, perform fast time domain and slow time domain feature extraction on the video image sequences respectively, and perform correlation analysis on the extracted features to obtain a feature sequence corresponding to at least one target object. An enhancement parameter determination module is used to determine feature enhancement parameters based on the distribution parameters of the target object in the feature sequence; The target sequence determination module is used to enhance the target object based on the feature enhancement parameters to obtain the enhanced target sequence; The search area determination module is used to obtain the motion parameters of the target object in multiple consecutive frames within the current tracking period of the target sequence, and determine the search area for the next tracking period based on the motion parameters. The target tracking and matching module is used to perform feature matching on the target object in the search area, and take the target area with the highest matching degree as the latest position of the target object in the next tracking cycle, so as to realize continuous tracking of the target object.
[0092] Optionally, the feature sequence determination module is further configured to perform foreground segmentation on the video image sequence and extract at least one candidate target region containing the target object; For each candidate target region, the difference image of adjacent frames in the video image sequence is obtained, and the instantaneous motion features of the target object are extracted as fast temporal features based on the gradient changes in the difference image. Multi-directional edge detection is performed on the target object within a preset time window to obtain the edge response spectrum of the target object, and the stable structural features of the target object are extracted as slow time domain features based on the temporal cumulative value of the edge response spectrum. The feature fusion weights are determined based on the instantaneous nature of the fast time-domain features and the stability of the slow time-domain features. The fast time-domain features and the slow time-domain features are then weighted and combined according to their respective feature fusion weights to obtain the feature sequences corresponding to each target object.
[0093] Optionally, the enhanced parameter determination module is also used to calculate the contrast distribution and texture complexity distribution of the target object in the feature sequence; A contrast adjustment coefficient is determined based on the contrast distribution, and a texture enhancement coefficient is determined based on the texture complexity distribution; the contrast adjustment coefficient and the texture enhancement coefficient are combined to obtain a feature enhancement parameter, wherein the feature enhancement parameter is used to enhance the contrast and texture features of the target object.
[0094] Optionally, the target sequence determination module is also used to decompose the target object into low-frequency components that characterize the overall features of the target and high-frequency components that characterize the detailed features of the target. The grayscale mapping range of the low-frequency component is determined based on the contrast adjustment coefficient, and nonlinear contrast enhancement processing is performed on the low-frequency component. The gain parameters of the high-frequency components are determined based on the texture enhancement coefficient, and the high-frequency components are then subjected to gain processing. The fusion weights of the processed low-frequency components and high-frequency components are calculated separately, and the processed low-frequency components and high-frequency components are combined according to the fusion weights to obtain the enhanced target sequence.
[0095] Optionally, the search region determination module is also used to extract the motion trajectory of the target object in multiple consecutive frames and calculate the displacement vector sequence between adjacent sampling points in the motion trajectory; Based on the displacement vector sequence, the tangent and normal directions of the motion trajectory at each sampling point are calculated to obtain the motion direction sequence of the target object; Based on the motion direction sequence, the curvature value and angle change of the motion trajectory are calculated to obtain the steering characteristic parameters of the target object; the shape parameters of the search area are calculated according to the steering characteristic parameters, and the shape parameters include the major axis, minor axis and center position of the search area; A search region is constructed based on the current position of the target object and the shape parameters, serving as the search range for the next tracking cycle.
[0096] Optionally, the search area determination module is also used to calculate the ratio of the arc length to the chord length of the motion trajectory; The center position of the search area is determined based on the ratio, and the center position is located at the center of the arc of the motion trajectory; The major axis length of the search region is calculated based on the arc length, and the minor axis length of the search region is calculated based on the chord length. The major axis length and minor axis length are scaled according to the current size of the target object, and the corrected major axis, minor axis and center position are used as shape parameters.
[0097] Optionally, the target tracking and matching module is also used to extract the edge response spectrum of the target object as a reference feature, and calculate the cross-correlation coefficient between the edge response spectrum of each local region within the search area and the reference feature; A frequency response diagram is constructed based on the cross-correlation coefficient, and multiple response points in the frequency response diagram whose response values exceed a preset threshold are extracted. Construct a candidate matching region centered on each of the aforementioned response points; Calculate the frequency feature similarity between each candidate matching region and the target object, and select the candidate matching region with the highest similarity as the latest position of the target object in the next tracking cycle.
[0098] It should be noted that the system provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0099] This application also provides a computer storage medium that can store multiple instructions. The instructions are adapted to be loaded by a processor and executed by a video image target recognition and tracking method according to the above embodiments. For the specific execution process, please refer to the detailed description of the above embodiments, which will not be repeated here.
[0100] Please refer to Figure 5 This application also discloses an electronic device. Figure 5 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.
[0101] The communication bus 302 is used to enable communication between these components.
[0102] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0103] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0104] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling data stored in the memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.
[0105] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include a non-transitory computer-readable storage medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. (Refer to...) Figure 5The memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application for identifying and tracking video image targets.
[0106] exist Figure 5 In the illustrated electronic device 300, the user interface 303 is mainly used to provide an input interface for the user and to acquire user input data; while the processor 301 can be used to call an application program stored in the memory 305 for a video image target recognition and tracking method. When executed by one or more processors 301, the electronic device 300 performs one or more methods as described in the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0107] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0108] In the various embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.
[0109] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0110] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0111] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0112] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will readily conceive of those skilled in the art upon consideration of the specification and the disclosure of practical truths.
[0113] This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A method for recognizing and tracking a target in a video image, characterized by, The method comprises: obtaining a video image sequence, respectively extracting features in fast time domain and slow time domain of the video image sequence, and performing correlation analysis on the extracted features to obtain a feature sequence corresponding to at least one target object; determining a feature enhancement parameter according to a distribution parameter of the target object in the feature sequence; performing enhancement processing on the target object based on the feature enhancement parameter to obtain an enhanced target sequence; obtaining motion parameters of the target object in consecutive multiple frames in a current tracking period of the target sequence, and determining a search area of a next tracking period according to the motion parameters; performing feature matching on the target object in the search area, and taking a target area with the highest matching degree as the latest position of the target object in the next tracking period to realize continuous tracking of the target object.
2. The video image target recognition tracking method according to claim 1, characterized in that, The method comprises: performing foreground segmentation on the video image sequence to extract at least one candidate target area containing a target object; for each candidate target area, obtaining a difference image of adjacent frames in the video image sequence, and extracting instantaneous motion features of the target object as fast time domain features according to gradient changes in the difference image; performing multi-directional edge detection on the target object within a preset time window to obtain an edge response spectrum of the target object, and extracting stable structure features of the target object as slow time domain features according to time domain cumulative values of the edge response spectrum; determining a feature fusion weight according to the instantaneous nature of the fast time domain features and the stability of the slow time domain features, and combining the fast time domain features and the slow time domain features according to the corresponding feature fusion weight to obtain a feature sequence corresponding to each target object.
3. The video image object recognition tracking method according to claim 1, wherein, The method comprises: calculating a contrast distribution and a texture complexity distribution of the target object in the feature sequence; determining a contrast adjustment coefficient according to the contrast distribution, and determining a texture enhancement coefficient according to the texture complexity distribution; combining the contrast adjustment coefficient and the texture enhancement coefficient to obtain a feature enhancement parameter, wherein the feature enhancement parameter is used to enhance the contrast and texture features of the target object.
4. The video image target recognition tracking method according to claim 3, wherein, The method comprises: decomposing the target object into a low-frequency component representing overall features of the target object and a high-frequency component representing detailed features of the target object; determining a gray mapping interval of the low-frequency component according to the contrast adjustment coefficient, and performing nonlinear contrast enhancement processing on the low-frequency component; determining a gain parameter of the high-frequency component according to the texture enhancement coefficient, and performing gain processing on the high-frequency component; calculating a fusion weight of the processed low-frequency component and the high-frequency component respectively, and combining the processed low-frequency component and the high-frequency component according to the fusion weight to obtain an enhanced target sequence.
5. The method of claim 1, wherein, The acquiring the motion parameters of the target object in continuous multiple frames in a current tracking period of the target sequence, and determining a search region of a next tracking period according to the motion parameters, comprises: extracting a motion trajectory of the target object in the continuous multiple frames, and calculating a displacement vector sequence between adjacent sampling points in the motion trajectory; calculating a tangent direction and a normal direction of the motion trajectory at each sampling point according to the displacement vector sequence, to obtain a motion direction sequence of the target object; calculating a curvature value and an angle change amount of the motion trajectory based on the motion direction sequence, to obtain a turning feature parameter of the target object; calculating a shape parameter of the search region according to the turning feature parameter, the shape parameter comprising a long axis, a short axis and a center position of the search region; constructing the search region based on the current position of the target object and the shape parameter, as a search range of the next tracking period.
6. The video image target recognition tracking method according to claim 5, wherein, The calculating the shape parameter of the search region according to the turning feature parameter comprises: calculating a ratio of an arc length to a chord length of the motion trajectory; determining the center position of the search region based on the ratio, the center position being located at an arc center position of the motion trajectory; calculating a long axis length of the search region according to the arc length, and calculating a short axis length of the search region according to the chord length; performing scale correction on the long axis length and the short axis length according to a current size of the target object, and taking the corrected long axis, the short axis and the center position as the shape parameter.
7. The method of claim 1, wherein, The performing feature matching on the target object in the search region, and taking a target region with the highest matching degree as a latest position of the target object in the next tracking period, comprises: extracting an edge response spectrum of the target object as a reference feature, and calculating a cross-correlation coefficient between an edge response spectrum of each local region in the search region and the reference feature; constructing a frequency response map according to the cross-correlation coefficient, and extracting a plurality of response points with a response value exceeding a preset threshold in the frequency response map; constructing a candidate matching region corresponding to each response point as a center; calculating a frequency feature similarity between each candidate matching region and the target object, and selecting a candidate matching region with the highest similarity as the latest position of the target object in the next tracking period.
8. A video image object recognition and tracking system, characterized by The system comprises: a feature sequence determination module, configured to acquire a video image sequence, perform feature extraction on the video image sequence in a fast time domain and a slow time domain respectively, and perform correlation analysis on the extracted features to obtain a feature sequence corresponding to at least one target object; an enhancement parameter determination module, configured to determine a feature enhancement parameter according to a distribution parameter of the target object in the feature sequence; a target sequence determination module, configured to perform enhancement processing on the target object based on the feature enhancement parameter, to obtain an enhanced target sequence; a search region determination module, configured to acquire motion parameters of the target object in continuous multiple frames in a current tracking period of the target sequence, and determine a search region of a next tracking period according to the motion parameters. A target tracking matching module is configured to perform feature matching on the target object in the search area, and take the target area with the highest matching degree as the latest position of the target object in the next tracking period, so as to realize continuous tracking of the target object.
9. A computer-readable storage medium, characterized in that, A computer readable storage medium stores a plurality of instructions, which are adapted to be loaded and executed by a processor to perform the method of any one of claims 1-7.
10. An electronic device, comprising: An electronic device includes a processor, a memory, a user interface, and a network interface, the memory is configured to store instructions, the user interface and the network interface are configured to communicate with other devices, and the processor is configured to execute the instructions stored in the memory to cause the electronic device to perform the method of any one of claims 1-7.