Image target tracking method and device based on deep learning

Through the deep learning-based image target tracking method, adaptive frame filtering and feature mapping combined with dynamic state estimation, the problems of tracking accuracy reduction and response delay in complex scenarios are solved, and high-precision and fast target tracking are achieved.

CN120339335AInactive Publication Date: 2025-07-18HOHEM TECHNOLOGY CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510557558.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In complex scenarios, traditional algorithms combined with gimbal control image target tracking technology have problems such as reduced tracking accuracy and system response delay.

Method used

The image target tracking method based on deep learning is adopted to generate feature frame sequences and keyframes through adaptive frame filtering, feature extraction and mapping are performed, global correlation features are calculated, dynamic target state estimation and cluster correlation analysis are performed, and trajectory optimization is performed in combination with deep learning models to generate tracked target trajectories.

Benefits of technology

Improves tracking accuracy and response speed, maintains a stable target tracking effect in complex scenarios, adjusts the camera perspective in real time, and ensures that the target is always within the monitoring range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339335A_ABST
    Figure CN120339335A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and provides an image target tracking method and device based on deep learning, and the method comprises the steps: obtaining an original frame sequence of a target video, carrying out the adaptive frame screening of the original frame sequence, and obtaining a plurality of feature frame sequences, key frames and common frames, performing feature extraction and feature mapping on all the feature frame sequences to obtain target trajectory nodes, performing trajectory optimization calculation on the key frames and the common frames by using all the target trajectory nodes to obtain key frame trajectories and common frame trajectories, and inputting the key frame trajectories and the common frame trajectories into a preset deep learning model to perform trajectory reconstruction optimization so as to obtain a target trajectory; and obtaining a tracking target trajectory. By performing adaptive frame screening on the original frame sequence and generating the feature frame sequence, the key frame and the common frame, the data processing efficiency and the information extraction accuracy are improved, the accurate generation of the target motion track is realized, and the problems of reduced tracking accuracy and delayed system response in a complex scene are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image processing, and particularly to an image target tracking method and device based on deep learning. Background Art

[0002] With the rapid development of artificial intelligence technology, computer vision technology based on deep learning has been widely applied in multiple fields, including autonomous driving, security monitoring, intelligent manufacturing, etc. In these application scenarios, image target tracking technology plays a key role. By acquiring and analyzing video data in real time, this technology can accurately identify and track dynamic targets, thereby improving the automation and intelligence level of the system.

[0003] In related technical means, traditional algorithms are often combined with pan-tilt control to complete the target tracking task. For example, target features in video frames are extracted through feature extraction methods (such as SIFT or HOG), and the target position is predicted using Kalman filtering or particle filtering. Then, the prediction results are converted into control commands for the pan-tilt, and the direction and angle of the camera are adjusted in real time. These methods can achieve relatively ideal tracking effects and realize continuous monitoring of the target area in scenarios with stable environments and relatively simple target motion laws.

[0004] For the above technical solutions, although target tracking can be achieved by combining traditional algorithms with pan-tilt control, in complex scenarios, such as when the target moves quickly, is partially occluded, or the environmental light changes, there are still problems of reduced tracking accuracy and system response delay. Summary of the Invention

[0005] In order to improve the problems of reduced tracking accuracy and system response delay in complex scenarios, this application provides an image target tracking method and device based on deep learning.

[0006] The present invention provides an image target tracking method based on deep learning, including: obtaining the original frame sequence of the target video, performing adaptive frame screening on the original frame sequence to obtain several feature frame sequences, key frames, and ordinary frames; performing feature extraction and feature mapping on all the feature frame sequences to obtain several enhanced feature mappings, calculating the cross-frame correlation degree of all the enhanced feature mappings to obtain several global correlation features; performing dynamic target state estimation on all the global correlation features to obtain several target state vectors, performing clustering correlation analysis on all the target state vectors to obtain the target trajectory nodes corresponding to each global correlation feature; using all the target trajectory nodes to perform trajectory optimization calculation on the key frames and the ordinary frames to obtain the key frame trajectory and the ordinary frame trajectory, and inputting the key frame trajectory and the ordinary frame trajectory into a preset deep learning model for trajectory reconstruction and optimization to obtain the tracking target trajectory.

[0007] As a preferred solution, the steps of obtaining the original frame sequence of the target video, adaptively filtering the original frame sequence to obtain a plurality of feature frame sequences, key frames, and ordinary frames include: collecting the original frame sequence of the target video, using a motion estimation algorithm to calculate the inter-frame optical flow of the original frame sequence of the target video to obtain an inter-frame motion vector and inter-frame change information, using the inter-frame motion vector to extract the change mode of the inter-frame change information to obtain a global change mode and a local change mode; constructing a multi-scale change feature map based on the global change mode and the local change mode, and performing feature level division on the multi-scale change feature map based on a preset multi-scale hierarchical sampling strategy to obtain a low-level feature region and a high-level feature region; performing information entropy evaluation on the low-level feature region and the high-level feature region to obtain a low-level information entropy distribution and a high-level information entropy distribution, and selecting a plurality of feature sampling points based on the low-level information entropy distribution and the high-level information entropy distribution; performing adaptive frame filtering on the original frame sequence based on the feature sampling points to obtain a feature frame sequence and candidate key frames, performing temporal consistency analysis on the feature frame sequence to obtain temporal consistency features, and performing motion stability analysis on the candidate key frames to obtain key frame stability vectors; using the temporal consistency features and the key frame stability vectors to filter the candidate key frames to obtain key frames, and removing the key frames from the feature frame sequence to obtain ordinary frames.

[0008] As a preferred solution, the steps of performing feature extraction and feature mapping on all the feature frame sequences to obtain a plurality of enhanced feature maps, and calculating the cross-frame correlation degree of all the enhanced feature maps to obtain a plurality of global correlation features include: performing multi-scale feature extraction on all the feature frame sequences to obtain a spatial feature set and a temporal feature set, using the spatial feature set to perform feature matching calculation on the temporal feature set to obtain matching feature pairs; performing local matching degree evaluation on all the matching feature pairs to obtain a local matching weight matrix, and performing cross-frame feature fusion on the temporal feature set based on the local matching weight matrix to obtain enhanced temporal features; using the spatial feature set to perform multi-layer feature mapping on the enhanced temporal features to obtain enhanced feature maps, performing inter-frame correlation calculation on all the enhanced feature maps to obtain an inter-frame correlation matrix and an inter-frame similarity matrix; calculating a global matching weight based on the inter-frame correlation matrix and the inter-frame similarity matrix, and performing associated feature aggregation on the enhanced feature maps using the global matching weight to obtain a plurality of global correlation features.

[0009] As a preferred solution, the step of evaluating the local matching degree for all the matching feature pairs to obtain a local matching weight matrix, and performing cross-frame feature fusion on the temporal feature set based on the local matching weight matrix to obtain enhanced temporal features includes: calculating the local gradient distribution of feature points based on all the matching feature pairs to obtain a feature point gradient matrix, calculating the local feature contrast of the feature point gradient matrix to obtain a contrast distribution matrix, and calculating the local matching degree based on the contrast distribution matrix to obtain a local matching weight matrix; using the local matching weight matrix to perform local feature aggregation on the temporal feature set to obtain a preliminary fusion feature, calculating the inter-frame feature constraint of the preliminary fusion feature to obtain an inter-frame feature constraint matrix, and using the inter-frame feature constraint matrix to perform cross-frame feature enhancement on the preliminary fusion feature to obtain enhanced temporal features.

[0010] As a preferred solution, the step of performing dynamic target state estimation on all the global association features to obtain a plurality of target state vectors, and performing clustering association analysis on all the target state vectors to obtain the target trajectory nodes corresponding to each global association feature includes: extracting target regions for all the global association features to obtain candidate target regions, analyzing the motion trend of the candidate target regions to obtain the target motion trend, and using the target motion trend to perform time series modeling on the candidate target regions to obtain a target time series; calculating the dynamic states of the target time series and the global association features to obtain a plurality of target state vectors, and performing feature clustering on all the target state vectors to obtain initial target clustering clusters; calculating the inter-cluster relationships of the initial target clustering clusters to obtain an inter-cluster association matrix and an inter-cluster separation matrix, and performing association adjustment on the initial target clustering clusters based on the inter-cluster association matrix and the inter-cluster separation matrix to obtain the target trajectory nodes.

[0011] As a preferred solution, the step of using all the target trajectory nodes to perform trajectory optimization calculation on the key frames and the ordinary frames to obtain a key frame trajectory and an ordinary frame trajectory, and inputting the key frame trajectory and the ordinary frame trajectory into a preset deep learning model for trajectory reconstruction optimization to obtain a tracking target trajectory includes: using all the target trajectory nodes to perform trajectory node screening on the key frames and the ordinary frames to obtain key frame trajectory points and ordinary frame trajectory points, and constructing trajectory segments for the key frame trajectory points and the ordinary frame trajectory points to obtain an initial key frame trajectory segment and an initial ordinary frame trajectory segment; performing trajectory smoothing calculation on the initial key frame trajectory segment and the initial ordinary frame trajectory segment to obtain a smoothed key frame trajectory segment and a smoothed ordinary frame trajectory segment, and performing trajectory consistency calculation on the smoothed key frame trajectory segment and the smoothed ordinary frame trajectory segment to obtain a trajectory consistency parameter; performing trajectory segment optimization on the trajectory consistency parameter to obtain a key frame trajectory and an ordinary frame trajectory, and inputting the key frame trajectory and the ordinary frame trajectory into the deep learning model for trajectory reconstruction calculation to obtain a tracking target trajectory.

[0012] As a preferred solution, the step of using all the target trajectory nodes to perform trajectory node screening on the key frames and the ordinary frames to obtain key frame trajectory points and ordinary frame trajectory points, and constructing trajectory segments for the key frame trajectory points and the ordinary frame trajectory points to obtain an initial key frame trajectory segment and an initial ordinary frame trajectory segment includes: performing spatial neighborhood analysis on the key frames and the ordinary frames based on all the target trajectory nodes to obtain spatial neighborhood features; performing temporal consistency evaluation on the spatial neighborhood features to obtain a temporal consistency matrix, and using the temporal consistency matrix to screen all the target trajectory nodes to obtain a preliminary trajectory node set; performing trajectory node allocation on the key frames and the ordinary frames based on the preliminary trajectory node set to obtain key frame trajectory points and ordinary frame trajectory points; performing motion direction constraint calculation on the key frame trajectory points and the ordinary frame trajectory points to obtain a trajectory direction matrix, and grouping the key frame trajectory points and the ordinary frame trajectory points based on the trajectory direction matrix to obtain a set of trajectory point groups; constructing trajectory segments based on the set of trajectory point groups to obtain an initial key frame trajectory segment and an initial ordinary frame trajectory segment.

[0013] The present application also provides an image target tracking device based on deep learning, comprising: an acquisition module, configured to acquire the original frame sequence of a target video, perform adaptive frame screening on the original frame sequence to obtain a plurality of feature frame sequences, key frames, and ordinary frames; a mapping module, configured to perform feature extraction and feature mapping on all the feature frame sequences to obtain a plurality of enhanced feature mappings, calculate the cross-frame correlation degree for all the enhanced feature mappings to obtain a plurality of global correlation features; an analysis module, configured to perform dynamic target state estimation on all the global correlation features to obtain a plurality of target state vectors, perform clustering correlation analysis on all the target state vectors to obtain the target trajectory nodes corresponding to each global correlation feature; a calculation module, configured to use all the target trajectory nodes to perform trajectory optimization calculation on the key frames and the ordinary frames to obtain key frame trajectories and ordinary frame trajectories, input the key frame trajectories and the ordinary frame trajectories into a preset deep learning model for trajectory reconstruction and optimization to obtain the tracking target trajectory.

[0014] Compared with the prior art, the present application has the following beneficial effects: high tracking accuracy and fast response speed. By performing adaptive frame screening on the original frame sequence to generate feature frame sequences, key frames, and ordinary frames, the efficiency of data processing and the accuracy of information extraction are greatly improved. Through feature extraction, enhanced feature mapping, and cross-frame correlation degree calculation, high-quality global correlation features are obtained, ensuring the coherence and consistency between features. The combination of the dynamic target state estimation algorithm and clustering correlation analysis realizes the accurate modeling of the target motion state and the precise generation of trajectory nodes. In the final trajectory optimization and reconstruction process, the combination of the deep learning model further improves the accuracy and integrity of the target trajectory, and generates pan-tilt control commands in real time, enabling the camera to quickly adjust the viewing angle according to the target motion trajectory to ensure that the target is always within the monitoring range, effectively solving the problems of target occlusion and light change in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0016] The structures, proportions, sizes, etc. depicted in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the conditions for the implementation of the present invention. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the objectives that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.

[0017] Figure 1 is a schematic flowchart of an image target tracking method based on deep learning provided by an embodiment of the present invention; Figure 2 is a schematic block diagram of the structure of an image target tracking device based on deep learning provided by an embodiment of the present invention.

[0018] Explanation of reference numerals: 10. Image target tracking device based on deep learning; 11. Acquisition module; 12. Mapping module; 13. Analysis module; 14. Calculation module. Detailed implementation manners

[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] The flowchart shown in the drawings is only an example illustration, and does not necessarily include all the content and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may be changed according to the actual situation.

[0021] It should also be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0022] It should be further understood that the term "and / or" used in this specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0023] Next, the technical solutions of the present invention will be further described in conjunction with the drawings and through specific implementation manners.

[0024] Example 1: As Figure 1 shown, this application provides an image target tracking method based on deep learning, including steps S100 to S400.

[0025] Step S100, obtain the original frame sequence of the target video, perform adaptive frame screening on the original frame sequence to obtain several feature frame sequences, key frames, and ordinary frames.

[0026] In this step, by parsing the original frame sequence of the target video and using the adaptive frame screening algorithm, feature frame sequences with higher information content are screened according to the change characteristics between frames and the target movement trend, while marking key frames with prominent motion features and ordinary frames with less background change. Specifically, the adaptive frame screening algorithm uses a deep learning model to evaluate the motion vectors and feature points of each frame in the original frame sequence, selects frames with significant features as feature frames, marks frames when the target moves rapidly or occlusion occurs as key frames, and marks relatively stable frames as ordinary frames.

[0027] For example, in a surveillance scenario, when the target suddenly starts to move rapidly from a stationary state, this frame will be screened as a key frame, while frames captured during the normal movement of the target are screened as ordinary frames.

[0028] Step S200, perform feature extraction and feature mapping on all feature frame sequences to obtain several enhanced feature mappings, and calculate the cross-frame correlation degree for all enhanced feature mappings to obtain several global correlation features.

[0029] In this step, the feature extraction module accurately locates the feature points of the feature frame sequence and generates enhanced feature mappings through the feature mapping layer of the deep learning model. Specifically, the feature extraction algorithm uses a convolutional neural network to perform high-precision feature extraction on the shape, texture, edges, etc. of the target, and amplifies the correlation of features between frames through the feature enhancement algorithm. Subsequently, the cross-frame correlation degree calculation module analyzes the similarity and correlation between feature frames to generate global correlation features for subsequent trajectory derivation.

[0030] For example, when detecting the movement of a pedestrian in different video frames, the feature extraction module can extract the texture features and contour features of the pedestrian's clothing, and the correlation degree calculation module connects similar features in the front and back frames into global correlation features.

[0031] Step S300, perform dynamic target state estimation on all global correlation features to obtain several target state vectors, and perform clustering correlation analysis on all target state vectors to obtain the target trajectory nodes corresponding to each global correlation feature.

[0032] In this step, based on the global correlation features, the dynamic target state estimation algorithm models and estimates the dynamic parameters of the target, such as position, velocity, and acceleration, to generate the target state vector. Specifically, the system models the motion state of the dynamic target through a recurrent neural network and combines the information of the previous and subsequent frames to predict the specific state of the target at each time point. Subsequently, all target state vectors are grouped through the clustering correlation analysis algorithm, and each group of vectors is associated with the global correlation features to generate a target trajectory node, providing accurate trajectory data for subsequent pan-tilt control.

[0033] For example, when the tracking module processes the video of a moving vehicle, the system can generate multiple trajectory nodes based on the motion characteristics of the vehicle and cluster the nodes to accurately segment and label the motion trajectories of different vehicles.

[0034] Step S400: Use all target trajectory nodes to perform trajectory optimization calculations on key frames and ordinary frames to obtain key frame trajectories and ordinary frame trajectories, and input the key frame trajectories and ordinary frame trajectories into a preset deep learning model for trajectory reconstruction and optimization to obtain the tracking target trajectory.

[0035] In this step, the trajectory optimization module performs interpolation and error correction on all target trajectory nodes and generates key frame trajectories and ordinary frame trajectories. Specifically, the system first performs polynomial fitting on the data between trajectory nodes to optimize the smoothness of the trajectory, and then inputs the optimized trajectory into the deep learning model for trajectory reconstruction and optimization to further improve the trajectory accuracy by learning historical trajectory patterns and environmental characteristics. Finally, the generated tracking target trajectory can not only accurately describe the motion path of the target but also generate real-time control commands for the pan-tilt to achieve dynamic tracking of the target.

[0036] For example, when tracking a fast-moving drone, through the trajectory optimization of the deep learning model, the system can generate a smooth and accurate tracking trajectory and adjust the direction of the pan-tilt to ensure that the camera always points at the drone.

[0037] In this embodiment, through adaptive frame screening of the original frame sequence of the target video, several feature frame sequences, key frames, and ordinary frames are generated, and feature extraction and feature mapping are performed on all feature frame sequences to obtain enhanced feature maps. Then, the cross-frame correlation degree of all enhanced feature maps is calculated to obtain global correlation features, and dynamic target state estimation is performed on them to generate target state vectors. Next, through clustering correlation analysis, all target state vectors are associated with the global correlation features to generate target trajectory nodes. Finally, by using the target trajectory nodes, the control commands of the pan-tilt are calculated, and at the same time, trajectory optimization and deep learning model trajectory reconstruction and optimization are performed on key frames and ordinary frames to generate an accurate tracking target trajectory and adjust the pan-tilt in real time to achieve efficient tracking of the target.

[0038] The effectiveness of video frame data processing is improved through adaptive frame screening. Meanwhile, by combining the feature extraction and trajectory reconstruction optimization of deep learning models, the accuracy of object tracking is significantly enhanced. In addition, this solution significantly enhances the robustness in complex scenarios. Even in situations such as rapid movement of the object, partial occlusion, and changes in lighting conditions, it can maintain a stable tracking effect, generate control commands for the pan-tilt head in real time, and improve the system's continuous tracking ability for the object by precisely adjusting the direction and angle of the camera, meeting the application requirements in various complex environments.

[0039] Embodiment 2: In step S100, the original frame sequence of the target video is collected. The inter-frame optical flow of the original frame sequence of the target video is calculated using the motion estimation algorithm to obtain the inter-frame motion vector and inter-frame change information. The change mode extraction is performed on the inter-frame change information using the inter-frame motion vector to obtain the global change mode and the local change mode.

[0040] By parsing the original frame sequence of the target video and combining the optical flow method to extract the motion characteristics between each pair of adjacent frames, the optical flow vector field is used to calculate the inter-frame motion vector and inter-frame change information. Specifically, the original frame sequence is analyzed frame by frame using the motion estimation algorithm to extract the direction change and speed characteristics of the object, while separating the background motion of the environment from the actual motion of the object. Based on the inter-frame motion vector, the system extracts the global change mode (reflecting the overall movement trend of the scene) and the local change mode (reflecting the fine motion characteristics of the target area) in the frame sequence through the change mode extraction module, providing important basic data for subsequent pan-tilt head control and object tracking.

[0041] For example, when shooting a pedestrian, when the tracking camera captures a pedestrian entering the monitoring area and starting to move, the motion estimation algorithm will generate an inter-frame motion vector according to the pedestrian's motion path, and simultaneously extract the local change mode including the pedestrian area and the global change mode caused by the camera in the entire scene.

[0042] A multi-scale change feature map is constructed based on the global change mode and the local change mode. The feature hierarchical division of the multi-scale change feature map is performed based on the preset multi-scale hierarchical sampling strategy to obtain the low-level feature area and the high-level feature area.

[0043] By hierarchically expressing the global change pattern and the local change pattern, a multi-scale change feature map is constructed to reflect the motion characteristics of the target at different scales. Specifically, through a preset multi-scale hierarchical sampling strategy, the system divides the feature map layer by layer, defines the low-resolution area as the low-level feature area (reflecting the overall dynamic change characteristics of the environment), and defines the high-resolution area as the high-level feature area (reflecting the specific characteristics of the target, such as boundaries and textures). These feature areas can provide accurate feature-level information support for pan-tilt control.

[0044] For example, in a video surveillance scenario, when the camera captures a vehicle driving on a crowded street, the multi-scale feature map can distinguish the overall traffic flow trend of the street (low-level feature area) from the detailed changes of the vehicle, such as the contour or license plate information (high-level feature area).

[0045] Evaluate the information entropy of the low-level feature area and the high-level feature area to obtain the low-level information entropy distribution and the high-level information entropy distribution, and select several feature sampling points based on the low-level information entropy distribution and the high-level information entropy distribution.

[0046] By calculating the information entropy of the information in the multi-scale feature area, the system evaluates the distribution of features and the richness of information in the area. Specifically, the system calculates the information entropy distribution of the low-level feature area and the high-level feature area respectively. The low-level information entropy reflects the motion characteristics in the global change pattern and is suitable for overall target positioning, while the high-level information entropy can describe the target detail characteristics in the local change pattern, laying a foundation for accurate trajectory tracking. Based on the distribution of the information entropy, the system selects the area with a higher feature density as the feature sampling point, providing an accurate feature point reference for subsequent frame screening.

[0047] For example, in a drone camera system, the low-level information entropy distribution can assist in identifying the overall dynamic change area on the ground (such as a crowd or traffic flow), while the high-level information entropy distribution can locate the specific motion details of the drone target, thereby improving the accuracy of target tracking in a complex scenario.

[0048] Based on the feature sampling points, perform adaptive frame screening on the original frame sequence to obtain a feature frame sequence and candidate key frames, perform temporal consistency analysis on the feature frame sequence to obtain temporal consistency features, and perform motion stability analysis on the candidate key frames to obtain key frame stability vectors.

[0049] By using the feature sampling points selected in the previous step, the dynamic features of each frame in the original frame sequence are deeply screened to generate a feature frame sequence and candidate key frames. Specifically, the adaptive frame screening combines the temporal continuity and spatial stability of the target region, dynamically adjusts the screening strategy, and preferentially retains the frames that play a key role in subsequent target tracking. Next, a temporal consistency analysis is performed on the feature frame sequence to extract the motion trajectory and continuity features of the target on the time axis in the video frame sequence, forming temporal consistency features. On this basis, for the candidate key frames, by calculating the change amplitude of their motion states within a specific time period, their stability degrees in space and time are evaluated, and a key frame stability vector is generated. The key frame stability vector is used to represent the motion reliability of the candidate key frames and provides support for the optimization analysis of subsequent frames.

[0050] For example, in the scenario of unmanned aerial vehicle (UAV) target tracking, through the adaptive frame screening algorithm, the frames with the most significant changes of the UAV can be screened out from the full-scene video as candidate key frames. At the same time, the continuous features of the UAV during flight, such as the stability of the flight path and the consistency of the direction, are extracted through temporal consistency analysis, and the motion stability of the UAV under complex airflows is further analyzed, thereby generating a key frame stability vector.

[0051] The candidate key frames are screened using the temporal consistency features and the key frame stability vector to obtain the key frames, and the key frames are removed from the feature frame sequence to obtain ordinary frames.

[0052] By combining the temporal consistency features and the key frame stability vector, the candidate key frames are deeply screened to preferably select the key frames with the highest temporal continuity and spatial stability, and the remaining frames are removed. Specifically, the system uses a multi-objective characteristic evaluation algorithm to perform a weight analysis on the temporal consistency features and stability vectors of the candidate key frames, calculates the priority of the key frames, and screens out the key frames according to the priority. At the same time, during the process of removing the key frames from the feature frame sequence, the system classifies the remaining frames and defines them as ordinary frames, thereby distinguishing the priority levels of the frames and providing a basis for subsequent trajectory analysis and pan-tilt control.

[0053] For example, in security monitoring, when the camera tracks a fast-moving suspect, the screening algorithm will automatically select the frames when the suspect enters the monitoring area and suddenly changes direction as key frames, while the remaining frames of the travel trajectory are marked as ordinary frames, ensuring that the monitoring system focuses resources on capturing key motion details and at the same time providing a clear trajectory basis for the pan-tilt adjustment instructions of the camera.

[0054] In step S200, multi-scale feature extraction is performed on all feature frame sequences to obtain a spatial feature set and a temporal feature set, and the spatial feature set is used to perform feature matching calculations on the temporal feature set to obtain matching feature pairs.

[0055] By performing multi-scale analysis on all feature frame sequences using a deep learning model, the system extracts the spatial feature set and the temporal feature set of the feature frame sequences. Specifically, the spatial feature set conducts multi-level analysis on the geometric structure, boundary texture, color distribution, etc. of the target region in the video frame through a multi-scale feature extraction algorithm; the temporal feature set extracts dynamic features of the motion trajectory and time series changes between target frames through a temporal modeling algorithm. Subsequently, the system performs feature matching calculations on the spatial feature set and the temporal feature set to pair the feature points of similar targets between frames and generate matching feature pairs, which provide a basis for further feature fusion and enhancement.

[0056] For example, in the pan-tilt monitoring scenario, through multi-scale feature extraction, the spatial features of the target can be extracted, such as the shape of the vehicle and the boundary of the license plate. At the same time, based on the temporal feature set, the moving path and speed change of the vehicle are captured to generate matching feature pairs for analysis.

[0057] The local matching degree of all matching feature pairs is evaluated to obtain a local matching weight matrix, and cross-frame feature fusion is performed on the temporal feature set based on the local matching weight matrix to obtain enhanced temporal features.

[0058] By evaluating the local matching degree of all matching feature pairs, the system quantifies the matching correlation degree between target feature points and generates a local matching weight matrix to represent the priority and reliability of feature point matching. Specifically, the local matching degree evaluation combines the spatial relative position, gradient change, and contrast difference between feature points to calculate the local weight of each pair of matching features. Subsequently, cross-frame feature fusion is performed on the temporal feature set based on the weight matrix to integrate the matching feature points in consecutive frames and generate enhanced temporal features, thereby enhancing the feature consistency and stability of the target in the time dimension.

[0059] For example, in real-time target tracking, the local matching weight matrix can preferentially fuse the feature points in frames with better lighting conditions and at the same time eliminate invalid matching points caused by occlusion or background complexity, thereby ensuring the accuracy of temporal features.

[0060] The enhanced temporal features are subjected to multi-layer feature mapping using the spatial feature set to obtain enhanced feature maps, and inter-frame correlation calculations are performed on all enhanced feature maps to obtain an inter-frame correlation matrix and an inter-frame similarity matrix.

[0061] By combining the spatial feature set and enhanced temporal features, the system generates a high-dimensional enhanced feature map using a multi-layer feature mapping algorithm to represent the comprehensive characteristics of the target region. Specifically, the enhanced feature map performs multi-layer mapping on the geometry, boundaries, and motion states in the spatial feature set and the trajectories, speeds, etc. in the temporal feature set through a feature fusion module. Then, through an inter-frame correlation calculation module, the system generates an inter-frame correlation matrix and an inter-frame similarity matrix, which are used to quantify the global correlation degree and local similarity between feature maps respectively. These matrices provide important data support for subsequent global feature aggregation.

[0062] For example, in a complex surveillance scenario, the inter-frame similarity matrix can effectively analyze the changes in the characteristics of dynamic targets that appear in consecutive frames, such as the shape and motion trend of a vehicle when it moves from one camera to another.

[0063] Calculate the global matching weight based on the inter-frame correlation matrix and the inter-frame similarity matrix, and use the global matching weight to perform associated feature aggregation on the enhanced feature map to obtain several globally associated features.

[0064] By performing weighted analysis on the inter-frame correlation matrix and the inter-frame similarity matrix, the system calculates the global matching weight of the target region to represent the overall matching relationship of the target features between different frames. Specifically, the system uses a global feature aggregation algorithm to perform associated feature aggregation on the enhanced feature map, integrating cross-frame target characteristics to generate several globally associated features. These associated features can reflect the motion state and visual characteristics of the target in the video sequence, providing a solid feature basis for dynamic state calculation and trajectory derivation.

[0065] For example, in a UAV tracking scenario, the globally associated features can integrate the flight data of the UAV captured by multiple cameras, including flight path, attitude, and environmental interference, etc., to provide key support for the precise tracking and pan-tilt control of the UAV.

[0066] Among them, the steps of evaluating the local matching degree for all matching feature pairs to obtain a local matching weight matrix and performing cross-frame feature fusion on the temporal feature set based on the local matching weight matrix to obtain enhanced temporal features include: calculating the local gradient distribution of feature points based on all matching feature pairs to obtain a feature point gradient matrix, performing local feature contrast calculation on the feature point gradient matrix to obtain a contrast distribution matrix, and calculating the local matching degree based on the contrast distribution matrix to obtain a local matching weight matrix.

[0067] Through the analysis of all matching feature pairs, the system first generates a feature point gradient matrix using the local gradient changes of feature points, thereby capturing the motion direction and detailed change characteristics of the target area in the video frame. Specifically, the system uses a gradient calculation model to evaluate the gradient distribution pattern between feature points frame by frame to extract the gradient matrix reflecting the target motion trajectory and edge features. Then, the local feature contrast of the feature point gradient matrix is calculated, and a contrast distribution matrix is generated by quantifying the texture and illumination differences between feature points to express the relative stability and local change intensity of the target area. Based on the contrast distribution matrix, the system further calculates the local matching degree to evaluate the matching reliability of feature points between frames and form a local matching weight matrix, ultimately providing a weight basis for subsequent cross-frame feature fusion.

[0068] For example, in the UAV monitoring scenario, when the camera captures the subtle changes in the texture of the UAV boundary, the local matching degree evaluation can generate a gradient matrix by analyzing the gradient changes in the flight direction, and calculate the contrast of the texture characteristics in combination with the illumination conditions to generate a local matching weight matrix, providing accurate data support for dynamic monitoring.

[0069] Using the local matching weight matrix to perform local feature aggregation on the temporal feature set to obtain the preliminary fusion feature, performing inter-frame feature constraint calculation on the preliminary fusion feature to obtain the inter-frame feature constraint matrix, and using the inter-frame feature constraint matrix to perform cross-frame feature enhancement on the preliminary fusion feature to obtain the enhanced temporal feature.

[0070] By combining the local matching weight matrix, the system aggregates the key feature points in the temporal feature set to integrate the dynamic characteristics of the target in consecutive frames and generate the preliminary fusion feature. Specifically, the system performs inter-frame feature constraint calculation on the preliminary fusion feature, uses inter-frame correlation analysis to constrain the stability and continuity of the preliminary fusion feature, and generates an inter-frame feature constraint matrix to represent the consistency and correlation degree of feature points during the inter-frame change process. Finally, the inter-frame feature constraint matrix is used to perform further cross-frame feature enhancement on the preliminary fusion feature, making the temporal features of the target more stable and having higher coherence in the video sequence, generating the enhanced temporal feature to support subsequent trajectory optimization and association calculation.

[0071] For example, in dynamic target tracking, the system can optimize the flight path features of the UAV through the inter-frame feature constraint matrix, aggregate the feature points that change multiple times in consecutive frames into cross-frame enhanced temporal features, and at the same time eliminate the non-target characteristics caused by wind speed changes, thereby improving the coherence of the target trajectory.

[0072] In step S300, target region extraction is performed on all global associated features to obtain candidate target regions. Motion trend analysis is carried out on the candidate target regions to obtain the target motion trend. The target motion trend is used to perform time series modeling on the candidate target regions to obtain the target time series.

[0073] By analyzing all global associated features, the system uses a target region extraction algorithm to separate dynamic targets in the video frame and generate candidate target regions. Specifically, the system combines background separation technology and region screening methods to identify target regions that match the global associated features, while removing irrelevant background and noise data. Then, motion trend analysis is carried out on the candidate target regions. By calculating the motion direction and speed changes of the target in consecutive frames, the target motion trend is generated to characterize the dynamic characteristics of the target. Subsequently, based on the target motion trend, the system constructs a time series model, predicts the trajectory changes of the target through a recurrent neural network, generates the target time series, and provides the time dimension information for subsequent dynamic calculations.

[0074] For example, in an outdoor security monitoring scenario, when the camera captures a fast-moving pedestrian, the system can separate the pedestrian region through the target region extraction algorithm, and at the same time identify the moving direction and speed of the pedestrian through motion trend analysis, and generate a time series model to predict the potential next moving position of the pedestrian.

[0075] Dynamic state calculation is performed on the target time series and global associated features to obtain several target state vectors. Feature clustering is carried out on all target state vectors to obtain the initial target clustering clusters.

[0076] By using the collaborative analysis of the target time series and global associated features, the system models the dynamic state of the target and generates target state vectors. Specifically, the system combines a dynamic prediction algorithm to evaluate the motion parameters of the target, such as position, speed, acceleration, etc., and generates multi-dimensional state vectors to represent the dynamic change characteristics of the target in the video sequence. Subsequently, the system performs feature clustering analysis on all target state vectors. By using a clustering algorithm, the state vectors with similar motion characteristics and visual characteristics are grouped to generate the initial target clustering clusters, which are used to mark the distribution of the initial trajectory points of the target.

[0077] For example, in an unmanned aerial vehicle (UAV) tracking application, the system can generate the motion parameters of the UAV at different time points through dynamic state calculation, and use feature clustering to separate the flight state of the UAV from background noise data to obtain the initial target clustering clusters.

[0078] Cluster relationship calculation is carried out on the initial target clustering clusters to obtain the inter-cluster association matrix and the inter-cluster separation matrix. Based on the inter-cluster association matrix and the inter-cluster separation matrix, the initial target clustering clusters are adjusted for association to obtain the target trajectory nodes.

[0079] Through the analysis of the initial target clustering clusters, the system further calculates the relationships between clusters to optimize the distribution of target trajectory points. Specifically, the system uses the correlation analysis algorithm to calculate the inter-cluster correlation matrix to quantify the connectivity and similarity between clustering clusters, and at the same time calculates the inter-cluster separation matrix to evaluate the independence and trajectory differences between clustering clusters. Based on the correlation matrix and the separation matrix, the system performs dynamic association adjustment on the initial target clustering clusters to optimize the spatial and temporal distribution relationships between clustering clusters, and finally generates target trajectory nodes to accurately mark the dynamic path of the target.

[0080] For example, in a complex monitoring scenario, the system can optimize the trajectory nodes of multiple vehicles through the calculation of inter-cluster relationships, cluster the vehicles on the same path into one cluster, and at the same time separate the vehicles on different paths to generate accurate target trajectory nodes, providing accurate target position guidance for pan-tilt control.

[0081] In step S400, all target trajectory nodes are used to screen the key frames and ordinary frames to obtain key frame trajectory points and ordinary frame trajectory points, and trajectory segments are constructed for the key frame trajectory points and ordinary frame trajectory points to obtain the initial key frame trajectory segments and initial ordinary frame trajectory segments.

[0082] By screening the target trajectory nodes, the system extracts key frame trajectory points and ordinary frame trajectory points according to the characteristics of different frame types. Specifically, the system combines the time distribution and spatial position of the trajectory nodes, classifies the nodes with similar target motion characteristics as key frame trajectory points to accurately reflect the key motion trajectory of the target, and the remaining nodes are classified as ordinary frame trajectory points to assist in describing the global motion trend of the target. Subsequently, trajectory segments are constructed based on these trajectory points, and the trajectory points are connected into initial key frame trajectory segments and initial ordinary frame trajectory segments according to the time series and spatial position relationships to provide the basic framework of the target motion trajectory.

[0083] For example, in video monitoring, when the system tracks a high-speed moving car, it can generate key frame trajectory points when the car turns and decelerates through trajectory node screening, and define the trajectory points during the remaining smooth driving process as ordinary frame trajectory points, thus constructing a complete trajectory segment.

[0084] Perform trajectory smoothing calculation on the initial key frame trajectory segments and initial ordinary frame trajectory segments to obtain smooth key frame trajectory segments and smooth ordinary frame trajectory segments, and perform trajectory consistency calculation on the smooth key frame trajectory segments and smooth ordinary frame trajectory segments to obtain trajectory consistency parameters.

[0085] By performing trajectory smoothing calculations on the initial key-frame trajectory segments and initial ordinary-frame trajectory segments, the system corrects the noise points and abnormal points in the trajectory segments to improve the smoothness and continuity of the trajectories. Specifically, interpolation processing and curvature adjustment are performed on the trajectory segments using a smoothing algorithm to eliminate the sudden changes in the trajectories, generating smooth key-frame trajectory segments and smooth ordinary-frame trajectory segments. Subsequently, the system performs consistency calculations on the smoothed trajectory segments. By analyzing the consistency of the trajectory segments in terms of time and space, trajectory consistency parameters are generated, which serve as important indicators for measuring the optimization degree of the trajectory segments.

[0086] For example, in the scenario of UAV tracking, the smoothing calculation can eliminate the trajectory fluctuations caused by airflow interference during flight, and evaluate the smoothness of the UAV flight trajectory through consistency calculation for subsequent optimization.

[0087] Perform trajectory segment optimization on the trajectory consistency parameters to obtain the key-frame trajectory and the ordinary-frame trajectory. Input the key-frame trajectory and the ordinary-frame trajectory into a deep learning model for trajectory reconstruction calculation to obtain the tracking target trajectory.

[0088] By analyzing the trajectory consistency parameters, the system further optimizes the key-frame trajectory segments and ordinary-frame trajectory segments to generate the final key-frame trajectory and ordinary-frame trajectory. Specifically, the trajectory segment optimization combines a multi-objective optimization algorithm to globally adjust the spatio-temporal distribution of the trajectories, thereby improving the accuracy and robustness of the trajectories. Finally, the system inputs the optimized key-frame trajectory and ordinary-frame trajectory into a deep learning model for trajectory reconstruction calculation. By learning the trajectory patterns and background changes in complex scenarios through the model, a highly accurate and coherent tracking target trajectory is generated, providing comprehensive trajectory information for real-time pan-tilt adjustment.

[0089] For example, in a complex urban surveillance system, the trajectory reconstruction model can accurately reproduce the complete driving path of a vehicle in a crowded street and adjust the camera direction in real time through a pan-tilt to ensure that the vehicle is always within the surveillance range.

[0090] Among them, the steps of using all target trajectory nodes to screen key frames and ordinary frames to obtain key-frame trajectory points and ordinary-frame trajectory points, and constructing trajectory segments for the key-frame trajectory points and ordinary-frame trajectory points to obtain the initial key-frame trajectory segments and initial ordinary-frame trajectory segments include: performing spatial neighborhood analysis on the key frames and ordinary frames based on all target trajectory nodes to obtain spatial neighborhood features.

[0091] By performing spatial neighborhood analysis on all target trajectory nodes, the system identifies the spatial distribution characteristics around the trajectory nodes and extracts spatial neighborhood features. Specifically, the neighborhood analysis combines the positional relationship and distribution density of the trajectory nodes to analyze the local spatial patterns of each node to reflect the motion characteristics of the target in different regions, providing a reference basis for the subsequent screening process.

[0092] For example, in target tracking applications, spatial neighborhood analysis can identify the turning areas of vehicles at intersections and extract relevant trajectory nodes as spatial neighborhood features to optimize the construction of trajectories.

[0093] Perform a temporal consistency evaluation on the spatial neighborhood features to obtain a temporal consistency matrix, and use the temporal consistency matrix to filter all target trajectory nodes to obtain a preliminary set of trajectory nodes.

[0094] Through temporal consistency evaluation, the system combines spatial neighborhood features with temporal dimension information to generate a temporal consistency matrix, which is used to represent the temporal continuity of trajectory nodes. Specifically, temporal consistency evaluation analyzes the time series and motion trends of trajectory nodes, filters out nodes with high temporal coherence, and forms a preliminary set of trajectory nodes for further refining the distribution of trajectory points.

[0095] For example, in highway surveillance, temporal consistency evaluation can filter out a stable set of nodes from the continuously traveling trajectories of vehicles and eliminate invalid nodes caused by occlusion or camera switching.

[0096] Based on the preliminary set of trajectory nodes, allocate trajectory nodes to key frames and ordinary frames to obtain key-frame trajectory points and ordinary-frame trajectory points.

[0097] By classifying the preliminary set of trajectory nodes, the system assigns them as key-frame trajectory points and ordinary-frame trajectory points according to the temporal distribution characteristics of the trajectory points. Specifically, the system preferentially selects nodes with drastic temporal changes or significant motion characteristics as key-frame trajectory points, while other stable trajectory points are used as ordinary-frame trajectory points, providing a clear classification structure for constructing trajectory segments.

[0098] For example, in an unmanned aerial vehicle (UAV) tracking task, the system can preferentially assign the nodes during the sharp turns of the UAV as key-frame trajectory points, and the nodes during straight flight as ordinary-frame trajectory points to generate a trajectory point set with a clear structure.

[0099] Perform a motion direction constraint calculation on the key-frame trajectory points and ordinary-frame trajectory points to obtain a trajectory direction matrix, and based on the trajectory direction matrix, group the key-frame trajectory points and ordinary-frame trajectory points to obtain a set of trajectory point groups.

[0100] Through motion direction constraint calculation, the system analyzes the changes in the motion directions of key-frame trajectory points and ordinary-frame trajectory points to generate a trajectory direction matrix, which is used to represent the direction distribution of each trajectory point. Specifically, the system groups the trajectory points based on the trajectory direction matrix, and assigns the trajectory points with similar motion directions to the same set of trajectory point groups, laying a foundation for subsequent trajectory segment construction.

[0101] For example, in dynamic target tracking, trajectory point grouping can effectively distinguish different direction segments of a drone in a complex flight path to optimize the directional structure of the trajectory.

[0102] Based on the set of trajectory point groups, construct trajectory segments to obtain initial key-frame trajectory segments and initial ordinary-frame trajectory segments.

[0103] By processing the set of trajectory point groups, the system constructs initial key-frame trajectory segments and initial ordinary-frame trajectory segments according to the spatio-temporal relationship of the trajectory points. Specifically, the trajectory segment construction algorithm combines the time series and spatial distribution of the trajectory points, connects adjacent trajectory points one by one, and generates a preliminary motion trajectory framework of the target to support further optimization of the trajectory.

[0104] For example, in video surveillance, the system generates the motion path of a pedestrian in a crowded crowd through trajectory segment construction, and takes the key turning points of the pedestrian's motion as key-frame trajectory segments to provide input for subsequent deep learning trajectory optimization.

[0105] In this embodiment, by using a motion estimation algorithm to calculate the inter-frame optical flow of the original frame sequence of the target video, extract the inter-frame motion vectors and inter-frame change information, and extract the global change pattern and local change pattern, a multi-scale change feature map is constructed, so as to achieve precise division of the feature level and determine the low-level feature region and high-level feature region. Select feature sampling points through information entropy evaluation, and perform adaptive frame screening based on the feature sampling points to generate a feature frame sequence and candidate key frames. At the same time, perform temporal consistency analysis on the feature frame sequence to obtain temporal consistency features; perform motion stability analysis on the candidate key frames to generate key frame stability vectors to accurately screen key frames and distinguish ordinary frames. Further, through multi-scale feature extraction technology, perform feature matching calculation on the spatial feature set and the temporal feature set, generate matching feature pairs and perform local matching degree evaluation, generate a local matching weight matrix, and generate an enhanced feature map through cross-frame feature fusion and feature mapping to obtain an inter-frame correlation matrix and an inter-frame similarity matrix. Subsequently, perform associated feature aggregation on the enhanced feature map through the global matching weight to form global associated features. Based on the global associated features, the system generates a target state vector through dynamic state calculation, combines feature clustering and inter-cluster relationship calculation to optimize the initial target clustering clusters, and finally generates target trajectory nodes. Use the target trajectory nodes to construct and optimize the trajectory segments of the key frames and ordinary frames, and perform trajectory reconstruction through a deep learning model to generate a high-precision tracking target trajectory. This method effectively improves the robustness and accuracy of target tracking, and realizes real-time dynamic tracking and pan-tilt intelligent control in complex scenarios.

[0106] Embodiment 3: Such as Figure 2As shown in the figure, the present application also provides an image target tracking device 10 based on deep learning, including an acquisition module 11, a mapping module 12, an analysis module 13, and a calculation module 14.

[0107] The acquisition module 11 is mainly used to acquire the original frame sequence of the target video, perform adaptive frame screening on the original frame sequence, and obtain a plurality of feature frame sequences, key frames, and ordinary frames.

[0108] The mapping module 12 is mainly used to perform feature extraction and feature mapping on all feature frame sequences to obtain a plurality of enhanced feature maps, calculate the cross-frame correlation degree of all enhanced feature maps, and obtain a plurality of global correlation features.

[0109] The analysis module 13 is mainly used to perform dynamic target state estimation on all global correlation features to obtain a plurality of target state vectors, perform clustering correlation analysis on all target state vectors, and obtain the target trajectory nodes corresponding to each global correlation feature.

[0110] The calculation module 14 is mainly used to perform trajectory optimization calculation on the key frames and ordinary frames by using all target trajectory nodes to obtain the key frame trajectory and the ordinary frame trajectory, input the key frame trajectory and the ordinary frame trajectory into a preset deep learning model for trajectory reconstruction optimization, and obtain the tracking target trajectory.

[0111] In this embodiment, by using the acquisition module 11 to perform adaptive frame screening on the original frame sequence of the target video, the system can effectively screen out a plurality of feature frame sequences, key frames, and ordinary frames with key dynamic information, providing efficient and accurate frame data input for subsequent processing. The mapping module 12 performs feature extraction and feature mapping on all feature frame sequences, generates a plurality of enhanced feature maps by combining the multi-scale feature extraction method, and extracts the global correlation features between video frames by using the cross-frame correlation degree calculation method, significantly improving the correlation and expression ability of the features. The analysis module 13 performs dynamic modeling and state prediction of the target area on the global correlation features through dynamic target state estimation, generates a plurality of target state vectors; at the same time, uses the clustering correlation analysis method to group the target state vectors, and optimizes to obtain the target trajectory nodes corresponding to each global correlation feature, thereby providing high-quality basic data for the construction of the target trajectory. The calculation module 14 performs trajectory optimization calculation on the target trajectory nodes, divides the optimized trajectory points into key frame trajectories and ordinary frame trajectories, and inputs them into the deep learning model for trajectory reconstruction optimization, and finally generates an accurate tracking target trajectory. This device comprehensively utilizes the characteristics of deep learning and multi-module cooperation, improves the target tracking accuracy while enhancing the robustness in complex scenarios, and realizes real-time dynamic tracking of the target through pan-tilt control, meeting various actual application requirements.

[0112] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module can refer to the corresponding processes in the foregoing Embodiment 1 and will not be elaborated herein.

[0113] The structures, proportions, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the conditions under which the present invention can be implemented. Therefore, they do not have any technical substance. Any modification of the structure, change of the proportional relationship or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed in the present invention.

[0114] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image target tracking method based on deep learning, characterized in that, Including: Obtain the original frame sequence of the target video, perform adaptive frame screening on the original frame sequence to obtain a number of feature frame sequences, key frames, and ordinary frames; Perform feature extraction and feature mapping on all the feature frame sequences to obtain a number of enhanced feature mappings, and perform cross-frame correlation degree calculation on all the enhanced feature mappings to obtain a number of global correlation features; Perform dynamic target state estimation on all the global correlation features to obtain a number of target state vectors, and perform clustering correlation analysis on all the target state vectors to obtain target trajectory nodes corresponding to each global correlation feature; Use all the target trajectory nodes to perform trajectory optimization calculation on the key frames and the ordinary frames to obtain key frame trajectories and ordinary frame trajectories, and input the key frame trajectories and the ordinary frame trajectories into a preset deep learning model for trajectory reconstruction optimization to obtain the tracking target trajectory.

2. The image target tracking method based on deep learning according to claim 1, characterized in that The step of obtaining the original frame sequence of the target video, performing adaptive frame screening on the original frame sequence to obtain a number of feature frame sequences, key frames, and ordinary frames includes: Collect the original frame sequence of the target video, perform inter-frame optical flow calculation on the original frame sequence of the target video using a motion estimation algorithm to obtain inter-frame motion vectors and inter-frame change information, and use the inter-frame motion vectors to extract change patterns from the inter-frame change information to obtain global change patterns and local change patterns; Construct a multi-scale change feature map based on the global change pattern and the local change pattern, and perform feature level division on the multi-scale change feature map based on a preset multi-scale hierarchical sampling strategy to obtain a low-level feature region and a high-level feature region; Perform information entropy evaluation on the low-level feature region and the high-level feature region to obtain a low-level information entropy distribution and a high-level information entropy distribution, and select a number of feature sampling points based on the low-level information entropy distribution and the high-level information entropy distribution; Perform adaptive frame screening on the original frame sequence based on the feature sampling points to obtain a feature frame sequence and candidate key frames, perform temporal consistency analysis on the feature frame sequence to obtain temporal consistency features, and perform motion stability analysis on the candidate key frames to obtain key frame stability vectors; Use the temporal consistency features and the key frame stability vectors to screen the candidate key frames to obtain key frames, and remove the key frames from the feature frame sequence to obtain ordinary frames.

3. The image target tracking method based on deep learning according to claim 1, wherein The step of performing feature extraction and feature mapping on all the feature frame sequences to obtain a number of enhanced feature mappings, and performing cross-frame correlation degree calculation on all the enhanced feature mappings to obtain a number of global correlation features includes: Perform multi-scale feature extraction on all the feature frame sequences to obtain a spatial feature set and a temporal feature set, and perform feature matching calculation on the temporal feature set using the spatial feature set to obtain matching feature pairs; Perform local matching degree evaluation on all the matching feature pairs to obtain a local matching weight matrix, and perform cross-frame feature fusion on the temporal feature set based on the local matching weight matrix to obtain enhanced temporal features; Perform multi-layer feature mapping on the enhanced temporal features using the spatial feature set to obtain enhanced feature maps, and calculate the inter-frame correlation for all the enhanced feature maps to obtain an inter-frame correlation matrix and an inter-frame similarity matrix; Calculate the global matching weight based on the inter-frame correlation matrix and the inter-frame similarity matrix, and use the global matching weight to perform associated feature aggregation on the enhanced feature maps to obtain a number of global associated features.

4. The image target tracking method based on deep learning according to claim 3, characterized in that, The step of evaluating the local matching degree for all the matching feature pairs to obtain a local matching weight matrix, and performing cross-frame feature fusion on the temporal feature set based on the local matching weight matrix to obtain enhanced temporal features includes: Calculate the local gradient distribution of feature points based on all the matching feature pairs to obtain a feature point gradient matrix, perform local feature contrast calculation on the feature point gradient matrix to obtain a contrast distribution matrix, and calculate the local matching degree based on the contrast distribution matrix to obtain a local matching weight matrix; Use the local matching weight matrix to perform local feature aggregation on the temporal feature set to obtain a preliminary fusion feature, perform inter-frame feature constraint calculation on the preliminary fusion feature to obtain an inter-frame feature constraint matrix, and use the inter-frame feature constraint matrix to perform cross-frame feature enhancement on the preliminary fusion feature to obtain enhanced temporal features.

5. The image target tracking method based on deep learning according to claim 1, characterized in that The step of performing dynamic target state estimation on all the global associated features to obtain a number of target state vectors, and performing clustering association analysis on all the target state vectors to obtain the target trajectory nodes corresponding to each global associated feature includes: Extract the target regions for all the global associated features to obtain candidate target regions, analyze the motion trends of the candidate target regions to obtain the target motion trends, and use the target motion trends to perform time series modeling on the candidate target regions to obtain target time series; Perform dynamic state calculation on the target time series and the global associated features to obtain a number of target state vectors, and perform feature clustering on all the target state vectors to obtain initial target clustering clusters; Calculate the inter-cluster relationship for the initial target clustering clusters to obtain an inter-cluster association matrix and an inter-cluster separation matrix, and perform association adjustment on the initial target clustering clusters based on the inter-cluster association matrix and the inter-cluster separation matrix to obtain target trajectory nodes.

6. The method for image object tracking based on deep learning according to claim 1, characterized in that, The step of performing trajectory optimization calculation on the key frames and the ordinary frames using all the target trajectory nodes to obtain key frame trajectories and ordinary frame trajectories, and inputting the key frame trajectories and the ordinary frame trajectories into a preset deep learning model for trajectory reconstruction optimization to obtain the tracking target trajectory includes: Use all the target trajectory nodes to perform trajectory node screening on the key frames and the ordinary frames to obtain key frame trajectory points and ordinary frame trajectory points, and construct trajectory segments for the key frame trajectory points and the ordinary frame trajectory points to obtain initial key frame trajectory segments and initial ordinary frame trajectory segments; Perform trajectory smoothing calculation on the initial key-frame trajectory segment and the initial ordinary-frame trajectory segment to obtain a smoothed key-frame trajectory segment and a smoothed ordinary-frame trajectory segment, and perform trajectory consistency calculation on the smoothed key-frame trajectory segment and the smoothed ordinary-frame trajectory segment to obtain trajectory consistency parameters; Perform trajectory segment optimization on the trajectory consistency parameters to obtain a key-frame trajectory and an ordinary-frame trajectory, and input the key-frame trajectory and the ordinary-frame trajectory into the deep learning model for trajectory reconstruction calculation to obtain a tracking target trajectory.

7. The method for image object tracking based on deep learning according to claim 6, characterized in that, The step of using all the target trajectory nodes to perform trajectory node screening on the key frames and the ordinary frames to obtain key-frame trajectory points and ordinary-frame trajectory points, and performing trajectory segment construction on the key-frame trajectory points and the ordinary-frame trajectory points to obtain an initial key-frame trajectory segment and an initial ordinary-frame trajectory segment includes: Perform spatial neighborhood analysis on the key frames and the ordinary frames based on all the target trajectory nodes to obtain spatial neighborhood features; Perform temporal consistency evaluation on the spatial neighborhood features to obtain a temporal consistency matrix, and use the temporal consistency matrix to screen all the target trajectory nodes to obtain a preliminary trajectory node set; Perform trajectory node assignment on the key frames and the ordinary frames based on the preliminary trajectory node set to obtain key-frame trajectory points and ordinary-frame trajectory points; Perform motion direction constraint calculation on the key-frame trajectory points and the ordinary-frame trajectory points to obtain a trajectory direction matrix, and perform trajectory point grouping on the key-frame trajectory points and the ordinary-frame trajectory points based on the trajectory direction matrix to obtain a set of trajectory point groups; Perform trajectory segment construction based on the set of trajectory point groups to obtain an initial key-frame trajectory segment and an initial ordinary-frame trajectory segment.

8. An image target tracking device based on deep learning, characterized in that, Including: An acquisition module, configured to acquire an original frame sequence of a target video, and perform adaptive frame screening on the original frame sequence to obtain a plurality of feature frame sequences, key frames, and ordinary frames; A mapping module, configured to perform feature extraction and feature mapping on all the feature frame sequences to obtain a plurality of enhanced feature mappings, and perform cross-frame correlation degree calculation on all the enhanced feature mappings to obtain a plurality of global correlation features; An analysis module, configured to perform dynamic target state estimation on all the global correlation features to obtain a plurality of target state vectors, and perform clustering correlation analysis on all the target state vectors to obtain target trajectory nodes corresponding to each of the global correlation features; A calculation module, configured to use all the target trajectory nodes to perform trajectory optimization calculation on the key frames and the ordinary frames to obtain a key-frame trajectory and an ordinary-frame trajectory, and input the key-frame trajectory and the ordinary-frame trajectory into a preset deep learning model for trajectory reconstruction optimization to obtain a tracking target trajectory.

Citation Information

Cited By

  • Video target segmentation method based on interactive click

    CN120580630A

  • Blind sidewalk identification method based on image processing

    CN121121690A

  • Multispectral and thermal imaging data fused mountain nighttime animal tracking method and system

    CN121147981A

  • People flow channel track detection method, device, equipment and medium

    CN121482096A

  • Trajectory tracking method and system based on permanent magnet image

    CN122199619A