Target tracking method based on spatial conversion network and convolutional network
By combining spatial conversion network and convolutional network in the target tracking method, processing frame images and performing advanced feature extraction, the problems of low tracking accuracy and poor robustness of traditional methods in complex scenarios are solved, and more accurate and stable target tracking is achieved.
Patent Information
- Application Number
- CN202510632425.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Traditional goal tracking methods have problems with low tracking accuracy and poor robustness when dealing with fast moving targets and complex backgrounds. Although deep learning methods have made progress, they still face the challenges of spatial perception bias, algorithm complexity and speed difficult to balance.
The target tracking method based on spatial conversion network and convolutional network is adopted, and the target tracking edge is optimized by performing feature extraction, spatial transformation and target area focus on the frame image, combined with advanced feature extraction and time series analysis, and the target tracking edge is optimized and adjusted by moving the boundary difference set.
Accurate identification and tracking of goals in complex scenarios, reducing errors and drift phenomena, and improving the accuracy and stability of tracking.
Smart Images

Figure CN120219442A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision, and particularly to an object tracking method based on a spatial transformation network and a convolutional network. Background Art
[0002] In the field of object tracking, traditional methods have problems such as low tracking accuracy and poor robustness when dealing with situations such as fast-moving objects and similar background interference. With the development of deep learning, methods based on convolutional neural networks have made certain progress, but still face challenges such as spatial perception deviation and difficulty in balancing algorithm complexity and speed. For example, in complex scenarios, the fast movement of objects can lead to inaccurate prediction of spatial positions, and similar backgrounds are likely to cause misjudgment of objects, affecting the tracking effect.
[0003] Therefore, the present invention proposes an object tracking method based on a spatial transformation network and a convolutional network. Summary of the Invention
[0004] The present invention provides an object tracking method based on a spatial transformation network and a convolutional network to solve the technical problems raised in the above background art.
[0005] The present invention provides an object tracking method based on a spatial transformation network and a convolutional network, including: Step 1: Moving and video shooting a specified object along different motion trajectories in a complex scenario, and splitting the captured video into frames; Step 2: Extracting features from the frame images using a convolutional network to obtain a first feature map, and performing spatial transformation and target region focusing on the first feature map based on the spatial transformation network to obtain a second feature map; Step 3: Performing high-level feature extraction on the second feature map based on the convolutional network to lock the predicted target; Step 4: Sorting the locked positions of all predicted targets in sequence according to the frame order to obtain a predicted trajectory and comparing and analyzing it with the corresponding moving trajectory to obtain a moving boundary difference set, where the moving boundary difference set includes the shape of the boundary difference region of each frame image; Step 5: Optimizing and adjusting the object tracking boundary based on all moving boundary difference sets.
[0006] Preferably, the complex scenario is pre-deployed, and there are N1 complex scenarios, and there are N2 pre-set motion trajectories in each complex scenario.
[0007] Preferably, splitting the captured video into frames includes: Determine the movement duration of each movement trajectory, and obtain the unit split duration consistent with the movement duration from the duration-split comparison table. Perform corresponding splitting of blank trajectories for complex scenarios without movement trajectories to obtain a number of original frames; Analyze the color distribution similar to the target color of the specified target in each original frame, and set a priority for each original frame; Split the captured video under the movement trajectory according to the unit split duration to obtain comparison frames, and analyze the visual effect and coherence of each comparison frame; Based on the priority, visual effect, and coherence of the same frame, determine the refinement quantity of the corresponding frame; According to the sum of all refinement quantities, readjust the unit split duration, and perform frame splitting on the captured video according to the adjusted duration.
[0008] Preferably, based on the priority, visual effect, and coherence of the same frame, determining the refinement quantity of the corresponding frame includes:
[0009] Among them, represents the refinement quantity of the corresponding frame, represents the priority of the corresponding frame, represents the preset level of the corresponding frame, represents the preset effect coefficient of the corresponding frame; represents the visual effect coefficient of the corresponding frame; represents the presettability of the corresponding frame; represents the coherence of the corresponding frame; represents the ceiling symbol; represents a constant with a value of 2.7; ln represents the logarithmic function symbol.
[0010] Preferably, based on the convolutional network, performing high-level feature extraction on the second feature map includes: Construct a deep residual convolutional module, where the deep residual convolutional module contains multiple residual blocks, and each residual block fuses the features of different layers through skip connections; At the same time, introduce multi-scale convolutional kernels, and process the second feature map in parallel through convolutional kernels of different sizes to obtain feature information at different scales; Use the time series analysis method, combine the position, speed, and acceleration of the specified target in the previous frame, and predict the position and state of the specified target in the current frame; Combine the extracted high-level features and the predicted target position information to lock the predicted target.
[0011] Preferably, based on all the moving boundary difference sets, optimizing and adjusting the target tracking edge includes: Based on the set of moving boundary differences, the actual boundary and the standard boundary of the locked position of each image frame are respectively mapped into a two-dimensional coordinate system, the shape of the boundary difference region between the actual boundary and the standard boundary is determined, and the difference area Si, the maximum difference distance L1i, and the minimum difference distance L2i of the shape of the boundary difference region are determined, where the actual boundary is the boundary of the prediction target; If , at this time, assign 0 to the corresponding image frame, otherwise, assign 1 to the image frame, where represents the area difference threshold; represents the radius of the target object; represents the boundary recognition accuracy; represents the number of image frames in the corresponding captured video; Regard the image frame assigned 1 as the first frame, and extract the background under the shape of the boundary difference region corresponding to the first frame from the complex background, regarded as the blurred background; If the blurred background is of the expanding type, at this time, convert the blurred background and the image of the specified target from the RGB color space to the HSV color space and lock the target edge; If the blurred background is of the shrinking type, at this time, extract the texture features of the blurred background and the image of the specified target respectively based on the gray-level co-occurrence matrix, and lock the target edge according to the texture features; If the blurred background is both of the expanding type and the shrinking type, at this time, perform edge self-drawing according to the known shape features of the specified target and in combination with the standard trajectory center point in the corresponding image frame, and lock the target edge; Perform a re-difference analysis on the target edge and the standard edge in the corresponding image frame to obtain an updated difference vector; Project the non-zero elements in the updated difference image vector into the two-dimensional coordinate system respectively. If the difference area of each non-zero element is less than or equal to and the variance of the difference areas of all non-zero elements in the updated difference vector is less than a1, at this time, retain the target edge of each image frame; Otherwise, manually frame the target edge of each first frame and construct the difference distance set of the first frame, and obtain the labeled image of the first frame, and optimize and train the neural network until the analysis of all moving boundary difference sets is completed.
[0012] Preferably, manually frame the target edge of each first frame and construct the difference distance set of the first frame, and obtain the labeled image of the first frame, including: Significantly display the manually framed target edge on the first frame with the first color; Significantly display the distance lines on the first frame according to the set of difference distances, where each distance line starts from the corresponding point of the target edge and ends at the other end of the corresponding difference distance; Obtain an annotated image based on the significant display result.
[0013] Preferably, the updated difference vectors correspond one-to-one with the image frames in the corresponding moving boundary difference set, and the non-zero elements corresponding to the first frame image are assigned to the zero elements of the corresponding zero-assigned image.
[0014] Compared with the prior art, the beneficial effects of the present application are as follows: By moving a specified target according to multiple motion trajectories in a complex scenario and performing frame splitting, a large amount of frame image data containing different target motion states and complex environment information can be obtained. The basic features of the frame images are extracted through a convolutional network to obtain a first feature map, providing preliminary feature information for subsequent analysis. Then, the first feature map is processed by a spatial transformation network to obtain a second feature map, making the performance of the target in the feature map more prominent and regular. By performing high-level feature extraction on the second feature map and realizing the prediction target locking, the category and position of the target can be accurately identified in the frame images of the complex scenario. By comparing and analyzing the prediction trajectory and the moving trajectory to obtain a moving boundary difference set, the accuracy and reliability of the prediction target locking can be quantitatively evaluated. After optimizing and adjusting the target tracking edge based on the moving boundary difference set, the target tracking edge can more accurately reflect the actual position and shape of the target, reducing errors and drift phenomena during the target tracking process. In a complex scenario, even if the target is subject to various interferences and changes, the optimized and adjusted target tracking edge can more stably track the target. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a flowchart of a target tracking method based on a spatial transformation network and a convolutional network provided by an embodiment of the present invention. Detailed Embodiments
[0017] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] The present invention provides a target tracking method based on a spatial transformation network and a convolutional network, as Figure 1 shown, including: Step 1: Move and shoot videos of a specified target along different movement trajectories in a complex scene, and split the captured videos into frames; Step 2: Use a convolutional network to extract features from the frame images to obtain a first feature map, and perform spatial transformation and target area focusing on the first feature map based on the spatial transformation network to obtain a second feature map; Step 3: Perform high-level feature extraction on the second feature map based on the convolutional network to lock the predicted target; Step 4: Sort the locked positions of all predicted targets in sequence according to the frame order to obtain a predicted trajectory, and compare and analyze it with the corresponding movement trajectory to obtain a moving boundary difference set, where the moving boundary difference set includes the shape of the boundary difference region of each frame image; Step 5: Optimize and adjust the target tracking edge based on all the moving boundary difference sets.
[0019] Preferably, the complex scene is pre-deployed, and there are N1 complex scenes, and there are N2 preset movement trajectories in each complex scene.
[0020] In this embodiment, the complex scene: is an environment with high complexity and interference that has been pre-planned and arranged, and is used to simulate various situations that are not conducive to target recognition and tracking in reality. For example, in a complex scene simulating security monitoring, it may contain a large number of similar objects, complex background patterns, different lighting conditions, and dynamic interference objects; it is set that there are N1 different complex scenes, and each scene has a unique combination of environmental layout and interference factors.
[0021] In this embodiment, the specified target is an object that needs to be concerned and locked in research or application, such as a suspicious person in security monitoring, a specific vehicle in autonomous driving, etc.
[0022] In this embodiment, the motion trajectory is the path of a specified target moving in a complex scenario. N2 motion trajectories are preset. For example, in a simulated security monitoring scenario, a person may move in an "S" shape through a crowd, walk in a circle around a certain area, or move tortuously among multiple obstacles.
[0023] In this embodiment, frame splitting refers to decomposing the captured video into a series of independent static image frames, and each frame represents the picture at a certain moment in the video. For example, for a video with 25 frames per second, after frame splitting, the video content per second becomes 25 independent images. The movement process of the target is video-recorded from different angles using a high-definition camera to ensure complete recording of the target's movement in a complex scenario. After shooting, a video processing software or algorithm is used to perform frame splitting on the captured video to obtain a series of consecutive frame images, providing basic data for subsequent analysis and processing.
[0024] In this embodiment, the first feature map is the result obtained after the convolutional network performs preliminary feature extraction on the frame image. It contains basic feature information of the image, such as lines, simple shapes, and local textures, etc. The second feature map is the feature map obtained after the spatial transformation network performs spatial transformation and target area focusing on the first feature map. Compared with the first feature map, the features of the target in the second feature map are more prominent, and the background interference is further weakened, which is more conducive to subsequent high-level feature extraction and locking of the target. The convolutional network performs feature extraction on the frame image through its internal convolutional layers, pooling layers, etc. After multiple layers of operations, the first feature map is generated. For example, using a ResNet-50 network, the image sequentially passes through multiple residual blocks and pooling layers to extract the basic features of the image. Then, the first feature map is input into the spatial transformation network. The spatial transformation network performs spatial transformation on the first feature map by learning the position and pose information of the target in the image, such as rotating the target to an upright direction, enlarging the target area, etc., and at the same time uses the attention mechanism to focus on the area where the target is located, suppressing background noise and irrelevant information, thereby obtaining the second feature map.
[0025] In this embodiment, high-level features are more abstract and semantic features than basic features, and can reflect the essential attributes and category information of the target. For example, for a pedestrian target, high-level features may include the overall structure of the human body, action postures, styles of clothing worn, etc.; for a vehicle target, high-level features include the external contour of the vehicle, the shape of the headlights, the structure of the wheels, etc.
[0026] In this embodiment, the target locking is predicted based on the extracted high-level features, and the position of the target in the image is determined by an algorithm, and the category of the target is determined, and the target is marked in the image with a border or other means to achieve the locking of the target. The second feature map is input into the convolutional network again, and the deeper structure of the convolutional network is used, such as adding more convolutional layers, fully connected layers, etc., to further extract and abstract the second feature map, and dig out the high-level semantic features of the target. For example, in the deep layer of the network, through the stacking of multiple convolutional layers and the action of nonlinear activation functions, the high-level features such as the overall structure and functional components of the target are gradually extracted. Then, based on the extracted high-level features, a classifier (such as a Softmax classifier) is used to judge the category of the target, and a regressor (such as a Bounding-Box regressor) is used to predict the position of the target in the image, determine the bounding box of the target, and achieve the prediction of target locking.
[0027] In this embodiment, the specific coordinate position of the target in the frame image obtained by predicting the target locking is usually represented by the coordinates of the upper left corner and the lower right corner of the bounding box. The predicted trajectory is the target motion trajectory formed by connecting the locked positions of the target in all frame images in the order of the frames, and the moving trajectory is the actual motion path of the designated target pre-set in step 1 in the complex scene.
[0028] In this embodiment, the moving boundary difference set is a result set obtained after comparative analysis, which includes the shape information of the boundary difference area between the predicted trajectory and the moving trajectory in each frame image. For example, in a certain frame, there may be an irregular difference area between the predicted target boundary box and the boundary box corresponding to the actual target moving path, and the shape information of the area is included in the moving boundary difference set.
[0029] The target tracking edge is a boundary line or bounding box used to define the target range and track the target movement during the target tracking process. It is updated and adjusted in different frame images as the target moves. Optimization adjustment is to modify and improve the position, shape, size and other parameters of the target tracking edge based on the information provided by the moving boundary difference set to improve the accuracy and stability of target tracking. For example, if it is found that the bounding box of the predicted trajectory in a certain frame is smaller than the bounding box of the actual moving trajectory, and the difference area is concentrated on the right side of the target, then the range of the target tracking edge on the right side can be appropriately expanded; if the difference area presents an irregular shape, it means that the posture or position of the target has changed significantly, and the target tracking edge needs to be adjusted more finely, such as adjusting the angle and shape of the bounding box.
[0030] The beneficial effects of the above technical solution are as follows: By moving a specified target along multiple motion trajectories and performing frame splitting in a complex scenario, a large amount of frame image data containing different target motion states and complex environmental information can be obtained. The basic features of the frame images are extracted through a convolutional network to obtain a first feature map, providing preliminary feature information for subsequent analysis. Then, a spatial transformation network is used to process the first feature map to obtain a second feature map, making the performance of the target in the feature map more prominent and regular. By performing high-level feature extraction on the second feature map and achieving prediction target locking, the category and position of the target can be accurately identified in the frame images of the complex scenario. By comparing and analyzing the predicted trajectory and the moving trajectory to obtain a moving boundary difference set, the accuracy and reliability of the prediction target locking can be quantitatively evaluated. After optimizing and adjusting the target tracking edge based on the moving boundary difference set, the target tracking edge can more accurately reflect the actual position and shape of the target, reducing errors and drift phenomena during the target tracking process. In a complex scenario, even if the target is subject to various interferences and changes, the optimized and adjusted target tracking edge can more stably track the target.
[0031] The present invention provides a target tracking method based on a spatial transformation network and a convolutional network, which performs frame splitting on a captured video, including: Determine the motion duration of each motion trajectory, and obtain the unit splitting duration consistent with the motion duration from the duration-splitting comparison table. Perform corresponding splitting of blank trajectories on the complex scenario without a moving trajectory to obtain a number of original frames; Analyze the color distribution similar to the target color of the specified target in each original frame, and set a priority for each original frame; Split the captured video under the moving trajectory according to the unit splitting duration to obtain comparison frames, and analyze the visual effect and coherence of each comparison frame; Based on the priority, visual effect, and coherence of the same frame, determine the refinement quantity of the corresponding frame; According to the sum of all refinement quantities, readjust the unit splitting duration, and perform frame splitting on the captured video according to the adjusted duration.
[0032] Preferably, based on the priority, visual effect, and coherence of the same frame, determining the refinement quantity of the corresponding frame includes:
[0033] Wherein, represents the refinement quantity of the corresponding frame, represents the priority of the corresponding frame, represents the preset level of the corresponding frame, represents the preset effect coefficient of the corresponding frame; represents the visual effect coefficient of the corresponding frame; Represents the presupposition of the corresponding frame; Represents the coherence of the corresponding frame; Represents the ceiling symbol; Represents a constant with a value of 2.7; ln represents the natural logarithm function symbol.
[0034] In this embodiment, the values of presupposition and presupposition effect coefficient are both 1, and the value of presupposition level is 0.5. In video frame processing, the importance and quality of different frames vary. Simply relying on a single factor (such as only according to priority, or only according to visual effects, etc.) to determine the refinement quantity of frames cannot comprehensively and accurately reflect the actual processing requirements of frames. This formula comprehensively considers multiple key factors such as priority (Yx), presupposition level (R0), presupposition effect coefficient (G0), visual effect coefficient (Gx), presupposition (U0), and coherence (Ux): Priority: Reflects the importance of the frame in the entire video processing. Important frames often require more attention and processing. Visual effect coefficient and presupposition effect coefficient: Compare the actual visual effect with the expected effect, measure the quality of the visual effect. Frames with poor effects require more refinement processing to improve quality. Coherence and presupposition: Reflect the continuity of the frame in the video sequence and the difference from the expected continuity. Frames with poor coherence also require more processing to ensure video fluency. Through this formula that comprehensively considers multiple factors, the refinement quantity of each frame can be determined more accurately and comprehensively, making video frame processing more targeted and reasonable, avoiding over-processing or under-processing situations, and improving the overall video processing effect and efficiency.
[0035] Visual effect coefficient (Gx) = current clarity / best clarity, where the best clarity is set in advance.
[0036] Coefficient of coherence (Ux) = 1 - ((distance between matching feature points - minimum distance) / (maximum distance - minimum distance)).
[0037] In this embodiment, the duration-splitting comparison table: is a pre-established mapping table that records the corresponding relationship between different motion durations and unit splitting durations. For example, when the motion duration is 1 - 5 seconds, the unit splitting duration is 0.1 second, and when it is 5 - 10 seconds, the unit splitting duration is 0.2 second, etc., which is used to guide subsequent video splitting operations.
[0038] In this embodiment, the unit splitting duration is the basic unit of time interval when splitting the video into frames.
[0039] In this embodiment, the blank trajectory refers to a virtual trajectory concept set in a complex scenario where there is no actual moving trajectory of the specified target. It is used to uniformly process video splitting operations and obtain the initial frame images after splitting the complex scenario video. For example, in a simulated shopping mall surveillance scenario, the movement trajectory of pedestrian A lasts for 8 seconds. By referring to the comparison table, the unit splitting duration is determined to be 0.2 seconds. The video recording the movement of pedestrian A is split according to this duration to obtain the original frames. For the scene areas where there is no pedestrian movement, the video is also split according to the unit splitting duration of 0.2 seconds under the blank trajectory.
[0040] In this embodiment, the target color refers to the color attribute of the specified target. For example, for a red car, its target color is red, and similar colors such as dark red and light red affect the color of the red car itself.
[0041] In this embodiment, Priority = Area occupied by approximate color in the original frame / Total area of the original frame.
[0042] In this embodiment, the comparison frame is the frame image obtained after splitting the captured video under the moving trajectory according to the unit splitting duration for comparative analysis.
[0043] In this embodiment, the refinement quantity: a quantitative index for the degree of further processing (such as adding details, optimizing display, etc.) of the frame determined based on factors such as the priority, visual effect, and coherence of the original frame. For example, if the refinement quantity is 3, it means that the frame needs to be refined in 3 levels or steps.
[0044] In this embodiment, if the total sum of the refinement quantities is large, it indicates that the frames as a whole need more delicate processing, and the unit splitting duration may be appropriately shortened; otherwise, it may be appropriately lengthened. Then, based on the adjusted unit splitting duration, the captured video is split into frames again to obtain the final frame images for subsequent analysis or application.
[0045] The beneficial effects of the above technical solution are as follows: The video in the complex scenario is split into original frames at a reasonable time interval, providing basic materials for subsequent in-depth analysis and processing of the frame images. By setting priorities for the original frames, the importance of different frames in subsequent processing can be distinguished. By analyzing the visual effect and coherence of the comparison frames, the quality and content connection of the video frames can be comprehensively understood, providing an important basis for determining the refinement quantity of the frames. By calculating the refinement quantity, a personalized degree of further processing can be determined for each frame, making the subsequent processing of the frames more targeted. By adjusting the unit splitting duration according to the total sum of the refinement quantities and splitting the frames again, the actual processing requirements of the frames in the video can be dynamically adapted, and the granularity of frame splitting can be optimized.
[0046] The present invention provides an object tracking method based on a spatial transformation network and a convolutional network. Advanced feature extraction is performed on the second feature map based on the convolutional network, including: Construct a deep residual convolutional module, where the deep residual convolutional module contains multiple residual blocks, and each residual block fuses features of different layers through skip connections; Meanwhile, introduce multi-scale convolutional kernels, and process the second feature map in parallel through convolutional kernels of different sizes to obtain feature information at different scales; Using time series analysis method, combining the position, speed, and acceleration of the specified target in the previous frame, predict the position and state of the specified target in the current frame; Combine the extracted advanced features and the predicted target position information to lock the predicted target.
[0047] In this embodiment, the second feature map generation module processes the basic feature map, introduces a spatial attention mechanism, and weights different regions of the basic feature map according to the position and importance of the target in the image, highlighting the features of the region where the target is located, and generating the second feature map. The advanced feature extraction module uses the constructed deep residual convolutional module and multi-scale convolutional kernels to process the second feature map. Multiple residual blocks in the deep residual convolutional module fuse features of different layers through skip connections to learn more complex advanced semantic features; the multi-scale convolutional kernels respectively use convolutional kernels of different sizes such as 3×3, 5×5, 7×7 to process the second feature map in parallel to obtain feature information of the target at different scales.
[0048] The target motion prediction module collects the position information of the target in historical frames and inputs it into the LSTM network. The LSTM network predicts the position and motion state of the target in the next few frames by learning the motion laws in historical data.
[0049] The predicted target locking module combines the extracted advanced features and the target motion prediction results, judges the category of the target through a classifier, such as judging whether it is a pedestrian, a vehicle, etc., and precisely regresses the position of the target through a regressor to determine the specific coordinates of the target. When the classification confidence of the target is higher than a set threshold, such as 0.8, the system considers that the target is successfully locked and marks the target with a border in the picture.
[0050] The beneficial effects of the above technical solution are: by extracting the advanced semantic features of the target and combining target motion prediction, more accurate and timely target locking is achieved.
[0051] The present invention provides an object tracking method based on a spatial transformation network and a convolutional network. Optimize and adjust the target tracking edge based on all mobile boundary difference sets, including: Based on the moving boundary difference set, the actual boundary and the standard boundary of the locked position of each image frame are respectively mapped into a two-dimensional coordinate system, the shape of the boundary difference region between the actual boundary and the standard boundary is determined, and the difference area Si, the maximum difference distance L1i, and the minimum difference distance L2i of the shape of the boundary difference region are determined, where the actual boundary is the boundary of the prediction target; If , at this time, assign 0 to the corresponding image frame, otherwise, assign 1 to the image frame, where represents the area difference threshold; represents the radius of the target object; represents the boundary recognition accuracy; represents the number of image frames in the corresponding captured video; Regard the image frame assigned 1 as the first frame, and extract the background under the shape of the boundary difference region corresponding to the first frame from the complex background, regarded as the blurred background; If the blurred background is of the expanding type, at this time, convert the blurred background and the image of the specified target from the RGB color space to the HSV color space, and lock the target edge; If the blurred background is of the shrinking type, at this time, extract the texture features of the blurred background and the image of the specified target respectively based on the gray-level co-occurrence matrix, and lock the target edge according to the texture features; If the blurred background is both of the expanding type and the shrinking type, at this time, perform edge self-drawing according to the known shape features of the specified target and in combination with the standard trajectory center point under the corresponding image frame, and lock the target edge; Perform a second difference analysis based on the target edge and the standard edge under the corresponding image frame to obtain an updated difference vector; Project the non-zero elements in the updated difference image vector into the two-dimensional coordinate system respectively. If the difference area of each non-zero element is less than or equal to and the variance of the difference areas of all non-zero elements in the updated difference vector is less than a1, at this time, retain the target edge of each image frame; Otherwise, manually select the target edge of each first frame and construct the difference distance set of the first frame, and obtain the labeled image of the first frame, and optimize the training of the neural network until the analysis of all moving boundary difference sets is completed.
[0052] Preferably, manually selecting the target edge of each first frame and constructing the difference distance set of the first frame, and obtaining the labeled image of the first frame includes: Significantly display the manually selected target edge on the first frame with the first color; Significantly display the distance lines on the first frame using a second color display according to the set of difference distances, where each of the distance lines starts from the corresponding point of the target edge and ends at the other end of the corresponding difference distance; Obtain an annotation image based on the significant display result.
[0053] Preferably, the updated difference vectors correspond one-to-one with the image frames in the corresponding moving boundary difference set, and the non-zero elements corresponding to the first frame image are given zero elements corresponding to the zero image.
[0054] In this embodiment, the actual boundary refers to the boundary range of the target obtained by the target locking algorithm in the image frame, usually represented in the form of a rectangular box, polygon, etc. to indicate the position and size of the target in the image. The standard boundary is a pre-set ideal boundary representing the true position and shape of the target, which can be determined based on manual annotation, high-precision sensor data or other reliable references and is used for comparison and evaluation with the actual boundary.
[0055] In this embodiment, the geometric shape presented by the area where there is a difference between the actual boundary and the standard boundary may be various shapes such as an irregular polygon or an arc, reflecting the spatial inconsistency between the predicted boundary and the true boundary. The difference area Si is the area size of the boundary difference region and is used to quantify the degree of difference between the actual boundary and the standard boundary. The larger the difference area, the greater the deviation of the target locking. The maximum difference distance L1i is the maximum distance from a point on the actual boundary to the corresponding point on the standard boundary within the boundary difference region, reflecting the maximum deviation in distance between the two. The minimum difference distance L2i is the minimum distance from a point on the actual boundary to the corresponding point on the standard boundary within the boundary difference region, reflecting the minimum deviation in distance between the two.
[0056] In this embodiment, the area difference threshold is a pre-set critical value for determining whether the area of the boundary difference region is acceptable, with a value of 5 square centimeters. The value of the boundary recognition accuracy is 0.9.
[0057] In this embodiment, the blurred background is the background part related to the target boundary difference region extracted from the complex background. This background may be difficult to distinguish from the target due to inaccurate target locking, background interference, etc., and needs further processing to achieve accurate locking of the target edge.
[0058] In this embodiment, the expanded type of blurred background means that the range of the blurred background extends outward relative to the actual range of the target, that is, the blurred background includes some areas that do not belong to the target but are similar to the target in terms of features such as color and texture, resulting in difficulty in accurately defining the target edge. The contracted type of blurred background is the opposite of the expanded type. The range of the blurred background shrinks inward relative to the actual range of the target, causing some edges of the target to be misidentified as the background, which also brings difficulties to locking the target edge. To determine the type of the blurred background, if it is the expanded type, the images of the blurred background and the specified target are converted from the RGB color space to the HSV color space. In the HSV space, the differences between the target and the background in the three components of hue, saturation, and value are analyzed, and the color segmentation algorithm (such as threshold segmentation, clustering segmentation, etc.) is used to lock the target edge according to the differences. If the blurred background is of the contracted type, the texture features of the blurred background and the specified target image are respectively extracted based on the gray-level co-occurrence matrix, and the eigenvalues of the gray-level co-occurrence matrix, such as contrast, correlation, etc., are calculated. By comparing the texture feature differences between the target and the background, the texture analysis algorithm (such as the classification algorithm based on eigenvalues) is used to lock the target edge. If the blurred background is both of the expanded type and the contracted type, according to the known shape features (such as circular, rectangular, etc.) of the specified target, combined with the standard trajectory center point in the corresponding image frame, the graphic drawing algorithm (such as Bezier curve drawing, polygon drawing, etc.) is used for self-drawing of the edge, so as to lock the target edge.
[0059] In this embodiment, when the analysis result of the updated difference vector does not meet the requirements, the accurate edge of the target is outlined on each first-frame image manually, and a specific drawing tool or software is used to prominently display the outlined target edge on the image with the first color (such as red). Then, according to the difference distance set, distance lines are drawn on the first-frame image with the second color (such as blue), starting from the corresponding point of the target edge and ending at the other end of the corresponding difference distance, to prominently display the difference distance situation. Based on these prominent display results, an annotated image is obtained. The annotated image is used as training data to optimize the training of the neural network, and the parameters of the neural network are continuously adjusted until the analysis of all moving boundary difference sets is completed, that is, the target edge locking results of all image frames reach a satisfactory accuracy.
[0060] The beneficial effects of the above technical solution are as follows: By determining the shape and related parameters of the boundary difference region, the accuracy of target locking can be visually and quantitatively evaluated. By assigning numerical values to the image frames, the image frames with large target locking deviations can be quickly screened out, and the image frames are divided into two categories, laying a foundation for subsequent targeted processing of image frames in different situations. By focusing on the blurred background, the problem that the target and the background have similar colors and are difficult to distinguish can be solved more effectively, creating conditions for accurately locking the target edge subsequently. By adopting corresponding methods to lock the target edge according to different types of blurred backgrounds, the characteristic differences between the target and the background in terms of color, texture, shape, etc. can be fully utilized, effectively solving the problem that the target and the blurred background are difficult to distinguish. By performing differential analysis again and processing the updated difference vectors, a comprehensive evaluation of the re-locked target edge can be carried out to ensure that the finally locked target edge has high accuracy. Manually selecting and constructing the annotation image can utilize human subjective judgment and experience to provide accurate training samples for the neural network. By optimizing the training of the neural network, the neural network can learn more accurate target edge locking methods, improving the performance of the neural network in dealing with complex backgrounds and target locking deviations.
[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A target tracking method based on a spatial transformer network and a convolutional network, characterized in that: include: Step 1: Move and shoot videos of designated targets according to different motion trajectories in complex scenes, and split the captured videos into frames; Step 2: extracting features from the frame image using a convolutional network to obtain a first feature map, and performing spatial transformation and target area focusing on the first feature map based on the spatial transformation network to obtain a second feature map; Step 3: extracting high-level features from the second feature map based on the convolutional network to predict target locking; Step 4: Based on the locked positions of all predicted targets and sorting them in frame order, the predicted trajectories are obtained and compared with the corresponding moving trajectories to obtain a moving boundary difference set, wherein the moving boundary difference set includes the boundary difference area shape of each frame image; Step 5: Optimize and adjust the target tracking edge based on all moving boundary difference sets.
2. The target tracking method based on the spatial transformer network and the convolutional network according to claim 1, characterized in that: The complex scenes are deployed in advance, and there are N1 complex scenes, and there are N2 pre-set motion trajectories in each complex scene.
3. The target tracking method based on spatial transformer network and convolutional network according to claim 1, characterized in that: Frame splitting of captured video, including: Determine the motion duration of each motion track, and obtain the unit split duration consistent with the motion duration from the duration-split comparison table, and perform corresponding splitting of blank tracks for complex scenes without moving tracks to obtain a number of original frames; Analyze the distribution of colors close to the target color of the designated target in each raw frame, and set a priority to each raw frame; Splitting the captured video under the moving trajectory according to the unit splitting time to obtain comparison frames, and analyzing the visual effect and coherence of each comparison frame; Determine the number of refinements of the corresponding frame based on the priority, visual effect and coherence of the same frame; According to the sum of all the refined quantities, the unit splitting duration is readjusted, and the captured video is frame split according to the adjusted duration.
4. The target tracking method based on spatial transformation network and convolutional network according to claim 3 is characterized in that: Based on the priority, visual effect, and coherence of the same frame, the number of refinements of the corresponding frame is determined, including: in, represents the number of refinements of the corresponding frame, Indicates the priority of the corresponding frame, Indicates the preset level of the corresponding frame, Indicates the preset effect coefficient of the corresponding frame; Represents the visual effect coefficient of the corresponding frame; Indicates the preset nature of the corresponding frame; Indicates the coherence of corresponding frames; Indicates the rounding up symbol; represents a constant, whose value is 2.7; ln represents the sign of logarithmic function.
5. The target tracking method based on spatial transformer network and convolutional network according to claim 1, characterized in that: Performing high-level feature extraction on the second feature map based on the convolutional network includes: Construct a deep residual convolution module, where the deep residual convolution module contains multiple residual blocks, and each residual block fuses the features of different layers through jump connections; At the same time, a multi-scale convolution kernel is introduced to process the second feature map in parallel through convolution kernels of different sizes to obtain feature information at different scales; Using a time series analysis method, combined with the position, velocity and acceleration of the designated target in the previous frame, the position and state of the designated target in the current frame are predicted; Combine the extracted high-level features with the predicted target location information to lock the predicted target.
6. The target tracking method based on spatial transformer network and convolutional network according to claim 1, characterized in that: Optimize the target tracking edge based on all moving boundary difference sets, including: Based on the moving boundary difference set, the actual boundary and the standard boundary of the locked position of each image frame are respectively mapped into a two-dimensional coordinate system, and the boundary difference area shape between the actual boundary and the standard boundary is determined, and the difference area Si, the maximum difference distance L1i and the minimum difference distance L2i of the boundary difference area shape are determined, wherein the actual boundary is the boundary of the predicted target; like , at this time, assign 0 to the corresponding image frame, otherwise, assign 1 to the image frame, where, represents the area difference threshold; Indicates the radius of the target object; Indicates the boundary recognition accuracy; Indicates the number of image frames in the corresponding captured video; The image frame assigned with 1 is regarded as the first frame, and the background of the first frame corresponding to the boundary difference area shape is extracted from the complex background and regarded as the blurred background; If the blurred background is of the expansion type, then the blurred background and the designated target image are converted from the RGB color space to the HSV color space, and the target edge is locked; If the blurred background is of the indented type, then the texture features of the blurred background and the image of the designated target are extracted based on the gray level co-occurrence matrix, and the target edge is locked according to the texture features; If the blurred background is both of the expansion type and the contraction type, at this time, the edge is self-drawn according to the known shape features of the designated target and combined with the standard track center point under the corresponding image frame to lock the target edge; Performing a difference analysis again based on the target edge and the standard edge in the corresponding image frame to obtain an updated difference vector; Project the non-zero elements in the updated difference image vector into the two-dimensional coordinate system respectively. If the difference area of each non-zero element is less than or equal to and the variance of the difference areas of all non-zero elements in the updated difference vector is less than a1, in which case the target edge of each image frame is retained; Otherwise, the target edge of each first frame is manually selected and the difference distance set of the first frame is constructed, and the labeled image of the first frame is obtained, and the neural network is optimized and trained until the analysis of all moving boundary difference sets is completed.
7. The target tracking method based on spatial transformer network and convolutional network according to claim 6, characterized in that: Artificially selecting the target edge of each first frame and constructing a difference distance set of the first frame, and obtaining a labeled image of the first frame, including: The edge of the manually selected target is displayed prominently in the first frame using the first color; According to the difference distance set, the distance lines are prominently displayed on the first frame using a second color display, wherein each distance line starts from a corresponding point of the target edge and ends at the other end of the corresponding difference distance; The labeled image is obtained based on the saliency display results.
8. The target tracking method based on spatial transformer network and convolutional network according to claim 6, characterized in that: The updated difference vector corresponds one-to-one to the image frames in the corresponding moving boundary difference set, and the non-zero elements corresponding to the first frame image are assigned with zero elements corresponding to the image with zero.
Citation Information
Patent Citations
Table tennis target tracking and trajectory predicting method and device, storage medium, and computer equipment
CN107481270A
Target tracking method and system based on triple convolutional network and perception interference learning
CN110349176A
Target tracking method based on graph convolution and trajectory convolution network learning
CN110660082A
Multi-target tracking method fusing spatial motion and apparent feature learning
CN115994929A
Intelligent camera target automatic identification and tracking method and system
CN119007106A