Multi-target smear-free tracking method and device in high dynamic scene
By employing a dual discriminator network and adaptive scale compensation parameters, combined with optical flow estimation and the Hungarian algorithm, the scale jump and motion blur problems in multi-target tracking under high dynamic scenes are solved, achieving high-precision and robust multi-target tracking.
Patent Information
- Application Number
- CN202511170442.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-12-12
AI Technical Summary
In highly dynamic scenarios, traditional multi-target tracking methods cannot effectively cope with abrupt changes in target scale, severe motion blur, complex scale interactions and identity associations among multiple targets, resulting in decreased tracking accuracy and insufficient continuity.
A dual discriminator network is used for global and local feature discrimination, adaptive scale compensation parameters are calculated, and multi-scale de-ghosting prediction templates and optical flow estimation are combined. Motion is eliminated by dense optical flow field calculation, and identity association is performed by combining the Hungarian algorithm. Camera exposure time and frame rate are optimized to improve image quality.
It effectively solves the problems of abrupt target scale changes and motion blur in highly dynamic scenes, improves tracking accuracy and continuity, and enhances robustness and accuracy in complex environments.
Smart Images

Figure CN121120690A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of target tracking, and in particular to a multi-target no-dragging tracking method and device in a high dynamic scene. BACKGROUND
[0002] In a high-speed motion environment, the camera needs to track multiple moving targets in real time and accurately to meet the needs of intelligent systems for target state monitoring and prediction. However, high dynamic scenes have characteristics such as fast target motion, dramatic scale changes, and complex imaging conditions, which pose great challenges to traditional multi-target tracking methods.
[0003] Existing multi-target tracking methods face serious technical bottlenecks in high dynamic scenes. Traditional methods generally use fixed scale kernel functions and single motion models, which cannot effectively deal with the problem of dramatic scale jumps, resulting in insufficient scale compensation parameters and a sharp decline in tracking accuracy. Secondly, the motion dragging phenomenon commonly existing in high-speed motion environments seriously affects target feature extraction and matching, and existing technologies lack effective drag elimination mechanisms, causing discontinuous target trajectory distribution and motion blur background interference. In addition, the scale interaction and identity association problems between multiple targets are more complex in high dynamic scenes, and traditional association methods cannot guarantee the continuity and accuracy of tracking. SUMMARY
[0004] The present application provides a multi-target no-dragging tracking method and device in a high dynamic scene, which significantly improves the tracking performance of traditional methods in motion blur conditions and effectively solves the identity switching and occlusion problems in high dynamic scenes.
[0005] In a first aspect, the present application provides a multi-target no-dragging tracking method in a high dynamic scene, which comprises: obtaining a sequence of continuous frame images of a camera in a high dynamic scene, and extracting motion parameters from the sequence of continuous frame images to obtain a sequence of multiple target speed changes and bounding box scale changes; inputting the sequence of multiple target speed changes and the sequence of bounding box scale changes into a double discriminator network for global and local feature discrimination to obtain a global discrimination feature vector and a local discrimination feature vector; calculating an adaptive scale compensation parameter according to the global discrimination feature vector and the local discrimination feature vector; performing multi-source scale fusion based on the adaptive scale compensation parameter to obtain a multi-scale de-dragging prediction template; performing optical flow estimation and drag elimination based on the multi-scale de-dragging prediction template to obtain a first multi-target tracking result, and performing scale consistency identity association on the first multi-target tracking result to output a second multi-target tracking result.
[0006] With reference to the first aspect, in a first implementation form of the first aspect of the present application, the method further includes: obtaining a continuous frame image sequence in a high dynamic scene, and performing motion parameter extraction on the continuous frame image sequence to obtain a multi-target speed change quantity and a bounding box scale change sequence, including: adjusting exposure time and frame rate parameters of the camera according to the detected target motion speed to obtain a continuous frame image sequence in a high dynamic scene; performing target detection and position tracking on the continuous frame image sequence to obtain a multi-target center position coordinate sequence and a bounding box coordinate sequence; calculating a target center position coordinate difference value between adjacent frames based on the multi-target center position coordinate sequence to obtain a multi-target speed change quantity, and calculating a bounding box scale change sequence based on the bounding box coordinate sequence.
[0007] With reference to the first aspect, in a second implementation form of the first aspect of the present application, the method further includes: inputting the multi-target speed change quantity and the bounding box scale change sequence into a dual discriminator network to perform global and local feature discrimination to obtain a global discrimination feature vector and a local discrimination feature vector, including: inputting the multi-target speed change quantity into a first discriminator of the dual discriminator network to perform global motion pattern feature extraction to obtain a global motion trend intermediate feature; inputting the bounding box scale change sequence into a second discriminator of the dual discriminator network to perform local scale change feature extraction to obtain a local scale change intermediate feature; performing four-layer down-sampling convolution processing on the global motion trend intermediate feature to obtain a global deep feature mapping; performing jump connection fusion on the local scale change intermediate feature, and performing feature reconstruction calculation through a decoder level to obtain a local reconstructed feature mapping; respectively performing scale perception attention mechanism analysis on the global deep feature mapping and the local reconstructed feature mapping to obtain a global discrimination feature vector and a local discrimination feature vector.
[0008] With reference to the first aspect, in a third implementation form of the first aspect of the present application, the method further includes: inputting the multi-target speed change quantity into a first discriminator of the dual discriminator network to perform global motion pattern feature extraction to obtain a global motion trend intermediate feature, including: performing time sequence arrangement on the multi-target speed change quantity, and constructing a time sequence motion vector sequence in combination with historical frame motion data; inputting the time sequence motion vector sequence and corresponding whole frame image data into an input layer of the first discriminator to perform data preprocessing to obtain a global motion input feature matrix; input the global motion input feature matrix into a multi-layer convolutional neural network, each layer of the multi-layer convolutional neural network uses a different size convolution kernel to perform spatial feature calculation, and a multi-layer global motion feature map is obtained; Based on the multi-layer global motion feature map, a global motion mode recognition analysis is performed to obtain a global motion trend intermediate feature.
[0009] In combination with the first aspect, in a fourth implementation manner of the first aspect of the application, the adaptive scale compensation parameter is calculated according to the global discriminant feature vector and the local discriminant feature vector, including: Based on the global discriminant feature vector, a historical scale change trend analysis is performed to obtain a scale adaptive coefficient, and a motion speed influence factor calculation is performed in combination with the local discriminant feature vector to obtain a motion speed attenuation factor; The scale adaptive coefficient and the motion speed attenuation factor are input into an exponential function calculation of a scale compensation factor dynamic update equation to obtain a scale compensation factor update value; Based on the scale compensation factor update value and current target scale data, a scale compensation prediction is performed to obtain a predicted scale compensation value; The predicted scale compensation value is subjected to a compensation boundary constraint detection and a hyperbolic tangent function smoothing processing to obtain an adaptive scale compensation parameter.
[0010] In combination with the first aspect, in a fifth implementation manner of the first aspect of the application, the scale adaptive coefficient and the motion speed attenuation factor are input into an exponential function calculation of a scale compensation factor dynamic update equation to obtain a scale compensation factor update value, including: Based on current frame target scale state data, a historical scale compensation factor is extracted to obtain a current scale compensation factor reference value; The scale adaptive coefficient and the motion speed attenuation factor are input into a parameter configuration of a scale compensation factor dynamic update equation to obtain a configured dynamic update equation; The product of the motion speed attenuation factor and the multi-target speed change amount in the configured dynamic update equation is subjected to an exponential attenuation calculation to obtain an exponential attenuation calculation result; The current scale compensation factor reference value is subjected to a multiplication operation with the scale adaptive coefficient and the exponential attenuation calculation result to obtain a scale compensation factor update value.
[0011] In combination with the first aspect, in a sixth implementation manner of the first aspect of the application, the multi-source scale fusion is performed based on the adaptive scale compensation parameter to obtain a multi-scale de-smearing prediction template, including: The adaptive scale compensation parameter is subtracted from current scale data of each target to obtain a scale credibility weight coefficient of each target; Feature pyramid multi-level extraction is performed on each target region in the continuous frame image sequence to obtain a multi-scale feature vector of each target, and a corresponding de-smearing gain coefficient is calculated based on a detection confidence of each target through a Sigmoid activation function; Element-wise multiplication is performed on the scale credibility weight coefficient, the multi-scale feature vector and the de-smearing gain coefficient, and the product results of all targets are summed and accumulated to obtain a multi-target scale fusion feature vector; The multi-target scale fusion feature vector and the multi-target velocity change amount are spliced and combined, and input into a generator of a Wasserstein generative adversarial network for multi-scale template analysis to generate a multi-scale de-smearing prediction template.
[0012] In combination with the first aspect, in a seventh implementation manner of the first aspect of the present application, the light flow estimation and smearing elimination are performed based on the multi-scale de-smearing prediction template to obtain a first multi-target tracking result, and scale consistency identity association is performed on the first multi-target tracking result to output a second multi-target tracking result, including: Dense optical flow field calculation and motion smearing region identification are performed on the continuous frame image sequence to generate de-smearing mask data; Template matching processing is performed on the multi-scale de-smearing prediction template and the de-smearing mask data, the best matching position of each target in the current frame is calculated through a normalized cross-correlation function, and target state information is updated to obtain a first multi-target tracking result; Scale consistency features are generated based on the first multi-target tracking result, and global optimal identity association matching is performed in combination with a Hungarian algorithm to obtain multi-target tracking data after association matching; Target life cycle management and trajectory continuity optimization are performed on the multi-target tracking data after association matching to generate a second multi-target tracking result.
[0013] In combination with the first aspect, in an eighth implementation manner of the first aspect of the present application, scale consistency features are generated based on the first multi-target tracking result, and global optimal identity association matching is performed in combination with a Hungarian algorithm to obtain multi-target tracking data after association matching, including: Appearance feature vectors and motion feature vectors of each target are extracted from the first multi-target tracking result, and scale consistency features are generated based on scale data difference values of each target; The appearance feature vectors, the motion feature vectors and the scale consistency features are weighted and fused to obtain a target identity association probability distribution matrix. construct a cost matrix of the Hungarian algorithm based on the target identity association probability distribution matrix and motion consistency measure data; input the cost matrix into the Hungarian algorithm for global optimal allocation calculation to obtain a minimum cost matching scheme, and generate multi-target tracking data after association matching according to the minimum cost matching scheme.
[0014] In a second aspect, the present application provides a multi-target no-dragging tracking device in a high dynamic scene, which comprises: An acquisition module is configured to acquire a continuous frame image sequence of a camera in a high dynamic scene, and extract motion parameters of the continuous frame image sequence to obtain a multi-target speed variation and a bounding box scale variation sequence; A feature discrimination module is configured to input the multi-target speed variation and the bounding box scale variation sequence into a double discriminator network for global and local feature discrimination to obtain a global discrimination feature vector and a local discrimination feature vector; A calculation module is configured to calculate an adaptive scale compensation parameter according to the global discrimination feature vector and the local discrimination feature vector; A multi-source scale fusion module is configured to perform multi-source scale fusion based on the adaptive scale compensation parameter to obtain a multi-scale no-dragging prediction template; A multi-target tracking module is configured to perform optical flow estimation and drag elimination based on the multi-scale no-dragging prediction template to obtain a first multi-target tracking result, and perform scale consistency identity association on the first multi-target tracking result to output a second multi-target tracking result.
[0015] In the technical scheme provided by the application, the adaptive scale compensation parameter calculation method based on the double discriminator features is constructed, the compensation strategy can be dynamically adjusted according to the target motion state and the scale change history, the limitations of the traditional fixed compensation coefficient method are overcome, and the tracking failure problem caused by the sharp scale jump of the target in the high dynamic scene is effectively solved. The double discrimination mechanism of the first discriminator processing global motion trend and the second discriminator processing local scale change can capture the global and local motion-scale correlation information at the same time, compared with the single discriminator structure, has stronger feature expression ability and discrimination precision, through the three-element weighted fusion processing of the scale reliability weight coefficient, the multi-scale feature vector and the de-smearing gain coefficient, the weight can be dynamically allocated according to the scale change reliability of each target, the scale interaction between multiple targets is effectively processed, and the feature conflict and information loss problems caused by the simple feature splicing method are avoided. The motion smearing elimination algorithm based on the dense optical flow field calculation and the space-time consistency constraint can accurately identify and eliminate the smearing interference caused by high-speed motion, realizes non-smearing tracking through the accurate matching of the de-smearing mask data and the prediction template, and significantly improves the tracking performance of the traditional method under the motion blur condition. The scale consistency feature is taken as an important basis for identity association, combined with the global optimal allocation strategy of the Hungarian algorithm, the identity switching and occlusion problems in the high dynamic scene can be effectively solved, compared with the traditional association method which only depends on the appearance feature, has stronger robustness and accuracy. According to the mechanism of adjusting the camera exposure time and frame rate parameters in real time according to the target motion speed, the imaging quality can be optimized at the hardware level, and the tracking conditions in the high dynamic scene are improved from the source.
[0016] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application will be realized and attained by the structure particularly pointed out in the description, claims and drawings.
[0017] In order to make the above-mentioned purpose, characteristics and advantages of the present application more obvious and easy to understand, the following preferred embodiments are specifically described below, and the accompanying drawings are shown as follows. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 An embodiment schematic diagram of the multi-target non-smearing tracking method in the high dynamic scene in the embodiment of the present application is shown. Figure 2 An embodiment schematic diagram of the multi-target non-smearing tracking device in the high dynamic scene in the embodiment of the present application is shown. DETAILED DESCRIPTION
[0019] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0020] The terms "comprising" and "having" and any variations thereof mentioned in the embodiments of the present application are intended to cover the inclusions not exclusively. For example, the processes, methods, systems, products or devices comprising a series of steps or units are not limited to the listed steps or units, but optionally further comprise other steps or units not listed, or optionally further comprise other steps or units inherent to the processes, methods, products or devices.
[0021] In order to facilitate the understanding of the embodiments, first, a high dynamic scene multi-target no-dragging tracking method disclosed by the embodiments of the present application is described in detail. As shown in Figure 1 The method comprises the following steps: 101. Obtain a continuous frame image sequence of a camera in a high dynamic scene, and perform motion parameter extraction on the continuous frame image sequence to obtain a multi-target speed change quantity and a bounding box scale change sequence; It can be understood that the execution subject of the present application can be a high dynamic scene multi-target no-dragging tracking device, and can also be a terminal or a server, and the specific execution subject is not limited here. The embodiments of the present application take the server as the execution subject for example.
[0022] Specifically, by detecting the motion speed information of targets in continuous frame images in real time, the camera's exposure time and frame rate are dynamically adjusted to avoid image ghosting caused by excessive exposure in high-speed motion scenes and motion blur caused by a fixed frame rate. The system compares the detected target motion speed in the current frame with a preset threshold. When the target speed exceeds the threshold, the exposure time is shortened accordingly. Compression is performed based on the relationship between the base exposure time, the speed sensitivity coefficient, and the maximum detection speed, thus achieving dynamic adjustment of the exposure time. Simultaneously, the frame rate is adjusted synchronously according to the motion speed, increasing the image sampling frequency, thereby forming a continuous frame image sequence with good image clarity under different motion speed conditions. Target detection and position tracking are performed on the continuous frame image sequence, extracting the center position coordinates and corresponding bounding box coordinates of each target in each frame. The target detection module uses a robust deep convolutional neural network structure, which can stably extract target position information in complex backgrounds and highly dynamic environments. Combined with multi-target tracking algorithms such as Kalman filtering or Hungarian matching, the consistency of target identity is maintained, forming a continuous and complete sequence of multi-target center position coordinates and bounding box coordinates. Based on this, the target center position coordinates between adjacent frames are differentially calculated. By analyzing the displacement changes between two adjacent frames, the displacement of the target between consecutive frames is obtained, thereby deriving the velocity changes of multiple targets. Simultaneously, based on the bounding box coordinate sequence of consecutive frames, the bounding box scale change between adjacent frames is calculated, recording the scale fluctuations caused by attitude changes, viewpoint changes, or distance changes during the target's motion, thus obtaining a scale change sequence.
[0023] 102. Input the multi-target velocity change and bounding box scale change sequence into a dual discriminator network to perform global and local feature discrimination, and obtain global and local discriminant feature vectors; Specifically, the multi-target speed change quantity is input into a first discriminator of the double discriminator network. The discriminator takes the speed change sequence as the input feature, and relies on the convolutional neural network to extract the global motion pattern features. Through the stacking of multiple layers of convolution and nonlinear activation units, the motion rules and trend information hidden in the speed change are gradually excavated, and the intermediate feature expression of the global motion trend is generated. These features can effectively reflect the acceleration, deceleration, turning and other change trends of the overall motion of the target. The global intermediate features are processed by four layers of down-sampling convolution. By gradually reducing the size of the feature map, not only the feature data volume is compressed and the network calculation efficiency is improved, but also the receptive field is gradually increased, so that the obtained global deep feature map can aggregate more extensive motion information, and the sensitivity and discrimination of the network to large-scale motion changes are enhanced. At the same time, the bounding box scale change sequence is input into a second discriminator of the double discriminator network. The discriminator analyzes and models the local scale change characteristics, and preliminarily extracts the intermediate features of the local scale change caused by the distance change, posture rotation and other reasons between consecutive frames. In order to retain more fine-grained scale change information, a skip connection mechanism is introduced in the processing process to fuse the features extracted by different convolution layers, avoiding the loss of local detail information in the deep feature extraction process. On the basis of the features after the skip connection, feature reconstruction calculation is performed through the decoder level, and the spatial resolution of the feature map is gradually restored, so that the local scale change information can be more completely retained and strengthened to form the local reconstructed feature map. Finally, scale perception attention mechanism analysis is performed on the global deep feature map and the local reconstructed feature map respectively. By introducing a dynamic weight adjustment strategy, the network can adaptively adjust the weight distribution of different scale features according to the scale change degree of the current target, so as to highlight the information with more discriminative value for the tracking task during feature fusion. The scale perception attention mechanism automatically allocates the attention degree of different feature channels by analyzing the scale change trend in the global deep feature map and the local reconstructed feature map, so that the network can focus more accurately on the target area with dramatic motion change or significant scale change. The global discriminative feature vector and the local discriminative feature vector are output respectively.
[0024] The multi-target speed change amount is arranged in time sequence, and the historical motion data of continuous multiple frames is combined to construct a time sequence motion vector sequence. Through the time sequence arrangement mode, the dynamic characteristics of the target in the time dimension are captured, and then the complex motion modes such as acceleration, deceleration and turning of the target are effectively represented. In order to enhance the adaptability of the model to high dynamic changes, while constructing the time sequence vector sequence, the speed change accumulation information of the historical frames is introduced, so that the motion features not only contain the information of the current frame, but also imply the motion trend in the past period of time, thereby ensuring that the input features have stronger time sequence continuity and dynamic description capability. The time sequence motion vector sequence and the corresponding whole frame image data are jointly input to form a global motion input feature matrix. Before the data enters the discriminator, preprocessing operations are performed through the input layer, including normalization processing, space-time alignment and feature enhancement. The normalization processing ensures that the motion features of targets with different scales and different speeds are within the same order of magnitude range, avoiding unstable training caused by large differences in feature amplitude; the space-time alignment ensures that the motion vector and the image data correspond in the time dimension, eliminating information misplacement caused by frame rate differences or synchronization errors; the feature enhancement introduces slight perturbations through data enhancement means to improve the robustness of the model. After preprocessing, a global motion input feature matrix is formed. The global motion input feature matrix is input into a multi-layer convolutional neural network for feature extraction. Different sizes of convolution kernels are used in each convolutional network to calculate spatial features. Smaller size convolution kernels are used to extract local motion detail features, such as small amplitude acceleration, slight turning and other subtle changes; while larger size convolution kernels focus on capturing large range motion trend information, such as rapid movement, overall displacement direction change and other global dynamics. Through the cooperation of multi-scale convolution kernels, the system can simultaneously perceive the target motion characteristics at different spatial scales, and construct multi-level global motion feature maps. These feature maps are stacked layer by layer inside the convolutional neural network, and after layer-by-layer convolution, activation and pooling operations, they gradually converge into a global motion feature set with high-dimensional expression capability. Finally, global motion pattern recognition analysis is performed based on the multi-level global motion feature map. By introducing high-level feature fusion and attention mechanism, the information weight between different feature levels is automatically adjusted to highlight the motion trend features with higher discriminative value and suppress redundant or noise information. Through this process, the intermediate feature expression representing the global motion trend of the target is finally obtained.
[0025] 103、according to the global discriminative feature vector and the local discriminative feature vector, calculating an adaptive scale compensation parameter; Specifically, based on the global discriminative feature vector, the historical scale change trend is analyzed, and relying on the trajectory data of the scale change of the target in the continuous frames, the statistical characteristics of the scale change are extracted, such as the mean, variance and change rate of the scale change, etc. Through the analysis of these statistical characteristics, the scale adaptive coefficient is dynamically generated. The coefficient can reflect the change trend and stability of the target scale in the historical stage. If the target scale changes smoothly, the adaptive coefficient tends to keep the original scale unchanged, and if the target scale changes dramatically, the coefficient is adjusted accordingly to enhance the response ability to the scale fluctuation. At the same time, combined with the local discriminative feature vector, the motion speed influence factor is calculated. Through the joint analysis of local scale change and motion state, the speed influence on the scale change of the target in the fast moving process is captured, and the motion speed attenuation factor is derived. The attenuation factor reflects the inhibition or amplification effect of the target motion speed on the scale change amplitude, and also enables the scale compensation mechanism to adaptively adjust the compensation amplitude according to the current motion state of the target, avoiding the dramatic fluctuation of scale estimation in high-speed motion. The scale adaptive coefficient and the motion speed attenuation factor are input into the scale compensation factor dynamic update equation together, and the exponential function operation is performed to obtain the update value of the scale compensation factor. Through the introduction of the exponential function, the compensation update process has nonlinear regulation ability, so that the compensation value can adaptively enhance or inhibit the compensation strength when facing different amplitude of motion change. Based on the update value of the scale compensation factor, combined with the scale data of the current target, the scale compensation prediction is executed to generate the predicted scale compensation value. The prediction value can dynamically adjust the target scale based on the full consideration of the historical scale evolution trend and the current motion state, and improve the accuracy of scale estimation and the robustness of tracking. In order to prevent the predicted scale compensation value from appearing excessive fluctuation or abnormal drift in the update process, a compensation boundary constraint detection mechanism is designed. Through detecting whether the compensation update amplitude exceeds the set threshold, if it exceeds, the smoothing processing mechanism is triggered, and the hyperbolic tangent function is used to smooth the scale compensation value. The hyperbolic tangent function has the characteristics of amplitude limitation and smooth transition, which can suppress the interference of abnormal values on the overall scale estimation without destroying the compensation trend, so as to obtain the final stable and adaptive scale compensation parameter.
[0026] The history scale compensation factor is extracted based on the target scale state data of the current frame. According to the scale evolution track of the target in the previous several frames, the scale compensation factor reference value corresponding to the current frame is extracted. Through the history reference extraction, the long-term trend of scale change is captured, and the influence of short-term scale fluctuation on the overall compensation strategy is prevented, thereby improving the smoothness and robustness of the scale updating process. The scale adaptive coefficient and the motion speed attenuation factor are input into the scale compensation factor dynamic updating equation for parameter configuration. The parameter configuration process is based on the comprehensive analysis of the current motion characteristics and scale evolution state of the target, dynamically determines the parameter weight and calculation logic in the updating equation, and ensures that the updating equation can accurately reflect the motion and scale characteristics of the target at the current time. This configuration not only involves assigning the scale adaptive coefficient and the motion speed attenuation factor to the corresponding positions in the updating equation, but also includes adjusting the internal coefficients in the equation according to the target motion complexity and scale fluctuation amplitude, further enhancing the adaptability of the dynamic updating mechanism to different scene changes, and ensuring that the updating stability and accuracy of the scale compensation factor can be maintained in the case of high-speed motion or dramatic scale changes. After completing the parameter configuration, the product of the motion speed attenuation factor and the multi-target speed change amount in the configured dynamic updating equation is calculated by exponential attenuation. Through the exponential function processing of the product, the nonlinear control of the scale compensation adjustment amplitude caused by the speed change is realized, and the overlarge fluctuation of the scale estimation caused by the dramatic change of the motion speed is avoided. Finally, the current scale compensation factor reference value is multiplied by the scale adaptive coefficient and the exponential attenuation calculation result to generate the final scale compensation factor update value. Through the multi-factor joint operation mode, the update value can comprehensively reflect the target history scale change trend, current scale stability, and the actual influence of motion speed on scale adjustment.
[0027] 104、based on adaptive scale compensation parameters, multi-source scale fusion is performed to obtain a multi-scale de-smearing prediction template; Specifically, the adaptive scale compensation parameter is subtracted from the scale data of each target to obtain a scale reliability weight coefficient by comparing the deviation between the compensated predicted scale and the actual detected scale. The scale reliability weight effectively reflects the reliability of the current scale estimation of each target. The smaller the scale difference, the better the compensation effect and the higher the reliability. The feature pyramid multi-level extraction is performed on each target region in the sequence of continuous frame images. The feature pyramid structure can extract rich spatial information at different scale levels, so that the system can obtain stable and diversified feature representation even in the case of severe target scale change. For each target, the multi-level feature vector at different scales is extracted, covering the multi-scale characteristics from the detail layer to the global structure layer. At the same time, the detection confidence of each target is combined to calculate a de-smearing gain coefficient by nonlinear mapping of the confidence through a Sigmoid activation function. The de-smearing gain coefficient reflects the importance of the target detection result in removing smearing. The higher the confidence of the target, the larger the gain coefficient, which can play a stronger guiding role in subsequent fusion. While the low-quality detection result is appropriately suppressed through the Sigmoid compression function, avoiding the negative impact of low-quality detection results on the overall de-smearing effect. The scale reliability weight coefficient, multi-scale feature vector and de-smearing gain coefficient are multiplied element by element. Through element-by-element operation, the feature strength of each dimension is dynamically adjusted in the feature space. The product results of all targets are summed and accumulated to generate a multi-target scale fusion feature vector after fusion. Finally, the multi-target scale fusion feature vector and the multi-target velocity change are spliced and combined to form a composite feature vector containing motion characteristics and scale information. The composite feature vector is input into the generator part of the Wasserstein generative adversarial network. The generator performs multi-scale template analysis based on the input fusion feature, learns and generates a multi-scale de-smearing prediction template. The generator continuously optimizes to improve the quality and de-smearing ability of the generated template. The final output template not only has accurate scale compensation effect, but also effectively eliminates the smearing phenomenon caused by high dynamic motion.
[0028] 105. Perform optical flow estimation and smearing elimination based on the multi-scale de-smearing prediction template to obtain a first multi-target tracking result, and perform scale consistency identity association on the first multi-target tracking result to output a second multi-target tracking result.
[0029] Specifically, dense optical flow field calculation is performed on the continuous frame image sequence. The high-precision optical flow estimation algorithm is used to analyze the pixel-level displacement change between adjacent frames to generate dense optical flow field data. The dense optical flow field can not only capture the overall motion trend of the target, but also reflect the slight motion change of the local area of the target. Based on the generated optical flow field, the motion smear area is identified, and by analyzing the amplitude and direction consistency of the optical flow vector, the blurred area formed due to high-speed motion is detected, and then the corresponding de-smear mask data is generated. The mask data marks the area in the image affected by the smear, which is used as a key constraint condition in the subsequent template matching process to ensure that only the effective image area is used for accurate matching and updating of the target position. The multiscale de-smear prediction template is combined with the de-smear mask data obtained by optical flow estimation, and the best matching position of each target in the current frame is determined through template matching processing. Specifically, the normalized cross-correlation function is used to calculate the similarity between the prediction template and the candidate area in the continuous frame image to find the most matching position of the template and the image area, thereby determining the accurate position coordinates of the target. After the matching is completed, the state information of the target is updated based on the matching position, including the center position, scale information, and tracking confidence, and the first-stage multi-target tracking result is obtained. In order to maintain the consistency of the identities of the targets in the multi-target tracking process, scale consistency features are generated based on the first multi-target tracking result. The scale consistency features are calculated based on the scale change rule of the target in the continuous frames to reflect the scale stability of the target in the time dimension. The scale consistency features, appearance features, and motion features are combined to construct a comprehensive feature vector, and the importance of each feature in identity discrimination is balanced through weighted combination. The Hungarian algorithm is used for global optimal identity association matching. The Hungarian algorithm considers the similarity of the features of each target to find the minimum cost matching path, ensuring that the identity assignment between the multi-targets is optimal as a whole, and avoiding identity drift or incorrect association caused by local optimal matching. Through the multi-feature fusion matching strategy combining scale consistency with motion and appearance features, the accuracy of identity association in complex scenes such as occlusion and crossing motion is improved. After the association matching is completed, the multi-target tracking data after association matching is subjected to target life cycle management and trajectory continuity optimization. The life cycle management module dynamically adjusts the activation and invalidation states of the targets according to the time sequence features of the appearance and disappearance of the targets, assigns a new identity number when a new target appears stably, and marks an existing target as invalid when it is not matched for a certain number of consecutive frames, preventing false detection results from interfering with the overall tracking process. At the same time, based on the Kalman filter or the long-short-term prediction model, the target trajectory is subjected to continuity optimization processing to smooth the trajectory changes, fill in the trajectory interruptions caused by short-term occlusion, and improve the overall coherence and physical rationality of the trajectory. After the above steps, the second-stage multi-target tracking result is finally generated.
[0030] The appearance feature vector and the motion feature vector of each target are extracted from the first multi-target tracking result. The appearance feature vector is mainly extracted from the target region by a deep convolutional network, which can reflect the static appearance attributes of the target such as color distribution, texture details and shape contour, and ensure that the targets can still be distinguished in a complex environment with many similar appearance targets. The motion feature vector is constructed by analyzing the dynamic characteristics such as the position change, speed and acceleration of the target between consecutive frames, which describes the motion trajectory and behavior pattern of the target in space-time, and provides additional dynamic behavior basis for identity association. At the same time of extracting the appearance and motion features, the scale consistency feature is generated according to the difference of the current scale data of each target. By calculating the difference of the scale change between adjacent frames or adjacent detection results, the stability index of the scale change is derived. The smaller the scale change amplitude is, the higher the scale consistency is, which indicates that the target has maintained stable physical size in the continuous time period, while the abnormal scale change indicates the problem of identity switching or detection abnormality. Through the introduction of the scale consistency feature, the identity distinguishing ability in the case of similar target appearance or motion trajectory intersection can be effectively enhanced, and the accuracy and stability of the overall association can be improved. The appearance feature vector, the motion feature vector and the scale consistency feature are weighted and fused to form a target identity association probability distribution matrix. In the process of weighted fusion, the weight distribution is dynamically adjusted according to the importance of each type of feature in different scenes, for example, the weight of the appearance feature is increased in the scene with high appearance feature discrimination, the weight of the motion feature is increased in the scene with high motion feature discrimination, and the scale consistency is introduced into the final score calculation as a stability guarantee. The probability distribution matrix obtained by weighted fusion represents the association confidence between a specific target and a historical trajectory. Based on the target identity association probability distribution matrix and the motion consistency measurement data, a cost matrix of the Hungarian algorithm is constructed. The motion consistency measurement data measures the similarity of the motion patterns between different targets based on the motion trend, direction consistency and speed change law of the target in the continuous time period. By considering the identity association probability and the motion consistency, the system reasonably matches the cost value of each pair of targets in the cost matrix. The smaller the cost value is, the higher the rationality of the matching is, which reflects the overall consistency between the targets in appearance, motion and scale. The cost matrix is input into the Hungarian algorithm for global optimal allocation calculation. The Hungarian algorithm finds the target correspondence relationship with the minimum overall cost by solving the optimal matching of the cost matrix, avoiding the problems of identity drift and false association caused by local optimization. Based on the minimum cost matching scheme, the system updates the identity label of each target and generates multi-target tracking data after association matching.
[0031] In a specific embodiment, the process of performing step 101 can specifically include the following steps: According to the detected target motion speed, the exposure time and frame rate parameters of the camera are dynamically adjusted, and a continuous frame image sequence under a high dynamic scene is obtained; Target detection and position tracking are performed on the continuous frame image sequence, and a multi-target center position coordinate sequence and a bounding box coordinate sequence are obtained. Based on the multi-target center position coordinate sequence, the target center position coordinate difference between adjacent frames is calculated, and the multi-target speed change quantity is obtained, and based on the bounding box coordinate sequence, the bounding box size change sequence is calculated.
[0032] Specifically, the imaging parameters of the camera are dynamically adjusted in real time according to the target motion state, ensuring that clear and continuous frame image sequences can still be obtained in a high-speed motion environment. By detecting the motion speed of the target in real time, the detected speed data is used as the basis for adjustment, and the exposure time and frame rate parameters of the camera are adjusted synchronously. When the detected motion speed of the target increases, the exposure time is shortened in time to reduce motion blur, because too long exposure time will cause the moving target in the image to have obvious trailing. At the same time, the frame rate is increased to increase the sampling frequency, ensuring that more continuous pictures are captured in unit time, thereby improving the temporal resolution of the image and enabling the fast moving track of the target to be recorded more completely. The shortening of the exposure time and the increase of the frame rate are coordinated within a certain range, and the system needs to dynamically balance the brightness and clarity of the image to avoid the image being too dark due to too short exposure time, and to prevent excessive rise of transmission bandwidth and storage pressure caused by too high frame rate. Through the dynamic adjustment strategy, the camera can adapt to the motion state of the target and continuously output high-quality image sequences that meet the subsequent processing requirements. The target detection and position tracking are performed on the continuous frame image sequences. The target detection module is based on a deep convolutional neural network, which uses its feature extraction and classification capabilities to accurately identify all moving targets in the image and generate a corresponding bounding box for each target on the image, recording the spatial position information and size characteristics of the target. In order to ensure the real-time and accuracy of detection, the detection network uses a lightweight structure to speed up the inference, and at the same time, a large number of high dynamic scene samples are used for reinforcement learning during the training stage to improve the detection robustness under high-speed motion conditions. After detection, based on algorithms such as Kalman filtering, Hungarian matching or deep feature correlation tracking, the time sequence correlation of the target position is performed, the motion trajectory of the target across frames is established, the consistency of the target identity is ensured, and the tracking loss caused by fast movement or short occlusion is avoided. Through continuous frame detection and tracking, the system can update the center position coordinate sequence and the bounding box coordinate sequence of each target in real time, forming a motion and scale change record. Based on the extracted multi-target center position coordinate sequence, the target center position difference between adjacent frames is calculated. For each target, the center position coordinates in the continuous frames are extracted, and the Euclidean distance or other distance functions suitable for measuring motion changes between the center position coordinates of adjacent two frames are calculated, thereby obtaining the displacement of the target. By combining the displacement with the time interval between frames, the speed change of the target in the time dimension is derived, reflecting the motion state change of the target in different time periods. The speed change can capture the overall motion trend of the target and reveal complex motion patterns such as acceleration, deceleration and turning. At the same time, the scale change sequence of the bounding box is calculated based on the bounding box coordinate sequence of the multi-target.By extracting the bounding box size parameters such as width and height of each target in consecutive frames, the change amplitude of the bounding box size of adjacent frames is calculated to obtain the scale change quantity, which reveals the apparent size fluctuation of the target in the motion process due to the change of distance from the camera, the transformation of viewing angle or the adjustment of posture. These changes are helpful to identify the real motion state and physical properties of the target. Especially in high dynamic scenes, the target often accompanies the behavior of quickly approaching or moving away from the camera, and the scale change sequence can effectively assist in judging whether the target is moving along the plane or there is rapid movement in the depth direction.
[0033] In a specific embodiment, the process of performing step 102 can specifically include the following steps: The multi-target speed change quantity is input into the first discriminator of the double discriminator network for global motion mode feature extraction to obtain global motion trend intermediate features; The bounding box scale change sequence is input into the second discriminator of the double discriminator network for local scale change feature extraction to obtain local scale change intermediate features; The global motion trend intermediate features are subjected to four-layer down-sampling convolution processing to obtain a global deep feature map; The local scale change intermediate features are subjected to jump connection fusion and feature reconstruction calculation through the decoder level to obtain a local reconstruction feature map; The global deep feature map and the local reconstruction feature map are respectively subjected to scale perception attention mechanism analysis to obtain a global discriminant feature vector and a local discriminant feature vector.
[0034] Specifically, the multi-target velocity change quantity is input into the first discriminator of the dual discriminator network. The task of the first discriminator is to model and analyze the overall motion trend of the target group. In order to fully capture the dynamic characteristics of the motion pattern, the velocity change quantity is arranged in time sequence in the input stage, maintaining its time continuity, ensuring that the input motion data not only contains the current speed state, but also retains the historical change trend. By inputting this time sequence motion information into the first discriminator, the network can understand the complex dynamic behaviors such as acceleration, deceleration, turning, etc. of the target on a longer time scale. The first discriminator internally uses a multi-layer convolutional neural network structure, with each layer of convolution kernel configured with different receptive field size, mining local patterns and global trends in motion changes through different scale convolution operations. Convolution processing not only improves the diversity of feature extraction, but also enhances the network's perception ability of different speed change patterns, thereby generating global motion trend intermediate features that can describe the motion pattern change. At the same time, the bounding box scale change sequence is input into the second discriminator of the dual discriminator network. The second discriminator focuses on capturing the dynamic characteristics of the local scale change of the target. Unlike motion patterns, scale changes reflect the relative distance changes, angle changes and posture transformations between the target and the camera, so the analysis of scale features helps to improve the model's adaptability to target shape changes. The scale change sequence is also arranged in time sequence to maintain continuity, ensuring that the network can understand the historical trend of scale changes. The second discriminator also uses a deep convolutional neural network internally, but its convolution kernel design is more biased towards capturing fine-grained changes, suitable for identifying small fluctuations in scale within a short period of time. Through a series of convolution and pooling operations, the second discriminator extracts local scale change intermediate features that can reflect the detailed information of the target's scale transformation in consecutive frames. The global motion trend intermediate features are subjected to four layers of down-sampling convolution processing. Each layer of down-sampling operation gradually reduces the feature map size through convolution and stride setting, expands the receptive field, so that the feature map can integrate more context information. The down-sampling process improves the abstraction level of the features and effectively suppresses the interference of local noise on the understanding of the motion pattern, obtaining a global deep feature map. The local scale change intermediate features are fused through jump connection, and the feature reconstruction calculation is performed through the decoder level. The decoder uses the jump connection mechanism to fuse the shallow features and deep features in the spatial dimension, preserving the details of the shallow layer and the abstract semantics of the deep layer, so that the decoder can restore the spatial resolution while taking into account the fine characteristics of the local scale change. Jump connection improves the expressiveness of features and improves the coherence of information flow in the feature reconstruction process, avoiding information loss caused by the singleization of deep features. Through layer-by-layer upsampling and convolution operations, the decoder gradually restores the feature map size, generating a local reconstructed feature map that completely retains the historical trajectory and current scale state of the target scale change.To enhance the dynamic adaptability of feature representation, scale-aware attention mechanisms are introduced for global deep feature mapping and local reconstructed feature mapping. The scale-aware attention mechanism learns the importance distribution of features at different scales adaptively and dynamically adjusts the feature weight distribution according to the current motion state and scale change characteristics of the target. The attention mechanism introduces a learnable weight matrix to perform weighted summation on feature channels, so that the network can pay more attention to the feature regions with large motion amplitude and severe scale change at high dynamic moments, and appropriately reduce the attention to irrelevant regions when the target motion is smooth or the scale is stable, reducing the interference of redundant features. Through the dynamic weighting strategy, the system generates global discriminative feature vectors and local discriminative feature vectors from the global motion pattern and local scale change dimensions respectively. The global discriminative feature vector can accurately describe the motion trend change of the target in a large range, while the local discriminative feature vector can precisely depict the fine-grained variation of the target in the scale dimension.
[0035] In a specific embodiment, the process of performing the step of inputting the multi-target speed change quantity into the first discriminator of the dual discriminator network for global motion pattern feature extraction to obtain the global motion trend intermediate feature can specifically include the following steps: The multi-target speed change quantity is arranged in time sequence, and the time sequence motion vector sequence is constructed in combination with the historical frame motion data; The time sequence motion vector sequence and the corresponding whole frame image data are input into the input layer of the first discriminator for data preprocessing to obtain a global motion input feature matrix; The global motion input feature matrix is input into a multi-layer convolutional neural network, and different sizes of convolution kernels are used in each layer of the multi-layer convolutional neural network to calculate spatial features, thereby obtaining multi-level global motion feature maps; Based on the multi-level global motion feature maps, global motion pattern recognition analysis is performed to obtain the global motion trend intermediate feature.
[0036] Specifically, the time correlation of the motion state between consecutive frames is established. For each detected target, the amount of speed change in consecutive frames is arranged in time sequence, maintaining the time sequence continuity of the speed change, thereby forming a preliminary speed sequence reflecting the dynamic change trajectory of the target. In order to enrich the time sequence information, the motion state data of each target in the history of several frames is jointly coded. These historical data contain speed information and can be extended to contain acceleration, displacement direction and other dynamic characteristics. By comprehensively constructing the time sequence motion vector sequence, the system can capture more delicate motion pattern changes of the target in the time dimension, including acceleration, deceleration, turning and nonlinear motion trend, so that the time sequence vector sequence is not only a simple stack of speed changes, but also a dynamic feature set describing the evolution process of target motion. The time sequence motion vector sequence and the corresponding whole frame image data are jointly processed and input into the input layer of the first discriminator for data preprocessing, so that data of different sources and different dimensions are unified into an input format suitable for subsequent convolution operations. The system normalizes the time sequence motion vector to ensure that the features in each dimension are within the same numerical range, avoiding feature deviation caused by inconsistent scales while maintaining the relative trend of motion changes. At the same time, the whole frame image data is processed through size standardization and pixel value normalization, etc. to ensure that the image features and motion features can be effectively aligned at the input stage. In addition, in order to improve the joint expression ability of spatio-temporal features, the motion vector features and image data are spliced or channel fused in the feature dimension during preprocessing to construct a global motion input feature matrix. The global motion input feature matrix is input into a multi-layer convolutional neural network for feature extraction. In order to fully capture the motion patterns at different scales, different sizes of convolution kernels are set in the multi-layer convolutional neural network for spatial feature calculation. Smaller size convolution kernels focus on capturing local detail changes and can sensitively perceive small amplitude motion shifts and local disturbance information, while larger size convolution kernels are used to extract large range motion trend features such as group movement, collective acceleration, etc. global dynamic change. By configuring convolution kernels of different scales at different network levels, the system simultaneously models multi-scale information during feature extraction. After each convolution processing, the spatial dimension of the feature map gradually decreases, and the semantic information of the feature gradually increases layer by layer. The system extracts fine-grained dynamic changes in the shallow network and aggregates long-term motion trends in the deep network to form a multi-level global motion feature map. Based on the multi-level global motion feature map, global motion pattern recognition analysis is performed. The analysis process gradually stacks or connects the feature maps extracted at different levels through feature fusion operations to form a comprehensive feature set rich in different spatial scale information. By introducing a global average pooling operation, the spatial dimension of the feature map is further compressed, and the spatial information is mapped to the feature channel, so that each channel can represent a typical motion pattern response.In order to enhance the sensitivity of the model to different motion patterns, an attention mechanism is introduced after pooling to adjust the weights of different feature channels, dynamically highlighting the motion feature channels most discriminative to the current motion scene, while suppressing the feature channels with lower relevance to the current scene. Through adaptive feature reweighting, the system can automatically adjust the focus of the feature space according to the input time series motion vector and image features, thereby more accurately identifying the global motion trend of the target. Through a series of deep convolution, feature fusion and attention enhancement operations, the global motion trend intermediate feature is obtained.
[0037] In a specific embodiment, the process of performing step 103 can specifically include the following steps: Based on the global discriminative feature vector, the history scale change trend is analyzed to obtain a scale adaptive coefficient, and the local discriminative feature vector is combined to calculate a motion speed influence factor to obtain a motion speed decay factor; The scale adaptive coefficient and the motion speed decay factor are input into an exponential function calculation of a scale compensation factor dynamic updating equation to obtain an updated value of the scale compensation factor; Based on the updated value of the scale compensation factor and the current target scale data, scale compensation prediction is performed to obtain a predicted scale compensation value; The predicted scale compensation value is subjected to compensation boundary constraint detection, and is smoothed by a hyperbolic tangent function to obtain an adaptive scale compensation parameter.
[0038] Specifically, the global discriminative feature vector is used to analyze the historical scale change trend. According to the target motion trend described by the global discriminative feature vector, combined with the scale change trajectory of the target in continuous multiple frames, the long-term stability and fluctuation characteristics of the scale change are analyzed. By statistically analyzing the mean, variance and change rate of the target scale change in the historical frames, the system can depict the overall trend of the scale change. If the scale change is stable, it means that the relative distance and posture change between the target and the camera are small, and the system sets a higher stability weight accordingly; if the scale fluctuation is severe, the system automatically reduces the stability expectation and appropriately relaxes the scale compensation range. Based on the historical scale change characteristics, a scale adaptive coefficient is generated, which can dynamically reflect the stability of the current target scale change. At the same time, combined with the local discriminative feature vector, the motion speed influence factor is calculated. The local discriminative feature vector mainly describes the change of the target in the local scale dimension, which is related to the target motion speed. By analyzing the scale change frequency, change amplitude and change direction in the local feature vector, the system derives the actual influence degree of the target motion speed on the scale change. The system defines a motion speed attenuation factor according to the sudden change of scale accompanied by fast motion. The motion speed attenuation factor reflects the influence degree of target speed change on the stability of scale estimation. The faster the speed, the larger the attenuation factor, and the larger the dynamic adjustment amplitude in the compensation process, so as to avoid the lag or oscillation of scale estimation caused by high-speed motion. The scale adaptive coefficient and the motion speed attenuation factor are input into the scale compensation factor dynamic updating equation for exponential function calculation. The dynamic updating equation based on the exponential function can nonlinearly adjust the updating amplitude of the scale compensation factor according to the stability of the scale change and the severity of the motion speed. The introduction of the exponential function makes the compensation factor updating tend to be linear in small amplitude change, ensuring the continuity and controllability of the scale change; while in large amplitude change, the updating speed is accelerated, quickly responding to the severe change of target motion, thereby improving the adaptability of the compensation mechanism in high-speed dynamic environment. Through this nonlinear adjustment, the system generates the scale compensation factor update value, which dynamically reflects the scale change trend and motion state of the target in the current frame and historical sequence. Based on the scale compensation factor update value, combined with the actual scale data of the target in the current frame, the scale compensation prediction is executed. In the prediction process, the system takes the scale data detected by the current target as the reference, adds the compensation adjustment amount calculated according to the update value to form the predicted scale compensation value. In order to enhance the stability and robustness of scale compensation, the predicted scale compensation value is subjected to compensation boundary constraint detection. A scale change tolerance threshold is set, when the predicted scale compensation value exceeds the threshold range, the system triggers the smoothing processing mechanism to prevent the system from being unstable due to the excessive scale update amplitude. In the smoothing process, the hyperbolic tangent function is used to adjust the compensation value.The hyperbolic tangent function has natural amplitude compression characteristics and good smoothness, which can effectively suppress abnormal change values while maintaining the trend of scale change, prevent severe fluctuations or abnormal jumps in the scale compensation process. The adaptive scale compensation parameter is generated.
[0039] In a specific embodiment, the process of inputting the scale adaptive coefficient and the motion speed attenuation factor into the scale compensation factor dynamic update equation for exponential function calculation to obtain the scale compensation factor update value can specifically include the following steps: Based on the current frame target scale state data, the historical scale compensation factor is extracted to obtain the current scale compensation factor reference value; Input the scale adaptive coefficient and the motion speed attenuation factor into the scale compensation factor dynamic update equation for parameter configuration to obtain the configured dynamic update equation; The product of the motion speed attenuation factor and the multi-target speed change amount in the configured dynamic update equation is calculated by exponential attenuation to obtain the exponential attenuation calculation result; The current scale compensation factor reference value is multiplied by the scale adaptive coefficient and the exponential attenuation calculation result to obtain the scale compensation factor update value.
[0040] Specifically, based on the target scale state data of the current frame, the history scale compensation factor is extracted to obtain the reference value of the current scale compensation factor. The system extracts the scale compensation history data of the target at each time node from the scale change record of the continuous frames, and through statistical analysis of these history scale compensation factors, the mean, variance and trend of the scale change are comprehensively considered to derive the scale compensation reference value corresponding to the current frame. The scale adaptive coefficient and the motion speed attenuation factor are input into the scale compensation factor dynamic updating equation for parameter configuration, and according to the current motion state and scale change characteristics of the target, the internal parameters of the updating equation are dynamically adjusted to make the updating mechanism adaptively respond to the motion intensity and scale change amplitude in different scenes. The scale adaptive coefficient is used to reflect the stability of the scale change, and the more stable the scale change is, the larger the coefficient is, indicating that the compensation update needs to maintain strong smoothness; while the motion speed attenuation factor measures the influence of the target motion speed on the scale change, and the faster the speed is, the larger the attenuation factor is, meaning that the compensation update needs to be more sensitive to quickly respond to the motion change. By taking these two coefficients as the core adjustment parameters, the dynamic updating equation can complete personalized parameter configuration according to the real-time state of the target at each frame, forming a highly adaptive scale compensation updating model. The parameterized motion speed attenuation factor is multiplied by the multi-target speed change amount, and an exponential decay calculation is performed based on the product result. The introduction of the exponential decay calculation is to realize the nonlinear control of the speed change influence in the compensation updating process, ensuring that the updating process has both sensitivity of fast response and inhibition of excessive adjustment caused by dramatic speed fluctuations. Specifically, the exponential decay function can keep a gentle change when the speed change is small, so that the adjustment range of the compensation factor is limited, avoiding large-scale fluctuations caused by small-scale motion; while when the speed change is significant, the exponential decay mechanism can quickly increase the response intensity to improve the adjustment speed of the compensation factor, so as to timely correct the scale error caused by fast motion. Through the exponential decay operation, the system generates an attenuation calculation result reflecting the current motion dynamic influence of the target. Finally, the current scale compensation factor reference value is multiplied by the scale adaptive coefficient and the exponential decay calculation result obtained above to generate the scale compensation factor update value. By taking the reference value as the basis for compensation update, combining the scale adaptive coefficient to guide the stability requirement of scale change, and introducing the motion dynamic adjustment factor from the exponential decay result, a scale compensation factor update value is formed, which takes into account both the history trend and the current dynamics.
[0041] In a specific embodiment, the process of performing step 104 can specifically include the following steps: The adaptive scale compensation parameters are subtracted from the current scale data of each target to obtain the scale reliability weight coefficient of each target; The feature pyramid multi-level extraction is performed on each target region in a continuous frame image sequence to obtain a multi-scale feature vector of each target, and a corresponding de-smearing gain coefficient is calculated based on a detection confidence of each target through a Sigmoid activation function; The scale credibility weight coefficient, the multi-scale feature vector and the de-smearing gain coefficient are subjected to element-by-element multiplication operation, and the product results of all targets are subjected to summation accumulation processing to obtain a multi-target scale fusion feature vector; The multi-target scale fusion feature vector and a multi-target speed change amount are spliced and combined, and input into a generator of a Wasserstein generative adversarial network for multi-scale template analysis to generate a multi-scale de-smearing prediction template.
[0042] Specifically, the scale confidence weight coefficient is calculated using the adaptive scale compensation parameter and the target current scale data. For each target, the actual scale data detected in the current frame is extracted, and the scale data is difference calculated with the adaptive scale compensation parameter. By calculating the absolute deviation of the current scale and the compensation scale, the system can quantify the degree of agreement between the current scale prediction and the actual detection. The smaller the difference, the higher the matching degree between the current detection scale and the compensation parameter, and the more stable the target state, so the scale confidence should be higher. Conversely, if the difference is large, it indicates that the target scale changes dramatically or the detection is biased, and the confidence should be reduced accordingly. In order to standardize this difference, it is mapped to a reasonable weight range through normalization operation to form the scale confidence weight coefficient corresponding to each target. The feature pyramid multi-level extraction is performed on each target region in the sequence of consecutive frame images. The feature pyramid network structure is used for target detection and recognition tasks, which can extract deep semantic features and shallow detail features of images at different spatial scales. For each target region, the system performs multiple convolution feature extraction at different scale resolutions to obtain multi-level feature vectors, which capture local texture information, shape contour features, and global spatial structure information of the target. At the same time, for the detection confidence of each target, the Sigmoid activation function is introduced for normalization processing to obtain the de-smearing gain coefficient. The introduction of the Sigmoid function effectively maps the original confidence to a smooth change interval, avoiding the influence of extreme values on subsequent fusion calculation. The higher the confidence of the target, the larger the gain coefficient after Sigmoid mapping, indicating that the target has a greater weight in feature fusion. The gain coefficient of the target with low confidence is appropriately compressed, reducing its influence on the fusion result, thereby effectively avoiding the introduction of too much noise by low-quality detection results. The scale confidence weight coefficient, multi-scale feature vector, and de-smearing gain coefficient are multiplied element by element. The scale confidence weight coefficient and the multi-scale feature vector of each target are multiplied element by element to enhance the feature expression of high-confidence targets and suppress the feature contribution of low-confidence targets. The product is multiplied element by element with the de-smearing gain coefficient to adjust the importance of each target feature according to the detection confidence. Through two consecutive element-by-element multiplication operations, the joint modeling of target scale stability and detection reliability is realized, and the contribution of each target to the final fusion feature is dynamically adjusted in the feature space. The sum of the product results of all targets is accumulated to obtain the multi-target scale fusion feature vector. The multi-target scale fusion feature vector and the multi-target velocity change are spliced and combined. Through the feature splicing operation, the target motion state information is introduced based on the fusion feature vector, so that the final input feature contains both static spatial features and dynamic temporal features. After splicing, the composite feature vector is input into the generator module based on the Wasserstein generative adversarial network.Compared with the traditional generative adversarial network, the Wasserstein generative adversarial network has better training stability and stronger generation quality control ability, and is suitable for processing high-dimensional complex feature generation problems. The generator receives the spliced features as input, and performs feature transformation and generation through a multi-layer neural network structure, and finally outputs a multi-scale de-smearing prediction template. In the generation process, the network not only learns the spatial distribution law of the target at different scales, but also can predict the optimal scale compensation value of the target at the next moment according to the motion change trend. The generated multi-scale de-smearing prediction template has diversified scale representation ability and can provide more adaptive tracking template for targets with different motion states and different scale changes.
[0043] In a specific embodiment, the process of performing step 105 can specifically include the following steps: Performing dense optical flow field calculation and motion smearing area identification on the sequence of continuous frame images to generate de-smearing mask data; Performing template matching processing on the multi-scale de-smearing prediction template and the de-smearing mask data, calculating the best matching position of each target in the current frame through a normalized cross-correlation function, and updating the target state information to obtain a first multi-target tracking result; Generating scale-consistent features based on the first multi-target tracking result, and performing global optimal identity association matching combined with the Hungarian algorithm to obtain multi-target tracking data after association matching; Performing target life cycle management and trajectory continuity optimization on the multi-target tracking data after association matching to generate a second multi-target tracking result.
[0044] Specifically, the dense optical flow field calculation and the motion blur area identification are performed based on the continuous frame image sequence. The high-precision dense optical flow estimation algorithm is used to calculate the displacement vector of each pixel in the continuous two frames of images to generate the optical flow field data. The optical flow field can reflect the pixel-level displacement change and reveal the dynamic characteristics such as motion direction and speed. After completing the optical flow calculation, the amplitude and direction consistency of the optical flow vector are analyzed to identify the motion blur area. Since the blur area is characterized by abnormal increase in the amplitude of the optical flow and irregular change in the direction, the system sets reasonable amplitude threshold and direction consistency standard, combines connected component analysis and morphological processing, and extracts continuous and clear motion blur area to generate preliminary de-blurring mask data. In order to improve the accuracy of the mask, the temporal smoothing and spatial regularization mechanisms are introduced to ensure the good spatio-temporal consistency of the mask between consecutive frames, providing accurate blur filtering area for subsequent template matching. The multi-scale de-blurring prediction template is matched with the de-blurring mask data generated above. The multi-scale prediction template has rich scale transformation and motion compensation ability through the pre-training of the generative adversarial network, which can adapt to the scale change and pose adjustment of the target under different motion states. During the template matching process, the matching range is limited by the de-blurring mask, and the template search is only performed in the effective area to avoid blur interference. The specific matching method uses the normalized cross-correlation function as the similarity measure standard, and the matching score of each position is obtained by sliding calculation of the prediction template and the candidate region of the current frame image. The normalized cross-correlation function can effectively eliminate the influence of local brightness change on similarity calculation, improving the robustness and accuracy of the matching. For each target, the position with the highest score is selected as the best matching position in the current frame, and the center coordinates, scale information and matching confidence of the target are updated based on the position, thereby forming the first-stage multi-target tracking result. The scale consistency feature is generated based on the first multi-target tracking result. The scale consistency feature quantifies the scale stability of the target in the time dimension by analyzing the stability and continuity of the scale change of the target between consecutive frames. The system differentiates the scale of each target in the current frame and the previous frames, and calculates the standard deviation and mean value of the scale change through the sliding window to obtain the scale consistency score. The more stable the scale change is, the higher the score is, reflecting the stable physical size and spatial structure of the target in the tracking process. The identity association feature vector is formed by combining the scale consistency feature with the appearance feature and the motion feature. The global optimal identity association matching is performed based on the Hungarian algorithm. The system constructs the identity association cost matrix according to the feature similarity between each tracking target and the historical trajectory. The smaller the matrix element value is, the higher the rationality of the target and trajectory matching is. The cost matrix considers the appearance similarity, motion trajectory consistency and scale consistency to ensure the reasonable fusion of multi-feature information in the matching process.The cost matrix is input into the Hungarian algorithm for solving to find a matching scheme with minimum cost and realize global optimal identity association. In this way, the system can effectively avoid identity drift or incorrect association caused by local optimal matching, especially in high dynamic environments with dense targets, frequent interaction or complex occlusion, greatly improving the accuracy of identity maintenance and the stability of tracking. After identity association, the multi-target tracking data after association matching is processed for target life cycle management and trajectory continuity optimization. The life cycle management module monitors the appearance and disappearance state of each target. When a new target is detected and stable tracking is maintained for a certain number of frames, the system assigns a new identity number; when an existing target is not detected for a certain number of frames, the system marks it as a disappearance state to avoid identity confusion caused by short-term loss. In addition, the system performs continuity optimization processing on the tracking trajectory based on Kalman filtering or a long-short trajectory prediction model, predicts and smooth interpolates the future position of the target, repairs the trajectory interruption caused by short-term occlusion or detection failure, and ensures the continuity and physical reasonableness of the trajectory. The second-stage multi-target tracking result is generated, including accurate target position, size, confidence information and complete continuous motion trajectory.
[0045] In a specific embodiment, the process of generating scale consistency features based on the first multi-target tracking result and performing global optimal identity association matching combined with the Hungarian algorithm to obtain the multi-target tracking data after association matching can specifically include the following steps: Extract the appearance feature vector and motion feature vector of each target from the first multi-target tracking result, and generate scale consistency features based on the scale data difference of each target; Weighted fusion processing of the appearance feature vector, motion feature vector and scale consistency feature is performed to obtain a target identity association probability distribution matrix; Construct a cost matrix of the Hungarian algorithm based on the target identity association probability distribution matrix and the motion consistency measurement data; Input the cost matrix into the Hungarian algorithm for global optimal allocation calculation to obtain a minimum cost matching scheme, and generate multi-target tracking data after association matching according to the minimum cost matching scheme.
[0046] Specifically, the appearance feature vector and motion feature vector of each target are extracted from the first multi-target tracking result, and the scale consistency feature is generated based on the scale data difference of each target. For each target region, a deep convolutional neural network is used to extract a high-dimensional appearance feature vector from the image. These appearance features can describe the static attributes of the target, such as color distribution, texture structure, and shape contour, ensuring good discrimination ability in complex scenes where similar targets coexist. At the same time, the system constructs a motion feature vector based on the center position change, velocity vector, and acceleration change of the target between consecutive frames. The motion feature effectively describes the dynamic behavior pattern of the target in the time dimension, including the direction of travel, speed, and trajectory change law. In addition, the system extracts the scale data between the current frame and the historical frames of the target, calculates the difference of the scale value, and generates the scale consistency feature by statistical analysis of the stability of the scale change. The scale consistency feature reflects the stability of the scale fluctuation of the target in the continuous tracking process. The smaller the scale change, the more stable the actual physical size of the target, and the higher the reliability of the identity association. The appearance feature vector, motion feature vector, and scale consistency feature are weighted and fused to generate a target identity association probability distribution matrix. The weighting and fusion process aims to dynamically adjust the weight of each feature according to its importance in different scenarios. In scenarios where the target appearance discrimination is high, the system gives more weight to the appearance feature. In environments where the motion characteristics are more prominent or the target appearance similarity is high, the weight of the motion feature is appropriately increased. At the same time, the scale consistency feature serves as a stability supplement, providing additional scale constraints in various complex situations and further improving the reliability of identity association. The weighted fusion forms a unified comprehensive feature vector by linearly weighting and summing each feature vector, and calculates the similarity score between targets based on this vector. The higher the score, the stronger the consistency of the two targets in terms of appearance, motion trajectory, and scale change. The system constructs a target identity association probability distribution matrix based on the similarity scores between each target. Each element in the matrix represents the confidence probability of matching between the current target and the historical trajectory. The higher the probability, the more reliable the association. Based on the target identity association probability distribution matrix and the motion consistency measurement data, a cost matrix for the Hungarian algorithm is constructed. The motion consistency measurement data measures the similarity of the motion patterns between targets by analyzing the motion trend, speed change, and trajectory smoothness of the target in consecutive time periods. Specifically, for each target and historical trajectory, the velocity vector angle, displacement direction consistency, and acceleration change trend are calculated to form a quantitative motion consistency score. By jointly modeling the identity association probability and the motion consistency measurement, the system comprehensively evaluates the matching rationality between target-trail pairs.To this end, each element in the cost matrix is defined as the inverse of the target identity association probability plus the weighted sum of the motion consistency score, the higher the probability and the better the motion consistency, the smaller the cost, thereby ensuring that the final matching scheme can maximize the identity continuity while taking into account the naturalness and continuity of the motion behavior. After completing the construction of the cost matrix, it is input into the Hungarian algorithm for global optimal allocation calculation. As a classic optimization matching algorithm, the Hungarian algorithm can efficiently find the matching scheme with the minimum total cost in the multi-target multi-track association task, ensuring the optimal overall matching quality and avoiding the identity drift or mismatch problem caused by local optimization. The algorithm models the association problem as a minimum weight matching problem, sequentially finds the optimal matching pair, and gradually approaches the optimal solution through row and column reduction, zero element covering and adjustment assignment, etc. Finally, a one-to-one matching relationship between each current target and historical track is obtained. Through the Hungarian algorithm, the system can ensure high-precision and high-stability identity tracking in a complex environment with dynamic changes of multiple targets, frequent occlusions and intensive interactions. Finally, based on the minimum cost matching scheme, the identity information of each target is updated according to the matching result to generate multi-target tracking data after association and matching. Each target not only retains the spatial position, scale data and detection confidence information in the preliminary tracking stage, but also adds an identity tag based on the matching result, ensuring the consistency of the target and the coherence of the track in the continuous tracking process. At the same time, the system combines the life cycle management strategy to assign a new identity to a newly appeared target, release the identity and terminate the track of a target that is out of service for more than a predetermined threshold, ensuring the integrity and rationality of the tracking data. Through the above steps, the complete identity association result of multi-target no-dragging tracking in a high dynamic scene is finally generated.
[0047] The above describes the multi-target no-dragging tracking method in a high dynamic scene in the embodiment of the application. The following describes a multi-target no-dragging tracking device in a high dynamic scene in the embodiment of the application. Please refer to Figure 2 An embodiment of the multi-target no-dragging tracking device in a high dynamic scene in the embodiment of the application includes: The acquisition module 201 is configured to acquire a continuous frame image sequence of a camera in a high dynamic scene, and extract motion parameters from the continuous frame image sequence to obtain a multi-target speed change quantity and a bounding box scale change sequence. The feature discrimination module 202 is configured to input the multi-target speed change quantity and the bounding box scale change sequence into a double discriminator network for global and local feature discrimination to obtain a global discrimination feature vector and a local discrimination feature vector. The calculation module 203 is configured to calculate an adaptive scale compensation parameter according to the global discrimination feature vector and the local discrimination feature vector. The multi-source scale fusion module 204 is configured to perform multi-source scale fusion based on the adaptive scale compensation parameter to obtain a multi-scale de-smearing prediction template. The multi-target tracking module 205 is configured to perform optical flow estimation and smearing elimination based on the multi-scale de-smearing prediction template to obtain a first multi-target tracking result, perform scale consistency identity association on the first multi-target tracking result, and output a second multi-target tracking result.
[0048] Through the cooperation of the above components, by constructing an adaptive scale compensation parameter calculation method based on a double discriminator feature, the compensation strategy can be dynamically adjusted according to the target motion state and scale change history, overcoming the limitations of the traditional fixed compensation coefficient method, and effectively solving the tracking failure problem caused by the rapid scale jump of the target in the high dynamic scene. The double discrimination mechanism of the first discriminator processing global motion trend and the second discriminator processing local scale change can capture global and local motion-scale correlation information at the same time, and has stronger feature expression ability and discrimination accuracy than the single discriminator structure. Through the three-element weighted fusion processing of the scale reliability weight coefficient, the multi-scale feature vector and the de-smearing gain coefficient, the weight can be dynamically allocated according to the scale change reliability of each target, effectively processing the scale interaction between multiple targets, and avoiding the feature conflict and information loss problem that may be caused by the simple feature splicing method. The motion smearing elimination algorithm based on dense optical flow field calculation and spatio-temporal consistency constraint can accurately identify and eliminate the smearing interference caused by high-speed motion, realize non-smearing tracking through accurate matching of the de-smearing mask data and the prediction template, and significantly improve the tracking performance of the traditional method under the motion blur condition. Taking the scale consistency feature as an important basis for identity association, combined with the global optimal allocation strategy of the Hungarian algorithm, the identity switching and occlusion problem in the high dynamic scene can be effectively solved, and the method has stronger robustness and accuracy than the traditional association method which only relies on the appearance feature. The mechanism of adjusting the camera exposure time and frame rate parameters in real time according to the target motion speed can optimize the imaging quality at the hardware level, and improve the tracking conditions in the high dynamic scene from the source.
[0049] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, the system and the unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0050] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the entire or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0051] The above description and the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features. These modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for multi-target motion-free tracking in high dynamic scenes, characterized in that, include: Acquire a sequence of consecutive frame images from a camera in a high dynamic scene, and extract motion parameters from the sequence of consecutive frame images to obtain a sequence of multi-target velocity changes and bounding box scale changes; The multi-target velocity change and the bounding box scale change sequence are input into a dual discriminator network for global and local feature discrimination to obtain global and local discriminant feature vectors. Calculate the adaptive scaling compensation parameters based on the global discriminant feature vector and the local discriminant feature vector; Multi-source scale fusion is performed based on the adaptive scale compensation parameters to obtain a multi-scale ghosting prediction template. Based on the multi-scale ghosting prediction template, optical flow estimation and ghosting elimination are performed to obtain a first multi-target tracking result. Scale-consistent identity association is then performed on the first multi-target tracking result to output a second multi-target tracking result.
2. The multi-target trail-free tracking method in high dynamic scenes according to claim 1, characterized in that, The process of acquiring a continuous frame image sequence from a camera in a high dynamic scene, and extracting motion parameters from the continuous frame image sequence to obtain a sequence of multi-target velocity changes and bounding box scale changes includes: The camera's exposure time and frame rate parameters are dynamically adjusted based on the detected target motion speed to obtain a continuous frame image sequence in a high dynamic scene. Target detection and position tracking are performed on the continuous frame image sequence to obtain a multi-target center position coordinate sequence and a bounding box coordinate sequence; The difference in target center position coordinates between adjacent frames is calculated based on the multi-target center position coordinate sequence to obtain the multi-target velocity change, and the bounding box scale change sequence is calculated based on the bounding box coordinate sequence.
3. The multi-target motion-free tracking method in high dynamic scenes according to claim 1, characterized in that, The step of inputting the multi-target velocity change and the bounding box scale change sequence into a dual discriminator network for global and local feature discrimination to obtain global and local discriminant feature vectors includes: The multi-target velocity change is input into the first discriminator of the dual discriminator network to extract global motion pattern features, thereby obtaining intermediate features of global motion trend. The bounding box scale change sequence is input into the second discriminator of the dual discriminator network to extract local scale change features, thereby obtaining intermediate features of local scale change. The intermediate features of the global motion trend are processed by four layers of downsampling convolution to obtain the global deep feature mapping; The intermediate features of the local scale changes are fused by skip connections, and feature reconstruction calculation is performed through the decoder layer to obtain the local reconstructed feature map; Scale-aware attention mechanism analysis is performed on the global deep feature map and the local reconstructed feature map respectively to obtain the global discriminative feature vector and the local discriminative feature vector.
4. The multi-target trail-free tracking method in high dynamic scenes according to claim 3, characterized in that, The process of inputting the multi-target velocity changes into the first discriminator of the dual discriminator network for global motion pattern feature extraction yields intermediate features of the global motion trend, including: The velocity changes of the multiple targets are arranged in a temporal sequence, and a temporal motion vector sequence is constructed by combining it with historical frame motion data; The temporal motion vector sequence and the corresponding whole-frame image data are input into the input layer of the first discriminator for data preprocessing to obtain the global motion input feature matrix. The global motion input feature matrix is input into a multi-layer convolutional neural network. Each layer of the multi-layer convolutional neural network uses convolutional kernels of different sizes to perform spatial feature calculations, resulting in a multi-level global motion feature map. Based on the multi-level global motion feature map, global motion pattern recognition analysis is performed to obtain intermediate features of global motion trends.
5. The multi-target trail-free tracking method in high dynamic scenes according to claim 1, characterized in that, The step of calculating the adaptive scaling compensation parameters based on the global discriminative feature vector and the local discriminative feature vector includes: Based on the global discriminative feature vector, historical scale change trend analysis is performed to obtain the scale adaptation coefficient, and combined with the local discriminative feature vector, motion speed influence factor is calculated to obtain motion speed attenuation factor; The scale adaptation coefficient and the motion velocity attenuation factor are input into the scale compensation factor dynamic update equation and calculated using an exponential function to obtain the updated value of the scale compensation factor. Based on the updated value of the scale compensation factor and the current target scale data, scale compensation prediction is performed to obtain the predicted scale compensation value. The predicted scale compensation value is subjected to compensation boundary constraint detection and smoothed using a hyperbolic tangent function to obtain adaptive scale compensation parameters.
6. The multi-target trail-free tracking method in high dynamic scenes according to claim 5, characterized in that, The step of inputting the scale adaptation coefficient and the motion velocity attenuation factor into the scale compensation factor dynamic update equation for exponential function calculation to obtain the scale compensation factor update value includes: Historical scale compensation factors are extracted based on the target scale state data of the current frame to obtain the baseline value of the current scale compensation factor; The scale adaptation coefficient and the motion velocity attenuation factor are input into the scale compensation factor dynamic update equation for parameter configuration, and the configured dynamic update equation is obtained. The product of the motion velocity decay factor and the velocity change of the multiple targets in the configured dynamic update equation is subjected to exponential decay calculation to obtain the exponential decay calculation result. The current scale compensation factor baseline value is multiplied by the scale adaptation coefficient and the exponential decay calculation result to obtain the scale compensation factor update value.
7. The multi-target trail-free tracking method in high dynamic scenes according to claim 1, characterized in that, The process of performing multi-source scale fusion based on the adaptive scale compensation parameters to obtain a multi-scale ghosting prediction template includes: The difference between the adaptive scale compensation parameter and the current scale data of each target is calculated to obtain the scale confidence weight coefficient of each target. Multi-level feature pyramid extraction is performed on each target region in the continuous frame image sequence to obtain the multi-scale feature vector of each target, and the corresponding de-ghosting gain coefficient is calculated by using the Sigmoid activation function based on the detection confidence of each target. The scale confidence weight coefficient, the multi-scale feature vector, and the de-ghosting gain coefficient are multiplied element-wise, and the product results of all targets are summed to obtain the multi-target scale fusion feature vector. The multi-target scale fusion feature vector is concatenated and combined with the multi-target velocity change, and then input into the generator of the Wasserstein generative adversarial network for multi-scale template analysis to generate a multi-scale de-ghosting prediction template.
8. The multi-target trail-free tracking method in high dynamic scenes according to claim 1, characterized in that, The process of optical flow estimation and ghosting elimination based on the multi-scale ghosting prediction template yields a first multi-target tracking result. Scale-consistent identity association is then performed on the first multi-target tracking result to output a second multi-target tracking result, including: Dense optical flow field calculation and motion blur region identification are performed on the continuous frame image sequence to generate de-blurring mask data; The multi-scale de-fuzzing prediction template and the de-fuzzing mask data are subjected to template matching processing. The optimal matching position of each target in the current frame is calculated by the normalized cross-correlation function, and the target state information is updated to obtain the first multi-target tracking result. Based on the first multi-target tracking result, scale consistency features are generated, and the Hungarian algorithm is used to perform global optimal identity association matching to obtain multi-target tracking data after association matching. The multi-target tracking data after association and matching is subjected to target lifecycle management and trajectory continuity optimization to generate a second multi-target tracking result.
9. The multi-target trail-free tracking method in high dynamic scenes according to claim 8, characterized in that, The process of generating scale-consistency features based on the first multi-target tracking result and performing globally optimal identity association matching using the Hungarian algorithm to obtain multi-target tracking data after association matching includes: Extract the appearance feature vector and motion feature vector of each target from the first multi-target tracking result, and generate scale consistency features based on the difference in scale data of each target; The appearance feature vector, the motion feature vector, and the scale consistency feature are weighted and fused to obtain the target identity association probability distribution matrix. The cost matrix of the Hungarian algorithm is constructed based on the target identity association probability distribution matrix and motion consistency measurement data; The cost matrix is input into the Hungarian algorithm for global optimal allocation calculation to obtain the minimum cost matching scheme, and multi-target tracking data after correlation matching is generated based on the minimum cost matching scheme.
10. A multi-target motion-free tracking device for high dynamic scenes, characterized in that, For performing the multi-target motion-free tracking method in a high-dynamic scene as described in any one of claims 1-9, the multi-target motion-free tracking device in the high-dynamic scene comprises: The acquisition module is used to acquire a continuous frame image sequence of the camera in a high dynamic scene, and to extract motion parameters from the continuous frame image sequence to obtain the multi-target velocity change and bounding box scale change sequence. The feature discrimination module is used to input the multi-target velocity change and the bounding box scale change sequence into the dual discriminator network for global and local feature discrimination, and obtain global discrimination feature vector and local discrimination feature vector; The calculation module is used to calculate adaptive scale compensation parameters based on the global discriminative feature vector and the local discriminative feature vector; The multi-source scale fusion module is used to perform multi-source scale fusion based on the adaptive scale compensation parameters to obtain a multi-scale ghosting prediction template. The multi-target tracking module is used to perform optical flow estimation and ghosting elimination based on the multi-scale ghosting prediction template to obtain a first multi-target tracking result, and to perform scale-consistent identity association on the first multi-target tracking result to output a second multi-target tracking result.
Citation Information
Cited By
Space-time consistency data generation method for visual target tracking
CN121527140A
Community abnormal behavior identification method and device based on video analysis
CN121982817A