Deep fusion of multi-modal information for target tracking

CN118823060BActive Publication Date: 2026-09-22XIAN AERONAUTICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410793660.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-19
Publication Date
2026-09-22
Estimated Expiration
2044-06-19

AI Technical Summary

Technical Problem

[0004]本发明的目的是提供多模态信息的深度融合目标跟踪方法,解决了现有技术中的单模态系统,因目标外观和环境变化容易导致跟踪失败的问题

Benefits of technology

[0006]本发明的有益效果是,通过构建颜色模型并引入双线性插值HOG特征,建立高密度、高维度形状模型;采用深层次自适应高置信度权衡策略和量纲层级归一化,实现模型内外多置信度的有效融合,提高目标跟踪的准确性;同时,针对模型更新设计分级非线性学习率曲线,保持跟踪稳定性,并通过粒子滤波提升跟踪速度和鲁棒性,显著优化了目标跟踪的准确度和适应性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118823060B_ABST
    Figure CN118823060B_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal information deep fusion target tracking method, and steps are as follows: step 1, for a given target, a multi-modal model is constructed, including a color model and a shape model; step 2, a particle swarm is randomly initialized; step 3, particle filtering search is performed on a candidate target to obtain an optimal solution, and the process is as follows: 3.1) color confidence and shape response confidence are calculated; 3.2) multi-confidence fusion is performed on the color model and the shape model; 3.3) particles are resampled; step 4, the model is updated; step 5, the steps are recycled, for the next frequency frame, steps 3 and 4 are recycled until tracking processing of all frequency frames is completed. The application belongs to the technical field of visual target tracking, meets real-time requirements of a tracking algorithm under general conditions, and has significant advantages in accuracy and success rate indexes in common specific scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of visual target tracking technology and relates to a deep fusion target tracking method based on multimodal information. Background Technology

[0002] In single-target tracking technology, building accurate target models, effective data fusion, and timely model updates are crucial. Designing models that adapt to different scenario requirements and accurately measuring their fusion and update strategies are key factors in ensuring tracking accuracy. Traditional target tracking technologies mainly rely on single-modal data input, such as video image sequences, which are easily affected by lighting, occlusion, and background interference in variable environments, thus reducing tracking accuracy. Furthermore, single-modal systems often exhibit low adaptability and flexibility in dynamic environments, especially when the target moves rapidly or the scene changes abruptly.

[0003] Existing multimodal target tracking methods still have room for improvement in information fusion processing. How to effectively integrate information from different modalities, and how to design a dynamically adjustable fusion strategy to adapt to environmental changes, are key to improving the performance of multimodal tracking systems. Furthermore, the real-time performance and resource consumption of the fusion algorithm need further optimization to meet the efficiency and computational resource requirements of practical applications. Summary of the Invention

[0004] The purpose of this invention is to provide a deep fusion target tracking method based on multimodal information, which solves the problem that existing single-modal systems are prone to tracking failure due to changes in target appearance and environment.

[0005] The technical solution adopted in this invention is a deep fusion target tracking method based on multimodal information, implemented according to the following steps: Step 1: For a given target, construct a multimodal model, including a color model and a shape model; Step 2: Randomly initialize the particle swarm; Step 3: Perform particle filter search on the candidate targets to obtain the optimal solution. The process is as follows: 3.1) Calculate the color confidence score and the shape response confidence score; 3.2) Perform multi-confidence fusion on the color model and shape model; 3.3) Particle resampling; Step 4: Update the model; Step 5: Repeat steps 3 and 4 for the next frequency frame until all frequency frames have been tracked.

[0006] The beneficial effects of this invention are that by constructing a color model and introducing bilinear interpolation HOG features, a high-density, high-dimensional shape model is established; a deep-level adaptive high-confidence trade-off strategy and dimensional hierarchical normalization are adopted to achieve effective fusion of multiple confidence levels inside and outside the model, thereby improving the accuracy of target tracking; at the same time, a hierarchical nonlinear learning rate curve is designed for model updates to maintain tracking stability, and particle filtering is used to improve tracking speed and robustness, significantly optimizing the accuracy and adaptability of target tracking. Attached Figure Description

[0007] Figure 1 This is a flowchart illustrating the implementation of the method of the present invention; Figure 2 Modeling the multimodal initialization in the method of this invention; Figure 3 This refers to the moving particle model used in the method of this invention; Figure 4 This is a schematic diagram of model fusion used in the method of the present invention; Figure 5 This is a diagram illustrating the Lemming process for an occluded scene in Example 1; Figure 6 This is a graph showing the tracking accuracy of the algorithm in an occluded scene under Lemming, as described in Embodiment 1 of the present invention. Figure 7 This is a success rate graph of Lemming in an occlusion scene according to Embodiment 1 of the present invention; Figure 8 This is a diagram of the jogging-2 process in Example 2, which involves leaving the field of view. Figure 9 This is a graph showing the tracking accuracy of the algorithm under jogging-2 out of view in Embodiment 2 of the present invention; Figure 10 This is a success rate graph of jogging-2 out of the field of view in Embodiment 2 of the present invention; Figure 11 This is a graph showing the tracking accuracy of the algorithm under the out-of-plane rotation scene shaking method in Embodiment 3 of the present invention. Figure 12 This is a success rate chart for the out-of-plane rotation scene shaking in Embodiment 3 of the present invention. Detailed Implementation

[0008] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0009] The deep fusion target tracking method based on multimodal information of the present invention comprises the following steps: Figure 1 As shown, Figure 1The frequency frame starts from t=2 and includes five steps: establishing a color model and a shape model for a given target, randomly initializing the particle swarm, performing particle filtering search on candidate targets to obtain the optimal solution, updating the model, and repeating steps 3 and 4 in the next frequency frame.

[0010] The deep fusion target tracking method based on multimodal information of the present invention is implemented according to the following steps: Step 1: For a given target, construct a multimodal model, including a color model and a shape model, such as... Figure 2 As shown, 1.1) Establish a color model, First, set the target of Color channels converted The color channels are analyzed, and the probability of each color appearing in each color channel is calculated. A color histogram is then built based on this probability and used as the target color model. The color model expression is as follows: (1) In equation (1), Yes The normalized values ​​of the color histogram The color histogram represents the first color. The value of the dimension; 1.2) Establish a shape model, Training was performed using bilinear interpolation HOG features, and the variable relationships in the frequency domain are as follows: (2) In equation (2), Represents the relevant response, and As a feature candidate region, Indicates a filter. It is a Fourier transform. express The complex conjugate; Feature enhancement will generate some HOG feature samples. and the corresponding responses , among them Set Fourier transform , , Then we have: (3) This was obtained by solving an optimization problem: (4) Therefore, the optimal solution is obtained. The expression: (5) In equation (5), yes .

[0011] Step 2: Randomly initialize the particle swarm. The target position in the current frame is predicted by the particle filtering method. Based on the position in the previous frame, it is assumed that the current position is near the position in the previous frame, so as to determine the search range and the particle moves only once in the search. particles The structural form is defined as ,in Let be the coordinates of the top-left corner of the rectangular region represented by the particle. The length and width of the area, For particles in and velocity in the direction of motion For the region in and The rate of change of dimension in the direction; like Figure 3 As shown, after acquiring the target in the initial frame, several particle boxes are allocated around the target position as the initial motion model, for example, 50 particle boxes can be set; subsequently, these particle boxes will be moved according to equation (6), as shown in the following expression: (6) In equation (6), , , It is the first Sub-images in the candidate box of a frame top left corner coordinate, Coordinates and width / height scaling factor; , and These are random numbers with a mean of 0 and a standard deviation of 1, 0.5, and 0.001, respectively.

[0012] Step 3: Perform particle filter search on the candidate targets to obtain the optimal solution. 3.1) Calculate the color confidence score and the shape response confidence score. like Figure 4 This step introduces a multi-confidence balancing strategy for target tracking, as follows: 3.1.1) Using Bach distance It measures the color confidence between the candidate box color model and the target color model; 3.1.2) In the shape model, the internal multi-confidence balancing mechanism combines two confidence indices, PSR and APCE, and uses a weighted average to ensure the authenticity and stability of the overall model confidence. This mechanism dynamically adjusts the weights of the two indices according to changes in the environment and the target, thereby maintaining high accuracy and stable target tracking performance under conditions such as occlusion and rapid target movement.

[0013] 3.1.3) Label the PSR and APCE indicators as follows: and The similarity of the preselected boxes is represented by the PSR and APCE metrics to enhance the reliability of model fusion. The expressions are as follows: (7) (8) In equation (7), r m,max The maximum value of the response graph, This represents the mean value of the side lobes. The standard deviation of the sidelobe; in equation (8), r m,min The minimum value of the response graph. To respond to the width and height of the image, r m,w,h To respond to each pixel value in the image; 3.2) Perform multi-confidence fusion on the color model and shape model. This step introduces a deep trade-off fusion method, which calculates the shape confidence (PSR and APCE) and color confidence (Bach distance) of candidate targets, and performs primary fusion of different confidence levels of the shape model within this framework; then, based on the highest confidence level, the shape confidence is further fused. and color confidence A deep weighted fusion with multiple confidence levels is performed between the shape model and the color model to obtain the final confidence level. The overall fusion strategy is as follows: Figure 4 As shown; Let the similarity limits of PSR confidence and APCE confidence be denoted as follows: and The arithmetic mean of the normalized values ​​is defined as the overall confidence score of the shape of the i-th candidate particle. sm i The expression is as follows: (9) Before fusing different models, it is also necessary to calculate the normalized color confidence score: (10) In equation (10), The confidence level of the color histogram. This represents the similarity limit for color confidence. The normalized confidence level; This step also proposes a high-confidence-based trade-off strategy, which is a method to dynamically adjust the weights of the shape model and the color model. By calculating the different confidence levels of each model separately, and then fusing the shape and color confidence levels through a weighted average, the accuracy and stability of target tracking are improved. The calculation formula is as follows: (11) In equation (11), and The first i Normalized color confidence and combined shape confidence of each candidate particle. The final confidence level after merging the two is... To integrate weight parameters, As can be seen from equation (11), combining the results of both and using the confidence level The form is applied to the particle filter candidate target search process, through analysis and The numerical value reflects the quality of the current tracking, thus determining the fusion parameters. The value of the input; based on this, this step proposes a new fusion parameter selection strategy, which prioritizes the use of inputs with higher confidence and gives them higher weights during the fusion process, thereby ensuring that the target in the current frame is determined with high confidence. The specific selection strategy is determined according to the following discrimination conditions: 1) If and This indicates that both trackers are working well, and weights will be assigned based on the difference in confidence between them. ; 2) If and This indicates that the shape tracker is fluctuating, while the color tracker is performing well. Therefore, the results from the color tracker will be used. The value is 0; 3) If and This indicates that the color tracker is fluctuating, while the shape tracker is performing well. Therefore, the results from the shape tracker will be used. The value is 1; For the three situations mentioned above, in the experiment The possible values ​​are as follows: (12) 4) If and This indicates that both trackers are fluctuating simultaneously, the target is lost or occluded, and fusion is necessary. The value is determined by the larger of the two; in this case, in the experiment... The formula for determining the value of is as follows: (13) in, Set to 0.16; Set to 0.12; 3.3) Particle resampling, Select the current frame The highest confidence particle box The particle frame The coordinates are used as the target position in the current frame. And at this position, particles Resampling.

[0014] Step 4: Update the model. This step proposes a nonlinear hierarchical balancing update mechanism to update two models separately. By dynamically adjusting the learning rate and combining the confidence scores of both models, a nonlinear hierarchical learning rate curve function is defined to control the step size of the model during learning and optimization. Under this mechanism, the learning rate curve adjusts the intensity of the learning rate according to different stages of model update and sets an upper limit for fluctuation, as shown in the following expression: (14) Equation (14) utilizes the optimal color confidence of each candidate particle in the current frame. Combined confidence level with optimal shape To evaluate the color model and shape model To assess the stability, the update frequencies of these two models are further calculated. and Different strategies are adopted based on the level of confidence. When the confidence level is high, it indicates that the model is stable and the model update speed can be accelerated; when the confidence level fluctuates, it indicates that the model has uncertainty and needs to be updated cautiously; when the confidence level is low, it indicates that the model is unstable and the update speed should be slowed down to maintain the robustness of the system.

[0015] 4.1) Color model update strategy, The optimal color model for the current frame is denoted as The color model update expression is: (15) In equation (15), To update the coefficients; (16) In formula (15) The first color in the histogram representing the color model The probability of a color. Indicates the current goal The corresponding color histogram; when When the color model is good, the model update is accelerated through a non-linear acceleration method, so that the acceleration gradually increases. when This indicates that the color model fluctuates, and although the update process maintains non-linear acceleration, the acceleration will gradually decrease. 4.2) Shape model update strategy, target of current frame Image after preprocessing Then, shape modeling is performed according to equation (17). Update: (17) Similar to the analysis of color model updates described above, the shape model update coefficients... The range of values ​​for is as follows: (18) when When the time is right, it indicates that the shape model is good. At this time, the model update is accelerated by nonlinearity, so that the acceleration gradually increases. when When this occurs, it indicates that the shape model is fluctuating. At this point, the model updates nonlinearly at an accelerated pace, and the acceleration gradually decreases.

[0016] Step 5, repeat the steps. For the next frame, steps 3 and 4 are repeated until tracking processing for all frames is complete. In each frame, the confidence weights of the shape model and color model are dynamically adjusted to ensure tracking stability and accuracy under different environments.

[0017] Through verification using multiple embodiments, the method of this invention demonstrates excellent tracking performance in various complex scenarios, effectively addressing changes in target appearance and environmental interference. The final results show that this method not only has high tracking accuracy and success rate but also good generalization ability across different tracking scenarios. This invention has significant application value and broad application prospects in the field of visual target tracking technology.

[0018] Example 1 exist Figure 5 In the scenario shown, tracking processing under occlusion scene lemming is performed according to the steps described above in the method of the present invention. In this scenario, as... Figure 6 , Figure 7As shown, the accuracy of this method reaches 0.799, and the success rate reaches 0.695. When the center error threshold is greater than 20 and the overlap rate threshold is less than 0.4, the accuracy and success rate of this algorithm are significantly better than the reference algorithm. It is evident that the tracking process according to the method of this invention results in errors within the allowable range, fully meeting the technical requirements.

[0019] Example 2 exist Figure 8 In the scenario shown, tracking processing under jogging-2 is performed according to the steps described above in the method of the present invention. In this scenario, as... Figure 9 , Figure 10 As shown, the accuracy of this method reached 0.723, and the correctness reached 0.718. When the center error threshold is greater than 10 and the overlap rate threshold is less than 0.8, the accuracy and success rate of this algorithm are significantly better than the reference algorithm. It is evident that the tracking process according to the method of this invention results in errors within the allowable range, fully meeting the technical requirements.

[0020] Example 3 Tracking processing under out-of-plane rotating scene shaking is performed according to the above steps of the method of the present invention. For example... Figure 11 , Figure 12 As shown, the accuracy of this method reaches 0.790, and the correctness reaches 0.700. When the center error threshold is greater than 20 and the overlap rate threshold is less than 0.4, the accuracy and success rate of this algorithm are significantly better than the reference algorithm. It is evident that the tracking process according to the method of this invention results in errors within the allowable range, fully meeting the technical requirements.

Claims

1. A deep fusion target tracking method based on multimodal information, characterized in that, Follow these steps: Step 1: For a given target, construct a multimodal model, including a color model and a shape model; Step 2: Randomly initialize the particle swarm; Step 3: Perform particle filter search on the candidate targets to obtain the optimal solution. The process is as follows: 3.1) Calculate the color confidence score and the shape response confidence score; 3.2) Perform multi-confidence fusion on the color model and shape model, specifically: A deep weighted fusion method is introduced, which calculates the shape confidence and color confidence of candidate targets and performs primary fusion of different confidence levels of the shape model within this framework; then, based on the highest confidence level, the shape confidence and color confidence are fused in a deep weighted fusion between the shape model and the color model to obtain the final confidence level. Let the similarity limits of PSR confidence and APCE confidence be denoted as follows: and The arithmetic mean of the normalized values ​​is defined as the overall confidence score of the shape of the i-th candidate particle. sm i The expression is as follows: (9) Before fusing different models, it is also necessary to calculate the normalized color confidence score: (10) In equation (10), The confidence level of the color histogram. This represents the similarity limit for color confidence. The normalized confidence level; This step also employs a high-confidence trade-off strategy, calculating the confidence levels of both shapes and colors separately, and then fusing the shape and color confidence levels by weighted averaging to improve the accuracy and stability of target tracking. The calculation formula is as follows: (11) In equation (11), and The first i Normalized color confidence and combined shape confidence of each candidate particle. The final confidence level after merging the two is... τ For fusion weight parameters; Based on this, this step proposes a new fusion parameter selection strategy, the specific selection strategy being determined according to the following criteria: 1) If and This indicates that both trackers are working well, and weights will be assigned based on the difference in confidence between them. ; 2) If and This indicates that the shape tracker is fluctuating, while the color tracker is performing well. Therefore, the results from the color tracker will be used. The value is 0; 3) If and This indicates that the color tracker is fluctuating, while the shape tracker is performing well. Therefore, the results from the shape tracker will be used. The value is 1; For the three situations mentioned above, in the experiment The possible values ​​are as follows: (12) 4) If and This indicates that both trackers are fluctuating simultaneously, the target is lost or occluded, and fusion is necessary. The value is determined by the larger of the two; in this case, in the experiment... The formula for determining the value of is as follows: (13) in, Set to 0.16; Set to 0.12; 3.3) Particle resampling; Step 4: Update the model; Step 5: Repeat steps 3 and 4 for the next frequency frame until all frequency frames have been tracked.

2. The deep fusion target tracking method based on multimodal information according to claim 1, characterized in that, In step 1, the specific process is as follows: 1.1) Establish a color model, First, set the target of Color channels converted The color channels are analyzed, and the probability of each color appearing in each color channel is calculated. A color histogram is then built based on this probability and used as the target color model. The color model expression is as follows: (1) In equation (1), Yes The normalized values ​​of the color histogram The color histogram represents the first color. The value of the dimension; 1.2) Establish a shape model, Training was performed using bilinear interpolation HOG features, and the variable relationships in the frequency domain are as follows: (2) In equation (2), Represents the relevant response. As a feature candidate region, Indicates a filter. It is a Fourier transform. express The complex conjugate; Feature enhancement will generate some HOG feature samples. and the corresponding responses , among them Set Fourier transform , , Then we have: (3) This was obtained by solving an optimization problem: (4) Therefore, the optimal solution is obtained. The expression: (5) In equation (5), yes .

3. The deep fusion target tracking method based on multimodal information according to claim 1, characterized in that, Step 2, the specific process is as follows: The target position in the current frame is predicted by the particle filtering method. Based on the position in the previous frame, it is assumed that the current position is near the position in the previous frame, so as to determine the search range and the particle moves only once in the search. particles The structural form is defined as ,in Let be the coordinates of the top-left corner of the rectangular region represented by the particle. The length and width of the area, For particles in and velocity in the direction of motion For the region in and The rate of change of dimension in the direction; After acquiring the target in the initial frame, several particle boxes are allocated around the target position as the initial motion model. These particle boxes will undergo position transfer according to equation (6), as shown in the following expression: (6) In equation (6), , , It is the first Sub-images in the candidate box of a frame top left corner coordinate, Coordinates and width / height scaling factor; , and These are random numbers with a mean of 0 and a standard deviation of 1, 0.5, and 0.001, respectively.

4. The deep fusion target tracking method based on multimodal information according to claim 1, characterized in that, In step 3.1), the specific process is as follows: A multi-confidence balancing strategy for target tracking is adopted, as follows: 3.1.1) Using Bach distance It measures the color confidence between the candidate box color model and the target color model; 3.1.2) In the shape model, the internal multi-confidence balancing mechanism combines the two confidence indices PSR and APCE and uses a weighted average to ensure the authenticity and stability of the overall confidence of the model. 3.1.3) Label the PSR and APCE indicators as follows: and The similarity of the preselected boxes is represented by the PSR and APCE metrics to enhance the reliability of model fusion. The expressions are as follows: (7) (8) In equation (7), r m,max The maximum value of the response graph, This represents the mean value of the side lobes. The standard deviation of the sidelobe; in equation (8), r m,min The minimum value of the response graph. To respond to the width and height of the image, r m,w,h This is in response to each pixel value in the image.

5. The deep fusion target tracking method based on multimodal information according to claim 1, characterized in that, In step 3.3), the specific process is as follows: Select the current frame The highest confidence particle box The particle frame The coordinates are used as the target position in the current frame. And at this position, particles Resampling.

6. The deep fusion target tracking method based on multimodal information according to claim 1, characterized in that, Step 4, the specific process is as follows: A non-linear hierarchical balancing update mechanism is adopted to update the two models separately. Under this mechanism, the learning rate curve adjusts the intensity of the learning rate according to different stages of model update, and an upper limit for fluctuation is set, as shown in the following expression: (14) Equation (14) utilizes the optimal color confidence of each candidate particle in the current frame. Combined confidence level with optimal shape To evaluate the color model and shape model To assess the stability, the update frequencies of these two models are further calculated. and Different strategies are adopted based on the level of confidence. When the confidence level is high, it indicates that the model is stable and the model update speed can be accelerated. When the confidence level fluctuates, it indicates that the model has uncertainty and needs to be updated cautiously. When the confidence level is low, it indicates that the model is unstable and the update speed should be slowed down to maintain the robustness of the system. 4.1) Color model update strategy, The optimal color model for the current frame is denoted as The color model update expression is: (15) In equation (15), To update the coefficients; (16) In equation (15), The first color in the histogram representing the color model The probability of a color. Indicates the current goal The corresponding color histogram; when When the color model is good, the model update is accelerated through a non-linear acceleration method, so that the acceleration gradually increases. when This indicates that the color model fluctuates, and although the update process maintains non-linear acceleration, the acceleration will gradually decrease. 4.2) Shape model update strategy, target of current frame Image after preprocessing Then, shape modeling is performed according to equation (17). Update: (17) Similar to the analysis of color model updates described above, the shape model update coefficients... The range of values ​​for is as follows: (18) when When the time is right, it indicates that the shape model is good. At this time, the model update is accelerated by nonlinearity, so that the acceleration gradually increases. when When this occurs, it indicates that the shape model is fluctuating. At this point, the model updates nonlinearly at an accelerated pace, and the acceleration gradually decreases.