Efficient robust target tracking method and application thereof in intelligent monitoring

By introducing deep learning, adaptive tracking technology, multimodal fusion and attention mechanisms into the target tracking method, and designing an end-to-end trainable tracking framework, the problem of insufficient accuracy and robustness of target tracking in the existing technology is solved, and real-time analysis and alarm functions in efficient and robust target tracking and intelligent monitoring applications are achieved.

CN119992282APending Publication Date: 2025-05-13HEPTAGON (SUZHOU) TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510063882.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing target tracking methods are inaccurate and robust in the face of target appearance changes, complex backgrounds and multiple sensor information, and are difficult to meet the needs of intelligent monitoring.

Method used

The initial model based on deep learning and adaptive tracking technology is adopted, combined with multimodal fusion tracking technology, an attention mechanism is introduced, and a fully end-to-end trainable tracking framework is designed to integrate steps such as feature extraction, object detection, tracking and model update.

Benefits of technology

It improves the accuracy and robustness of target tracking, enhances the adaptability and efficiency of the model, realizes continuous and accurate tracking of target objects, and has real-time analysis and alarm functions, enhancing the security and practicality of the intelligent monitoring system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992282A_ABST
    Figure CN119992282A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of target tracking, and provides a high-efficiency robust target tracking method and application thereof in intelligent monitoring, and the method comprises the steps: S1, building an initial model of a target based on the deep learning and adaptive tracking technology, and enabling the model to dynamically adapt to the change of the appearance of the target; s2, combining information from different sensors by using a multi-modal fusion tracking technology; s3, introducing an attention mechanism; according to the invention, by introducing deep learning and adaptive tracking technologies, the method can dynamically adapt to the change of the appearance of the target, and meanwhile, the multi-modal fusion tracking technology is combined with the information of different sensors, so that the tracking accuracy and robustness are improved; the attention mechanism is introduced, so that the model can process key information in the video in a centralized resource mode, and the tracking accuracy is further improved; in addition, all related steps are integrated by a completely end-to-end trainable tracking framework, and the adaptability and efficiency of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of target tracking, in particular to an efficient and robust target tracking method and its application in intelligent monitoring. Background Art

[0002] In the field of intelligent surveillance, target tracking is a key technology that involves accurately identifying and tracking target objects from video streams. However, existing target tracking methods face multiple challenges, such as changes in target appearance, interference from complex backgrounds, and diversity of sensor information. These problems lead to deficiencies in accuracy and robustness of existing methods, making it difficult to meet the needs of practical applications.

[0003] At present, traditional target tracking methods mainly rely on manually designed features and tracking algorithms, which often perform poorly when dealing with changes in target appearance and complex backgrounds; in addition, with the continuous development of sensor technology, a variety of sensors (such as RGB cameras, infrared cameras and radars) are widely used in intelligent monitoring systems.

[0004] However, existing methods can often only utilize the information of a single sensor and cannot fully utilize the complementarity of multimodal data to improve tracking performance.

[0005] To this end, those skilled in the art have proposed an efficient and robust target tracking method and its application in intelligent monitoring to solve the problems raised by the background technology. Summary of the invention

[0006] In order to solve the above technical problems, the present invention provides an efficient and robust target tracking method and its application in intelligent monitoring, so as to solve the problems that the prior art can only use the information of a single sensor and cannot fully utilize the complementarity of multimodal data to improve tracking performance.

[0007] An efficient and robust target tracking method, comprising:

[0008] S1. Based on deep learning and adaptive tracking technology, an initial model of the target is established, which can dynamically adapt to changes in the target's appearance;

[0009] S2. Use multimodal fusion tracking technology to combine information from different sensors to improve tracking accuracy and robustness;

[0010] S3. Introducing the attention mechanism enables the model to focus resources on processing key information in the video and improve tracking accuracy;

[0011] S4. Design a fully end-to-end trainable tracking framework that integrates all steps including feature extraction, target detection, tracking, and model updating to improve the adaptability and efficiency of the model.

[0012] S5. In intelligent monitoring applications, the above target tracking method is applied to real-time video stream processing to achieve continuous and accurate tracking of the target object.

[0013] Preferably, the adaptive tracking technology includes introducing a dynamic memory system that can store key information about target appearance changes during tracking and use this information to adjust the model, and the dynamic memory system includes a long short-term memory network (LSTM) model.

[0014] Preferably, the multimodal fusion tracking technology combines information from RGB cameras, infrared cameras and radars, and uses a deep learning model to learn how to effectively fuse features of different modalities by assigning a weight to each modality using a weighted summation approach.

[0015] Preferably, the attention mechanism focuses on the characteristics of the target itself and the interaction between the target and the background to improve the tracking performance in complex scenes. The attention mechanism introduces a model based on a soft attention mechanism to assign an attention weight to each position.

[0016] Preferably, the end-to-end trainable tracking framework uses deep reinforcement learning (including a cross-entropy loss function) to learn tracking strategies directly from tracking success rates.

[0017] Preferably, in the intelligent monitoring application, real-time analysis and alarm functions of the tracking results are also included. When the target object exhibits abnormal behavior or exceeds a preset range, the system can automatically alarm.

[0018] Compared with the prior art, the present invention has the following beneficial effects:

[0019] 1. The present invention introduces deep learning and adaptive tracking technology to establish an initial model that can dynamically adapt to changes in target appearance; this makes the present invention more adaptable and accurate when facing changes in target appearance.

[0020] 2. The present invention utilizes multimodal fusion tracking technology to combine information from different sensors, thereby improving tracking accuracy and robustness. By learning how to effectively fuse features of different modalities through a deep learning model, the present invention fully utilizes the complementarity of multimodal data and further improves tracking performance.

[0021] 3. The present invention introduces an attention mechanism, which enables the model to concentrate resources on processing key information in the video and improve the tracking accuracy; this mechanism focuses on the characteristics of the target itself and the interaction between the target and the background, thereby maintaining high accuracy in complex scenes.

[0022] 4. The present invention designs a fully end-to-end trainable tracking framework that integrates all steps such as feature extraction, target detection, tracking, and model updating; this improves the adaptability and efficiency of the model, making the present invention more practical and reliable in practical applications.

[0023] 5. In the intelligent monitoring application, the present invention realizes the real-time analysis and alarm function of the tracking results; when the target object exhibits abnormal behavior or exceeds the preset range, the system can automatically alarm, which further enhances the practicality and safety of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a flow chart of the efficient and robust target tracking method of the present invention. DETAILED DESCRIPTION

[0025] The following embodiments of the present invention are described in further detail in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0026] Embodiment: The present invention provides an efficient and robust target tracking method, such as Figure 1 As shown, including:

[0027] S1. Based on deep learning and adaptive tracking technology, an initial model of the target is established, which can dynamically adapt to changes in the target's appearance;

[0028] S2. Use multimodal fusion tracking technology to combine information from different sensors to improve tracking accuracy and robustness;

[0029] S3. Introducing the attention mechanism enables the model to focus resources on processing key information in the video and improve tracking accuracy;

[0030] S4. Design a fully end-to-end trainable tracking framework that integrates all steps including feature extraction, target detection, tracking, and model updating to improve the adaptability and efficiency of the model.

[0031] S5. In intelligent monitoring applications, the above target tracking method is applied to real-time video stream processing to achieve continuous and accurate tracking of the target object.

[0032] As can be seen from the above, by introducing deep learning and adaptive tracking technology, this method can dynamically adapt to changes in the target's appearance, and at the same time use multimodal fusion tracking technology to combine information from different sensors to improve the accuracy and robustness of tracking. The introduction of the attention mechanism enables the model to focus resources on processing key information in the video, further improving the accuracy of tracking. In addition, the fully end-to-end trainable tracking framework integrates all relevant steps, improving the adaptability and efficiency of the model. Finally, in real-time video stream processing, this method achieves continuous and accurate tracking of the target object, enhancing the practicality and security of the intelligent monitoring system.

[0033] Furthermore, the adaptive tracking technology includes introducing a dynamic memory system, which can store key information of target appearance changes during tracking and use this information to adjust the model. The dynamic memory system includes a long short-term memory network (LSTM) model. The formula of the long short-term memory network (LSTM) model includes:

[0034] h t =LSTM(x t ,h t-1 );

[0035] Among them, x t is the target feature of the current frame, h t is the hidden state (i.e. memory) of the current frame, h t-1 is the hidden state of the previous frame, and the formula describes how to update the memory of the current frame based on the current input and the memory of the previous frame.

[0036] As can be seen above, the introduction of a dynamic memory system including a long short-term memory network (LSTM) model has shown remarkable beneficial effects. The system can effectively store key information about the target's appearance changes during tracking and use this information to adjust the model in real time. The formula of the LSTM model accurately describes how to update the memory of the current frame based on the current input features and the memory of the previous frame, thereby enhancing the model's adaptability to changes in target appearance and tracking accuracy.

[0037] Furthermore, the multimodal fusion tracking technology combines information from RGB cameras, infrared cameras and radars, and uses a deep learning model to learn how to effectively fuse features of different modalities. It assigns a weight to each modality by using a weighted summation method. The algorithm formula of the weighted summation includes:

[0038]

[0039] Among them, F is the fused feature, F i is the feature of the i-th mode, w iis the weight of the ith mode, and Weight w i It is learned to maximize the tracking accuracy.

[0040] As can be seen from the above, this multimodal fusion tracking technology combines information from RGB cameras, infrared cameras and radars, and uses a deep learning model to learn how to effectively fuse the features of different modalities, showing significant beneficial effects. This technology assigns a learned weight to each modality through weighted summation to maximize the accuracy of tracking. This fusion strategy makes full use of the complementarity of multimodal data, improves the robustness and accuracy of tracking, and can significantly improve tracking performance, especially in complex environments or when the target appearance changes greatly.

[0041] Furthermore, the attention mechanism focuses on the characteristics of the target itself and the interaction between the target and the background, improving the tracking performance in complex scenes. The attention mechanism introduces a model based on the soft attention mechanism to assign an attention weight to each position, and its formula includes:

[0042]

[0043] Among them, α i is the attention weight of the ith position, e i is the energy value of the ith position (which can be calculated by models such as convolutional neural networks), and m is the number of all possible positions.

[0044] As can be seen from the above, the attention mechanism significantly improves the tracking performance in complex scenes by focusing on the characteristics of the target itself and the interaction between the target and the background. The mechanism introduces a model based on the soft attention mechanism, which assigns an attention weight calculated based on the energy value to each position. This refined weight allocation strategy enables the model to focus resources on processing key information in the video and effectively suppress background interference, thereby maintaining high-accuracy tracking in complex and changing environments, and enhancing the robustness and practicality of target tracking.

[0045] Furthermore, the end-to-end trainable tracking framework uses deep reinforcement learning (including cross entropy loss function) to learn the tracking strategy directly from the tracking success rate. The algorithm formula of the deep reinforcement learning includes:

[0046]

[0047] Where C is the number of categories y i is the one-hot encoding of the true category, is the class probability predicted by the model.

[0048] As can be seen above, this end-to-end trainable tracking framework uses deep reinforcement learning (including the cross entropy loss function) to learn tracking strategies directly from the tracking success rate, and can integrate all steps such as feature extraction, target detection, tracking, and model updating to form an efficient and adaptable whole. The deep reinforcement learning algorithm enables the model to continuously optimize its tracking strategy based on the tracking success rate, thereby improving the accuracy and efficiency of tracking. This learning method directly based on the tracking success rate enables the model to adapt to various complex scenarios more quickly in practical applications, improving the overall performance and practicality of target tracking.

[0049] Furthermore, in the intelligent monitoring application, real-time analysis and alarm functions of tracking results are also included. When the target object exhibits abnormal behavior or exceeds the preset range, the system can automatically alarm, that is, the alarm threshold formula is introduced to trigger the alarm condition, and the formula includes:

[0050]

[0051] Among them, Alarm is the alarm state (True means alarm, False means no alarm), Deviation is the deviation value between the target object and the preset range, and Threshold is the alarm threshold; this formula describes how to trigger the alarm condition based on the deviation value and the alarm threshold.

[0052] As can be seen from the above, the present invention has shown strong practicality and safety by introducing real-time analysis and alarm functions of tracking results. This function can automatically detect whether the target object has abnormal behavior or exceeds the preset range, and accurately trigger the alarm condition using the alarm threshold formula. When the deviation value of the target object exceeds the set alarm threshold, the system immediately enters the alarm state, thereby achieving timely response and processing of potential safety hazards, greatly enhancing the initiative and intelligence of the intelligent monitoring system, and providing strong technical support for various security monitoring scenarios.

[0053] Application example: Application of an efficient and robust target tracking method in intelligent monitoring, applicable to the above-mentioned efficient and robust target tracking method, including:

[0054] First, the staff will build an initial model for the monitored target based on deep learning and adaptive tracking technology. This model has strong dynamic adaptability and can automatically adjust as the target's appearance changes, ensuring the accuracy and continuity of tracking.

[0055] Then, using multimodal fusion tracking technology, the system can simultaneously receive and process information from multiple sensors such as RGB cameras, infrared cameras, and radars. Through the effective fusion of deep learning models, these different modal data can complement each other and further improve the robustness and accuracy of tracking.

[0056] Then, the introduction of the attention mechanism enables the model to process key information in the video more intelligently. It can not only focus on the characteristics of the target itself, but also keenly capture the interaction between the target and the background, so as to maintain high-accuracy tracking in complex scenes.

[0057] In addition, the application also uses a fully end-to-end trainable tracking framework. This framework integrates all steps such as feature extraction, target detection, tracking, and model updating, greatly improving the adaptability and efficiency of the model. Through deep reinforcement learning, the model can directly learn from the tracking success rate and optimize its tracking strategy.

[0058] Finally, in the intelligent monitoring application, the system also has the function of real-time analysis and alarm of tracking results. Once the target object shows abnormal behavior or exceeds the preset range, the system will immediately trigger the alarm condition and automatically alarm. This function greatly enhances the security and practicality of the intelligent monitoring system and provides a strong guarantee for various security monitoring scenarios.

[0059] Furthermore, the effects of an efficient and robust target tracking method of this embodiment and a conventional target tracking method (comparative example) are compared to obtain the following table:

[0060]

[0061] From the above table, it can be seen that the present invention is compared with the traditional technology as follows:

[0062] 1) Accuracy: The accuracy of traditional target tracking methods is often limited when facing target occlusion, deformation, etc. However, the method of this embodiment can significantly improve the tracking accuracy through deep learning and adaptive tracking technology, as well as the introduction of attention mechanism.

[0063] 2) Robustness: Traditional methods are not robust enough when facing complex scenes or changes in target appearance, and are prone to tracking failure. The method of this embodiment enhances the robustness of the model through multimodal fusion tracking technology and an end-to-end trainable framework, so that it can maintain stable tracking performance in various scenarios.

[0064] 3) Adaptability to changes in target appearance: Traditional methods usually rely on manually designed features and tracking algorithms, and are difficult to adapt to significant changes in target appearance. However, the method of this embodiment can store and utilize key information about changes in target appearance by introducing adaptive tracking technology of a dynamic memory system, thereby achieving dynamic adaptation to changes in target appearance.

[0065] 4) Anti-interference of complex background: Traditional methods are easily disturbed in complex backgrounds, resulting in tracking failure. However, the method of this embodiment uses the attention mechanism to focus on the interaction between the target and the background, effectively suppressing background interference and improving tracking performance in complex scenes.

[0066] 5) Multimodal data fusion capability: Traditional methods can only use information from a single sensor and cannot fully utilize the complementarity of multimodal data. However, the method of this embodiment uses multimodal fusion tracking technology to combine information from different sensors, significantly improving the accuracy and robustness of tracking.

[0067] 6) Real-time analysis and alarm function: Traditional methods usually do not have real-time analysis and alarm functions. The method of this embodiment introduces these functions in the intelligent monitoring application. When the target object exhibits abnormal behavior or exceeds the preset range, the system can automatically alarm, thereby enhancing the security and practicality of the intelligent monitoring system.

[0068] Working principle: Based on deep learning and adaptive tracking technology, the initial model of the target is established. Multimodal fusion tracking technology is used to combine information from different sensors to improve tracking accuracy and robustness. The attention mechanism is introduced to enable the model to concentrate resources to process key information in the video to improve tracking accuracy. A fully end-to-end trainable tracking framework is designed to integrate all relevant steps to improve the adaptability and efficiency of the model. Finally, real-time analysis and alarm functions of tracking results are realized in intelligent monitoring applications to achieve continuous and accurate tracking of the target object.

[0069] An embodiment of the present application provides an electronic device applicable to the above-mentioned efficient and robust target tracking method, including:

[0070] Memory, used to protect computer programs and data;

[0071] Processor, used to run system programs.

[0072] An embodiment of the present application provides a computer storage medium, which is applicable to the above-mentioned efficient and robust target tracking method, and performs hierarchical confidentiality management on the above-mentioned system and data in accordance with confidentiality management requirements.

[0073] Those skilled in the art will appreciate that the embodiments of the present application may be provided as a system or a computer program product. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0074] The present application is described with reference to the flowcharts and / or block diagrams of the devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0075] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0076] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0077] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0078] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0079] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0080] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, commodity or device including the elements.

[0081] The embodiments of the present invention are provided for the purpose of illustration and description. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations of the present invention. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. An efficient and robust target tracking method, characterized by: include: S1. Based on deep learning and adaptive tracking technology, an initial model of the target is established, which can dynamically adapt to changes in the target's appearance; S2, using multimodal fusion tracking technology to combine information from different sensors; S3, introduces the attention mechanism to enable the model to focus resources on processing key information in the video; S4. Design a fully end-to-end trainable tracking framework that integrates feature extraction, object detection, tracking, and model updating steps; S5. In intelligent monitoring applications, the above target tracking method is applied to real-time video stream processing.

2. An efficient and robust target tracking method as claimed in claim 1, characterized in that: The adaptive tracking technology includes the introduction of a dynamic memory system that can store key information about the target's appearance changes during the tracking process and use this information to adjust the model.

3. An efficient and robust target tracking method as claimed in claim 1, characterized in that: The multimodal fusion tracking technology combines information from RGB cameras, infrared cameras and radars, and uses a deep learning model to learn how to effectively fuse features of different modalities by assigning a weight to each modality using a weighted summation approach.

4. An efficient and robust target tracking method as claimed in claim 1, characterized in that: The attention mechanism focuses on the characteristics of the target itself and the interaction between the target and the background.

5. An efficient and robust target tracking method as claimed in claim 1, characterized in that: The proposed end-to-end trainable tracking framework adopts deep reinforcement learning to learn tracking policies directly from tracking success rates.

6. An efficient and robust target tracking method as claimed in claim 1, characterized in that: In intelligent monitoring applications, real-time analysis and alarm functions of tracking results are also included. When the target object exhibits abnormal behavior or exceeds the preset range, the system can automatically alarm.

Citation Information

Cited By

  • Light and small multispectral photoelectric recognition system and method based on intelligent algorithm

    CN121500326A