Visual target tracking method, system and device and storage medium

By enhancing image features in the Fourier frequency domain and combining time domain queries to generate spatiotemporal information decoding features, the tracking accuracy and efficiency issues of existing visual target tracking methods under extreme lighting and high-speed motion are solved, achieving more efficient target tracking.

CN120635148APending Publication Date: 2025-09-12UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510770326.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing visual target tracking methods based on event cameras fail to fully utilize the high-frequency spatiotemporal information of event signals, have difficulty in accurately tracking targets in harsh scenarios such as extreme lighting and high-speed motion, and have low computational efficiency.

Method used

By utilizing event voxel features in the Fourier frequency domain to enhance image features and combining them with time domain queries to generate spatiotemporal information decoding features, efficient target tracking is achieved.

Benefits of technology

It improves the accuracy and computational efficiency of target tracking in harsh scenarios, can better adapt to changes in target motion patterns, and capture the target's appearance details and historical motion information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635148A_ABST
    Figure CN120635148A_ABST
Patent Text Reader

Abstract

The invention discloses a visual target tracking method, system and device and a storage medium, which are corresponding schemes, in the scheme, visual target tracking is driven based on an event signal, and higher calculation efficiency is shown while the target tracking accuracy in a severe scene is ensured; moreover, the appearance details of the target are enhanced by fully utilizing the high-frequency characteristic of the event signal in the Fourier frequency domain space so as to adapt to target tracking in an extreme illumination scene; besides, target motion clues with high time resolution in the event signals are extracted in the time frequency domain space, the target is tracked robustly through space-time information decoding, the historical motion state of the fine granularity of the target can be memorized, reasoning and tracking are carried out, and the method can better adapt to challenging scenes in which the target motion mode changes. Generally speaking, the appearance details and the historical motion information of the moving target can be effectively captured by using the high-frequency spatio-temporal information of the event data, and the method can better adapt to a target tracking task in a severe scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual target tracking technology, and in particular to a visual target tracking method, system, device and storage medium. Background Art

[0002] In recent years, visual object tracking technology based on event cameras has rapidly developed and is widely used in various fields, including intelligent surveillance, autonomous driving, and drone navigation. Compared with visual object tracking technology based solely on RGB frame cameras, event-based visual object tracking models incorporate event modality data, which can obtain target motion information with higher temporal resolution (up to 1MHz) and finer target appearance details (such as local texture and edge contours), making them better able to adapt to harsh scenes.

[0003] Previous event-camera-based visual object tracking methods mostly compress event streams into frame-like images, neglecting the high-temporal resolution information inherent in the event signal itself. Furthermore, these tracking models focus solely on template similarity matching within frames, failing to leverage the rich inter-frame information in the event stream. Consequently, they are unable to memorize the target's fine-grained historical motion cues, and their ability to understand and predict target motion patterns is limited, making them incapable of adapting to harsh scenarios such as high-speed motion or background clutter. Finally, these tracking models perform feature fusion only in the spatial domain. This implicit fusion approach fails to fully exploit the high-frequency appearance details provided by the event, resulting in limited performance in challenging scenarios such as extreme lighting.

[0004] In summary, although existing methods have attempted to introduce event signals to drive visual target tracking tasks, they have failed to fully utilize the high-frequency spatiotemporal information of event signals and find it difficult to effectively memorize the fine-grained historical motion state of the target for reasoning. These unexplored directions and unresolved problems provide directions for further research. Summary of the Invention

[0005] The purpose of the present invention is to provide a visual target tracking method, system, device and storage medium, which have the advantages of better tracking effect, higher calculation efficiency and greater practical value.

[0006] The purpose of the present invention is achieved through the following technical solutions:

[0007] A visual target tracking method, comprising:

[0008] Step 1: Obtain a frame image and event voxels within the frame at the current moment, and extract image features and event voxel features therefrom; wherein the event voxels within the frame are obtained by voxelizing the event sequence within the frame;

[0009] Step 2: Enhance the image features using the event voxel features in the Fourier frequency domain to obtain the frequency-domain enhanced image features at the current moment;

[0010] Step 3: Combine the event voxel features and the time domain query at the previous moment to generate a time domain query that records the target's historical motion information at the current moment, and use the frequency domain enhanced image features at the current moment to perform target positioning based on the current time domain query, thereby obtaining the time domain query propagated to the next moment and the spatiotemporal information decoding features at the current moment;

[0011] Step 4: Use the spatiotemporal information decoding features at the current moment to obtain the predicted position of the target frame at the current moment.

[0012] A visual target tracking system for implementing the aforementioned method comprises: a visual target tracking neural network; the visual target tracking neural network comprises:

[0013] A feature extraction module is used to obtain the current frame image and the event voxels within the frame, and extract image features and event voxel features therefrom; wherein the event voxels within the frame are obtained by voxelizing the event sequence within the frame;

[0014] The frequency domain feature enhancement module is used to enhance the image features using the event voxel features in the Fourier frequency domain space to obtain the image features after frequency domain enhancement at the current moment;

[0015] a spatiotemporal information decoding module, configured to combine the event voxel features and the temporal query at the previous moment to generate a temporal query that records the target's historical motion information up to the current moment, and to guide the frequency-domain enhanced image features at the current moment to perform target positioning based on the temporal query at the current moment, thereby obtaining the temporal query propagated to the next moment and the spatiotemporal information decoding features at the current moment;

[0016] The prediction head is used to use the spatiotemporal information decoding features at the current moment to obtain the predicted position of the target box at the current moment.

[0017] A processing device comprising: one or more processors; a memory for storing one or more programs;

[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.

[0019] A readable storage medium stores a computer program, which implements the aforementioned method when the computer program is executed by a processor.

[0020] It can be seen from the technical solutions provided by the present invention that 1) visual target tracking driven by high-frequency spatiotemporal information of event signals shows higher computational efficiency while ensuring target tracking accuracy in harsh scenarios; 2) the high-frequency characteristics of event signals are fully utilized in the Fourier frequency domain to enhance the appearance details of the target to adapt to target tracking in extreme lighting scenarios; 3) target motion clues with high temporal resolution are extracted from event signals in the time-frequency domain, and the target is robustly tracked through spatiotemporal information decoding, which can more effectively memorize the fine-grained historical motion state of the target, and perform inference tracking based on this, which can better adapt to challenging scenarios where the target motion pattern changes (such as high-speed motion). In general, the present invention can achieve effective capture of the appearance details and historical motion information of the moving target by utilizing the high-frequency spatiotemporal information of event data, and better adapt to target tracking tasks in harsh scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 A flowchart of a visual target tracking method provided by an embodiment of the present invention;

[0023] Figure 2 A schematic diagram of a visual target tracking neural network provided by an embodiment of the present invention;

[0024] Figure 3 A schematic diagram of a frequency domain feature enhancement module provided in an embodiment of the present invention;

[0025] Figure 4 A schematic diagram of a spatiotemporal information decoding module provided in an embodiment of the present invention;

[0026] Figure 5 A schematic diagram of a processing device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0028] First, the following terms may be used in this article:

[0029] The terms "include," "comprises," "contains," "has," or other similar expressions should be interpreted as non-exclusive. For example, "including certain technical features (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products, or manufactured articles, etc.) should be interpreted as including not only the technical features explicitly listed, but also other technical features known in the art that are not explicitly listed.

[0030] The term "consisting of" excludes any technical features not explicitly listed. If used in a claim, this term renders the claim closed, excluding any technical features other than those explicitly listed, except for conventional impurities associated with them. If this term appears only in a clause of a claim, it limits only the elements explicitly listed in that clause; elements listed in other clauses are not excluded from the claim as a whole.

[0031] Unless otherwise specified or limited, the terms "mounted," "connected," "connect," and "fixed" should be interpreted broadly. For example, they can refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediary; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in this document based on specific circumstances.

[0032] The following describes in detail a visual target tracking method, system, device, and storage medium provided by the present invention. Any information not described in detail in the embodiments of the present invention is prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of the present invention, the procedures are performed in accordance with conventional conditions in the art or the conditions recommended by the manufacturer. Instruments used in the embodiments of the present invention, where the manufacturer is not specified, are all commercially available conventional products.

[0033] Example 1

[0034] The embodiment of the present invention provides a visual target tracking method, which is based on the high-frequency spatiotemporal information of event signals. By decoding the spatiotemporal information, the method can accurately track the target object in adverse scenes such as extreme lighting and high-speed motion. This method can make up for the defect that the existing target tracker based on event signals is difficult to adapt to adverse scenes such as extreme lighting and high-speed motion. Figure 1 As shown, it mainly includes the following steps:

[0035] Step 1: Obtain the current frame image and intra-frame event voxels, and extract image features and event voxel features therefrom; wherein the intra-frame event voxels are obtained by voxelizing the intra-frame event sequence.

[0036] The preferred implementation of this step includes: obtaining the template frame image at the current moment Search frame image Intra-frame template event voxel V1 T and search event voxel V within the frame t S ; Wherein, t represents the tth moment, which is the current moment; template frame image is the target area image captured from the first frame of the video sequence, is the frame image at the current time t in the video sequence, and the template event voxel V1 in the frame T Template frame image Corresponding to the event voxel information within the frame, search for event voxels V within the frame t S To search for frame images Corresponding to the voxelized information of the event in the frame. Patch embedding is performed through the patch embedding layer, and then feature extraction and feature association are performed through the visual converter to obtain the search frame image features at the current moment. and intra-frame search event features

[0037] Step 2: Enhance the image features using the event voxel features in the Fourier frequency domain to obtain the frequency-domain enhanced image features at the current moment.

[0038] The preferred implementation of this step is as follows:

[0039] (1) Search frame image features at the current moment and intra-frame search event features Perform two-dimensional fast Fourier transform processing respectively, and get Amplitude and phase as well as Amplitude and phase

[0040] (2) Phase fusion based on fully connected modules and Get the phase of the enhanced feature Expressed as:

[0041]

[0042] Among them, C(·) represents the convolutional layer, A(·) represents the ReLU activation function, and [·] represents channel cascade.

[0043] (3) Processing amplitude based on fully connected activation module and Get the amplitude after feature filtering Expressed as:

[0044]

[0045] Where J(a) is the amplitude filter, S(·) represents the Sigmoid function, and × represents element-by-element multiplication.

[0046] (4) Phase combination with enhanced features and the amplitude after feature filtering Through the two-dimensional inverse Fourier transform, the image feature F after frequency domain enhancement at the current moment is obtained t , expressed as:

[0047]

[0048] Among them, F -1 (·) represents inverse Fourier transform.

[0049] Step 3: Combine the event voxel features and the time domain query at the previous moment to generate a time domain query that records the historical motion information of the target at the current moment, and based on the current time domain query, guide the image features after frequency domain enhancement at the current moment to perform target positioning, obtain the time domain query propagated to the next moment, and the spatiotemporal information decoding features at the current moment.

[0050] In the embodiment of the present invention, the event voxel features and the time domain query at the previous moment are combined to generate a time domain query that records the target's historical motion information at the current moment, which is expressed as:

[0051]

[0052] in, is the time domain query of the previous moment, q t For the time domain query of the target's historical motion information at the current moment, for The query matrix, For the The query weight matrix, is the event voxel feature The bond matrix, for The bond weight matrix, V t E for The value matrix of for The value weight matrix of , (·) T represents the matrix transpose operation, softmax(·) is the normalized exponential function, and D represents the extracted feature dimension.

[0053] In the embodiment of the present invention, the image features after frequency domain enhancement at the current moment are guided to perform target positioning based on the current moment time domain query, and the time domain query propagated to the next moment and the spatiotemporal information decoding features at the current moment are obtained, which are expressed as:

[0054] B t =[q t ; F t ]

[0055]

[0056] Among them, B t The time domain query q for the target historical motion information at the current moment t and the frequency domain enhanced feature F at the current moment t cascade; B t The query matrix, B t The query weight matrix, B t The bond matrix, For B t The bond weight matrix, V t B For B t The value matrix of For B t The value weight matrix of , (·) T represents the matrix transpose operation, softmax(·) is the normalized exponential function, D represents the extracted feature dimension, G t is an intermediate variable, reflecting the features after initial attention enhancement, and FFN(·) is a feedforward neural network; For time domain queries that propagate to the next moment, Decode the spatiotemporal information features for the current moment.

[0057] Step 4: Use the spatiotemporal information decoding features at the current moment to obtain the predicted position of the target frame at the current moment.

[0058] The preferred implementation of this step is as follows: using the spatiotemporal information decoding features at the current moment, respectively predict the center coordinates, quantized offset, and length and width of the target frame, and after integration, obtain the predicted position of the target frame at the current moment.

[0059] Preferably: step 1 is implemented by a feature extraction module, step 2 is implemented by a frequency domain feature enhancement module, step 3 is implemented by a spatiotemporal information decoding module, and step 4 is implemented by a prediction head; the feature extraction module, frequency domain feature enhancement module, spatiotemporal information decoding module and prediction head constitute a visual target tracking neural network.

[0060] The visual target tracking neural network is trained in advance. During training, a training video image dataset X and its corresponding label Y and event sequence are obtained, and the event sequence is characterized to obtain a training video event dataset Z. The video image dataset X and the training video event dataset Z are input into the visual target tracking neural network to obtain the predicted position of the corresponding target frame, and a loss function is constructed in combination with the corresponding label Y. The loss function is used to optimize the parameters of the visual target tracking neural network.

[0061] Among them, the loss function L is:

[0062] L=L focal +λ1L1+λ2L GIOU

[0063] Among them, L focal is the focal classification loss, L1 is the absolute value loss, L GIOU is the generalized intersection-over-union loss, and λ1 and λ2 are two hyperparameters used to control the importance of absolute value loss and generalized intersection-over-union loss.

[0064] In order to more clearly demonstrate the technical solution and technical effects provided by the present invention, the method provided by the embodiment of the present invention is described in detail below with reference to specific embodiments.

[0065] 1. Overall overview of the plan.

[0066] In order to make up for the defect that existing event signal-based target trackers are difficult to adapt to harsh scenes such as extreme lighting and high-speed motion, a visual target tracking technology driven by high-frequency spatiotemporal information based on event signals is proposed. In the field of target tracking based on event signals, the present invention uses the high-frequency characteristics of event signals in the spatial Fourier frequency domain for the first time to enhance the appearance details of the target. At the same time, the present invention extracts high-temporal resolution target motion clues from event signals in the time-frequency domain for the first time, and guides the tracker to robustly track the target through spatiotemporal information decoding. Compared with traditional solutions, the present invention can fully capture the high-frequency details of the target appearance even in extreme lighting environments, and can better remember the fine-grained historical motion state of high-speed moving targets, and perform inference tracking based on this. The present invention not only achieves a significant improvement in visual effects, but also reaches the industry-leading level in multiple quantitative performance indicators.

[0067] 2. Detailed description of the plan.

[0068] In this embodiment of the present invention, a visual object tracking neural network was constructed, comprising a feature extraction module, a frequency domain feature enhancement module, a spatiotemporal information decoding module, and a prediction head. This neural network was trained using collected training data. Subsequently, the model was evaluated using a real-world object tracking dataset to demonstrate its visual object tracking capabilities and generalization performance. The following describes the processing of training data, the structure and principles of the visual object tracking neural network, and its training process.

[0069] 1. Processing of training data.

[0070] Obtain a training video image dataset X and its corresponding label Y and event sequence, and characterize the event sequence to obtain a training video event dataset Z.

[0071] (1) Obtain the training video image dataset X and its corresponding label Y.

[0072] Get the training video image dataset X={x1,x2,…,x i ,…,x M}, where x i Represents the i-th video sequence, i = 1, 2, ..., M, and M is the total number of videos in the dataset.

[0073] at the same time, Among them I j Represents the j-th frame image in the i-th video sequence, j = 1, 2, ..., N i , N i is the total number of frame images in the i-th video sequence.

[0074] The label set corresponding to X is recorded as Y = {y1,y2,...,y i ,...,y M}, where y i Represents the label of the i-th video sequence, i = 1, 2, ..., M, and M is the total number of videos in the dataset.

[0075] (2) Characterize the event sequence and obtain the training video event dataset Z.

[0076] Get the intra-frame event sequence U={u1,u2,...,u i ,...,u M}, where u i represents the intra-frame event sequence corresponding to the i-th video sequence, i = 1, 2, ..., M, and M is the total number of videos in the dataset.

[0077] at the same time, where z jRepresents the j-th intra-frame event sequence in the i-th video sequence, j = 1, 2, ..., N i , N i is the total number of intra-frame event sequences in the i-th video sequence.

[0078] For the intra-frame event sequence z j Perform voxel representation to obtain the intra-frame event voxel sequence corresponding to the i-th video sequence in, represents z j The corresponding event voxel in the frame, B represents the number of event voxel channels, H represents the image height, and W represents the image width;

[0079] Thus, we can obtain the training video event dataset Z = {o1, o2, ..., o i ,...,o M}, i=1,2,...,M.

[0080] 2. The structure and principle of visual target tracking neural network.

[0081] like Figure 2 As shown in the figure, the visual target tracking neural network mainly includes: feature extraction module, frequency domain feature enhancement module, spatiotemporal information decoding module and prediction head.

[0082] (1) Feature extraction module.

[0083] The feature extraction module includes a patch embedding layer and a visual converter, wherein the patch embedding layer is responsible for Search frame image Intra-frame template event voxel V1 T and search event voxel V within the frame t S Perform patch embedding. The visual converter is responsible for feature extraction and feature association of the above embedded patches, and finally outputs the search frame image features at the tth moment. and intra-frame search event features Where N represents the number of embedded patches and D represents the dimension of extracted features.

[0084] Specifically, given a video sequence, the first frame is taken and the 4-dimensional coordinates (x, y, w, h) of the target bounding box in the image are given (x and y are the horizontal and vertical coordinates of the target box center, respectively, and w and h are the width and height of the target box, respectively). Finally, a square area with an area of ​​w × h × factor is framed in the first frame, centered at (x, y). Factor is a set coefficient, for example, factor = 2. This square area is the target area image and serves as the template frame image, which marks the target to be tracked. All images at subsequent moments are search frame images. The visual single target tracking task is to track the target in the template frame image in subsequent video frames (search frame images) based on the given template frame image. During training, the in-frame template event voxels and in-frame search event voxels are both from the training video event dataset. They correspond to the template frame image and the search frame image, respectively.

[0085] (2) Frequency domain feature enhancement module.

[0086] The frequency domain feature enhancement module includes: Fourier transform layer, phase fusion layer, amplitude filter layer and inverse Fourier transform layer, which are responsible for utilizing the Fourier frequency domain space to enhance the The high frequency characteristics of Perform enhancement processing and finally output the image features after frequency domain enhancement at time t like Figure 3 As shown in FIG, it is a structural diagram of the frequency domain feature enhancement module.

[0087] (2.1) The Fourier transform layer is composed of Fourier transform modules, which will transform the input and Perform two-dimensional fast Fourier transform processing respectively, and get Amplitude and phase as well as Amplitude and phase

[0088] (2.2) The phase fusion layer is based on the fully connected module fusion phase and The phase of the enhanced feature after fusion is obtained using formula (1)

[0089]

[0090] Among them, C(·) represents the convolutional layer, A(·) represents the ReLU activation function, and [·] represents channel cascade.

[0091] (2.3) The amplitude filter layer is based on the fully connected activation module processing and Using formula (2) we can get the amplitude filter Then use formula (3) to get the amplitude after feature filtering through J(a)

[0092]

[0093] Where J(a) is the amplitude filter, S(·) represents the Sigmoid function, and × represents element-by-element multiplication.

[0094] (2.4) The inverse Fourier transform layer is based on the two-dimensional inverse fast Fourier transform to transform the enhanced amplitude and phase Transform from the frequency domain back to the image domain and use formula (4) to obtain the image features after frequency domain enhancement

[0095]

[0096] Among them, F -1 (·) represents inverse Fourier transform.

[0097] (3) Spatiotemporal information decoding module.

[0098] The spatiotemporal information decoding module includes a temporal query generation layer and a spatiotemporal information attention layer, wherein the temporal query generation layer is responsible for generating a temporal query that records the historical motion information of the target up to the tth moment. The spatiotemporal information attention layer is responsible for t Guide the image feature F after frequency domain enhancement t Execute target positioning and finally output the time domain query for the next moment and the spatiotemporal information decoding features at time t like Figure 4 As shown in FIG, it is a structural diagram of the spatiotemporal information decoding module.

[0099] (3.1) The temporal query generation layer is composed of a “temporal query-event” mutual attention module, which uses the temporal query propagated from the t-1th moment (the previous moment) To record the intra-frame search event characteristics The target historical motion information in the time domain query at time t is obtained using equations (5) to (8)

[0100]

[0101] in, for The query matrix, for The query weight matrix, for The key matrix, for The key weight matrix, for The value matrix of for The value weight matrix of , (·) T represents the matrix transpose operation, and softmax(·) represents the activation function (normalized exponential function).

[0102] (3.2) The spatiotemporal information attention layer is composed of the self-attention module of “time domain query-frequency domain enhancement feature”, which generates the time domain query q t and the frequency domain enhanced feature F t Features obtained after cascading Use equations (9) to (14) to obtain the time domain query for propagation to the next moment and the spatiotemporal information decoding features at time t

[0103] B t =[q t ; F t ] (9)

[0104]

[0105] in, For B t The query matrix, For B t The query weight matrix, For B t The key matrix, For B t The key weight matrix, For B t The value matrix of For B t The value weight matrix of , (·) T represents the matrix transpose operation, G t is an intermediate variable, reflecting the features after initial attention enhancement, and FFN(·) is a feedforward neural network.

[0106] (4) Prediction head.

[0107] In the embodiment of the present invention, the prediction head includes: a center coordinate prediction head, an offset prediction head and a size prediction head. The above three prediction heads are composed of convolutional layers and are responsible for decoding features based on spatiotemporal information. The center coordinates, quantized offset, and length and width of the target frame are predicted respectively. Combining the above three pieces of information, the predicted coordinates of the target frame are finally obtained.

[0108] 3. Training process.

[0109] The training data obtained in the first part (i.e., the training video image dataset X and the training video event dataset Z) are input into the visual target tracking neural network. After the calculation in the second part, the predicted coordinates of the target box are obtained. Combined with the corresponding label Y, the loss function L shown in formula (15) is calculated:

[0110] L=L focal +λ1L1+λ2L GIOU (15)

[0111] Among them, L focal is the focal classification loss, L1 is the absolute value loss, L GIOU is the generalized intersection-over-union loss, λ1 and λ2 are two hyperparameters used to control the importance of absolute value loss and generalized intersection-over-union loss. For example, λ1 can be set to 5 and λ2 can be set to 2.

[0112] The visual target tracking neural network is trained using the gradient descent method, and the loss function L is calculated to update the network parameters. When the number of training iterations reaches the set number or the loss function L converges, the training stops, thereby obtaining the optimal visual target tracking model for accurately tracking targets in harsh scenarios.

[0113] Through the description of the above embodiments, those skilled in the art will clearly understand that the above embodiments can be implemented through software or by using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) and includes a number of instructions for causing a computer device (such as a personal computer, a server, or a network device) to execute the methods described in the various embodiments of the present invention.

[0114] Example 2

[0115] The present invention also provides a visual target tracking system, which is mainly used to implement the method provided in the above embodiment. The system mainly includes: a visual target tracking neural network; the visual target tracking neural network includes:

[0116] A feature extraction module is used to obtain the current frame image and the event voxels within the frame, and extract image features and event voxel features therefrom; wherein the event voxels within the frame are obtained by voxelizing the event sequence within the frame;

[0117] The frequency domain feature enhancement module is used to enhance the image features using the event voxel features in the Fourier frequency domain space to obtain the image features after frequency domain enhancement at the current moment;

[0118] a spatiotemporal information decoding module, configured to combine the event voxel features and the temporal query at the previous moment to generate a temporal query that records the target's historical motion information up to the current moment, and to guide the frequency-domain enhanced image features at the current moment to perform target positioning based on the temporal query at the current moment, thereby obtaining the temporal query propagated to the next moment and the spatiotemporal information decoding features at the current moment;

[0119] The prediction head is used to use the spatiotemporal information decoding features at the current moment to obtain the predicted position of the target box at the current moment.

[0120] Considering that the main processing procedures involved in the system have been introduced in detail in the previous embodiments, they will not be repeated here.

[0121] Those skilled in the art will clearly understand that for the convenience and brevity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.

[0122] Example 3

[0123] The present invention also provides a processing device, such as Figure 5 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the aforementioned embodiment.

[0124] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.

[0125] In the embodiment of the present invention, the specific types of the memory, input device, and output device are not limited; for example:

[0126] The input device can be a touch screen, image acquisition device, physical button or mouse;

[0127] The output device may be a display terminal;

[0128] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.

[0129] Example 4

[0130] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.

[0131] In the embodiments of the present invention, the computer-readable storage medium may be provided in the aforementioned processing device, for example, as a memory in the processing device. Alternatively, the computer-readable storage medium may be a USB flash drive, a removable hard drive, a read-only memory (ROM), a magnetic disk, or an optical disk, among other media capable of storing program code.

[0132] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims. The information disclosed in the background technology section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that the information constitutes prior art already known to those skilled in the art.

Claims

1. A visual target tracking method, characterized in that: include: Step 1: Obtain a frame image and event voxels within the frame at the current moment, and extract image features and event voxel features therefrom; wherein the event voxels within the frame are obtained by voxelizing the event sequence within the frame; Step 2: Enhance the image features using the event voxel features in the Fourier frequency domain to obtain the frequency-domain enhanced image features at the current moment; Step 3: Combine the event voxel features and the time domain query at the previous moment to generate a time domain query that records the target's historical motion information at the current moment, and use the frequency domain enhanced image features at the current moment to perform target positioning based on the current time domain query, thereby obtaining the time domain query propagated to the next moment and the spatiotemporal information decoding features at the current moment; Step 4: Use the spatiotemporal information decoding features at the current moment to obtain the predicted position of the target frame at the current moment.

2. A visual target tracking method according to claim 1, characterized in that: The step of obtaining the current frame image and event voxels within the frame and extracting image features and event voxel features therefrom includes: Get the template frame image at the current moment Search frame image Intra-frame template event voxel V1 T and search event voxel V within the frame t S ; Wherein, t represents the tth moment, which is the current moment; template frame image is the target area image captured from the first frame of the video sequence, is the frame image at the current time t in the video sequence, and the template event voxel V1 in the frame T Template frame image Corresponding to the event voxel information within the frame, search for event voxels V within the frame t S To search for frame images Corresponding intra-frame event voxel information Patch embedding is performed through the patch embedding layer, and then feature extraction and feature association are performed through the visual converter to obtain the search frame image features at the current moment and intra-frame search event features 3. A visual target tracking method according to claim 2, characterized in that: The image features are enhanced using event voxel features in the Fourier frequency domain space to obtain the image features after frequency domain enhancement at the current moment, including: Search frame image features at the current moment and intra-frame search event features Perform two-dimensional fast Fourier transform processing respectively, and get Amplitude and phase as well as Amplitude and phase Phase fusion based on fully connected modules and Get the phase of the enhanced feature Expressed as: Where C(·) represents the convolutional layer, A(·) represents the ReLU activation function, and [·] represents the channel cascade; Processing amplitude based on fully connected activation modules and Get the amplitude after feature filtering Expressed as: Where J(a) is the amplitude filter, S(·) represents the Sigmoid function, and × represents element-by-element multiplication. Phase combined with enhanced features and the amplitude after feature filtering Through the two-dimensional inverse Fourier transform, the image feature F after frequency domain enhancement at the current moment is obtained t , expressed as: Among them, F -1 (·) represents inverse Fourier transform.

4. A visual target tracking method according to claim 1 or 2, characterized in that: The event voxel features and the time domain query at the previous moment are combined to generate a time domain query that records the target's historical motion information at the current moment, which is expressed as: in, is the time domain query of the previous moment, q t For the time domain query of the target's historical motion information at the current moment, for The query matrix, For the The query weight matrix, is the event voxel feature The bond matrix, for The bond weight matrix, V t E for The value matrix of for The value weight matrix of , (·) T represents the matrix transpose operation, D represents the extracted feature dimension, and softmax(·) is the normalized exponential function.

5. A visual target tracking method according to claim 1 or 2, characterized in that: The time domain query at the current moment guides the image features after frequency domain enhancement at the current moment to perform target positioning, obtains the time domain query propagated to the next moment, and the spatiotemporal information decoding features at the current moment, which are expressed as: B t =[q t ;F t ] Among them, B t The time domain query q for the target historical motion information at the current moment t and the frequency domain enhanced feature F at the current moment t cascade; For B t The query matrix, For B t The query weight matrix, For B t The bond matrix, For B t The bond weight matrix, V t B For B t The value matrix of For B t The value weight matrix of , (·) T represents the matrix transpose operation, D represents the extracted feature dimension, G t is an intermediate variable, softmax(·) is a normalized exponential function, and FFN(·) is a feedforward neural network; For time domain queries that propagate to the next moment, Decode the spatiotemporal information features for the current moment.

6. A visual target tracking method according to claim 1, characterized in that: The method of using the spatiotemporal information decoding feature at the current moment to obtain the predicted position of the target frame at the current moment includes: By using the spatiotemporal information decoding features at the current moment, the center coordinates, quantized offset, and length and width of the target frame are predicted respectively. After integration, the predicted position of the target frame at the current moment is obtained.

7. A visual target tracking method according to claim 1, characterized in that: Also includes: Step 1 is implemented by a feature extraction module, step 2 is implemented by a frequency domain feature enhancement module, step 3 is implemented by a spatiotemporal information decoding module, and step 4 is implemented by a prediction head; the feature extraction module, frequency domain feature enhancement module, spatiotemporal information decoding module and prediction head constitute a visual target tracking neural network; Pre-training the visual target tracking neural network, during which a training video image dataset X and its corresponding labels Y and event sequences are obtained, and the event sequences are characterized to obtain a training video event dataset Z; The video image dataset X and the training video event dataset Z are input into the visual target tracking neural network to obtain the predicted position of the corresponding target box, and the loss function is constructed in combination with the corresponding label Y. The loss function is used to optimize the parameters of the visual target tracking neural network. Among them, the loss function L is: L=L focal +λ1L1+λ2L GIOU Among them, L focal is the focal classification loss, L1 is the absolute value loss, L GIOU is the generalized intersection-over-union loss, and λ1 and λ2 are two hyperparameters used to control the importance of absolute value loss and generalized intersection-over-union loss.

8. A visual target tracking system, characterized in that: The method for implementing any one of claims 1 to 7 comprises: a visual target tracking neural network; the visual target tracking neural network comprises: A feature extraction module is used to obtain the current frame image and the event voxels within the frame, and extract image features and event voxel features therefrom; wherein the event voxels within the frame are obtained by voxelizing the event sequence within the frame; The frequency domain feature enhancement module is used to enhance the image features using the event voxel features in the Fourier frequency domain space to obtain the image features after frequency domain enhancement at the current moment; a spatiotemporal information decoding module, configured to combine the event voxel features and the temporal query at the previous moment to generate a temporal query that records the target's historical motion information up to the current moment, and to guide the frequency-domain enhanced image features at the current moment to perform target positioning based on the temporal query at the current moment, thereby obtaining the temporal query propagated to the next moment and the spatiotemporal information decoding features at the current moment; The prediction head is used to use the spatiotemporal information decoding features at the current moment to obtain the predicted position of the target box at the current moment.

9. A processing device, characterized in that include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.