Target tracking method and device, equipment, medium and product
By encoding and decoding image sequences, combined with a motion-sensing memory library and prompting information, the problem of insufficient target tracking by the network robot in dynamic environments is solved, achieving more efficient and accurate target recognition and tracking.
Patent Information
- Application Number
- CN202511557131.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-01-13
AI Technical Summary
In dynamic environments, the branch robots lack the ability to track dynamic targets of customers, resulting in significant bottlenecks. In particular, they exhibit poor robustness in occluded or complex environments and have weak cross-domain generalization capabilities, leading to the easy loss of targets and an inability to provide efficient and accurate customized services.
By acquiring image sequences, feature data is extracted using an image encoder and a memory time-series attention mechanism module. Combined with the decoding processing of a motion-aware memory library and a cue encoder, the accuracy of target tracking is improved.
It improves the correlation and accuracy of target tracking over time, ensures stable tracking of targets in complex environments, and enhances the efficiency and precision of robot services.
Smart Images

Figure CN121330318A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology or other related fields, and in particular to a target tracking method, apparatus, device, medium and product. Background Technology
[0002] With the deep integration of fintech and digital transformation, bank branches are transforming from traditional service models to intelligent, scenario-based, and contactless service models. Among these, branch service robots, as the core carrier of smart banking construction, have gradually taken on functions such as customer guidance, business diversion, precision marketing, and risk control.
[0003] In some solutions, when it is necessary for network robots or other devices to track dynamic targets, the dynamic target tracking capability still has significant bottlenecks, and there is an urgent need to propose a solution to improve the accuracy of dynamic target tracking. Summary of the Invention
[0004] This application provides a target tracking method, apparatus, device, medium, and product to improve the accuracy of target tracking.
[0005] In a first aspect, this application provides a target tracking method, comprising: acquiring a first image sequence; the first image sequence comprising image frame data of the target to be tracked in a time series;
[0006] The first image sequence is input into an image encoder for encoding processing to obtain a first feature data set.
[0007] The first feature data set and the motion-aware memory library are input into the memory time-series attention mechanism module for feature extraction to obtain the second feature data set; the motion-aware memory library includes feature information corresponding to image frame data at historical time points.
[0008] The prompt information corresponding to the first image sequence is input into the prompt encoder for encoding processing to obtain a prompt feature data set.
[0009] The second feature data set and the prompt feature data set are input into the mask decoder for decoding to obtain the first decoded image set; and the target to be tracked is determined based on the first decoded image set.
[0010] Secondly, this application provides a target tracking device, comprising:
[0011] The acquisition module is used to acquire a first image sequence; the first image sequence includes image frame data of the target to be tracked in a time series.
[0012] The first encoding module is used to input the first image sequence into the image encoder for encoding processing to obtain the first feature data set;
[0013] The extraction module is used to input the first feature data set and the motion-aware memory library into the memory time-series attention mechanism module for feature extraction to obtain the second feature data set; the motion-aware memory library includes feature information corresponding to image frame data at historical time points;
[0014] The second encoding module is used to input the prompt information corresponding to the first image sequence into the prompt encoder for encoding processing to obtain a prompt feature data set.
[0015] The decoding module is used to input the second feature data set and the prompt feature data set into the mask decoder for decoding processing to obtain a first decoded image set; and to determine the target to be tracked based on the first decoded image set.
[0016] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0017] The memory stores computer-executed instructions;
[0018] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0019] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0020] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0021] The target tracking method, apparatus, device, medium, and product provided in this application involve acquiring a first image sequence, which includes image frame data of the target to be tracked in a time series. The first image sequence is input into an image encoder for encoding processing to obtain a first feature data set. The first feature data set and a motion-aware memory library are input into a time-series attention mechanism module for feature extraction to obtain a second feature data set. The motion-aware memory library includes feature information corresponding to image frame data from historical time periods. The prompt information corresponding to the first image sequence is input into a prompt encoder for encoding processing to obtain a prompt feature data set. The second feature data set and the prompt feature data set are input into a mask decoder for decoding processing to obtain a first decoded image set. The target to be tracked is determined based on the first decoded image set. This solution, by encoding the image sequence of the target to be tracked in a time series and combining the feature data of image frames already identified in historical time periods during the tracking and recognition process, improves the temporal correlation of target tracking, thereby improving the accuracy of target tracking. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0023] Figure 1 An example of a target tracking method flowchart is shown. Figure 1 ;
[0024] Figure 2 An example illustration shows a schematic diagram of image style transfer based on a recurrent generative adversarial network;
[0025] Figure 3 An example is shown illustrating the flowchart of a target tracking method. Figure 2 ;
[0026] Figure 4 An example diagram of an inter-frame difference method is shown;
[0027] Figure 5 This is a schematic diagram illustrating the target tracking process as an example of this application;
[0028] Figure 6 An exemplary schematic diagram of a target tracking device is shown;
[0029] Figure 7 The diagram above illustrates the structure of an electronic device.
[0030] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0031] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0032] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning. The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise indicated. It should be understood that such terms can be used interchangeably where appropriate, for example, to be implemented in an order other than those given in the illustrations or descriptions of the embodiments of this application. The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application, are intended to be omnipresent but not exclusive. For example, a product or device that comprises a series of components is not necessarily limited to those components that are explicitly listed, but may include other components that are not explicitly listed or are inherent to such products or devices. The term "module" as used in this application refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing the functions associated with that element.
[0033] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0034] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0035] It should be noted that the target tracking method, apparatus, equipment, storage medium and product provided in this application can be used in the fintech field, or in any field other than fintech. The application field of the target tracking method, apparatus, equipment, medium and product in this application is not limited.
[0036] With the deep integration of fintech and digital transformation, bank branches are transforming from traditional service models to intelligent, scenario-based, and contactless service models. Branch service robots, as a core component of smart banking, are gradually taking on functions such as customer guidance, business triage, targeted marketing, and risk control.
[0037] However, significant bottlenecks remain in some target recognition and tracking solutions for dynamic environments, necessitating breakthroughs in computer vision technology to achieve more efficient and accurate service adaptation. In complex branch environments, robots lack the ability to capture and respond to customer dynamic behavior in real time, limiting the efficiency of robot services.
[0038] Some solutions rely primarily on multimodal sensor fusion, deep learning model optimization, and 3D pose estimation techniques for target tracking. However, these solutions suffer from limitations such as high hardware constraints, poor robustness to occlusion or complex environments, and weak cross-domain generalization.
[0039] Among these, the biggest constraint is hardware conditions: the camera's pixel count has a significant impact, as low-pixel cameras produce images with insufficient resolution, which reduces the accuracy of failures.
[0040] Its robustness to occlusion or complex environments is poor: it is prone to failure in extreme occlusion or multi-person scenarios. Furthermore, its target recognition accuracy varies under different lighting conditions in the same environment.
[0041] Cross-domain generalization ability is limited because current training datasets are mostly from the sports or medical fields, and there are few training schemes with higher general applicability. The same model may exhibit accuracy differences in different scenarios.
[0042] Conventional target recognition often uses fixed cameras to capture images of fixed areas, or can only identify targets in a single input image. This fails to meet the high responsiveness requirements of robots that need to constantly move closer to targets, and the complex environment within the branch can also cause targets to be easily lost. Furthermore, when identifying specific individuals, the method for obtaining target customer information typically uses facial recognition, which requires the customer's face to be very close to the robot's camera without any obstruction. In the one-on-one service scenario of branch robots, customers may move around frequently and cannot always face the robot's camera directly. This makes it difficult for most existing branch robots to fixate on a single target and continuously follow a specific customer to provide customized one-on-one service.
[0043] The target tracking method, apparatus, device, medium, and product provided in this application aim to solve the aforementioned technical problems of the prior art. The target tracking method, apparatus, device, medium, and product provided in this application acquire a first image sequence; the first image sequence includes image frame data of the target to be tracked in a time series; the first image sequence is input into an image encoder for encoding processing to obtain a first feature data set; the first feature data set and a motion-aware memory library are input into a memory-based time series attention mechanism module for feature extraction to obtain a second feature data set; wherein, the motion-aware memory library includes feature information corresponding to image frame data at historical time points; the prompt information corresponding to the first image sequence is input into a prompt encoder for encoding processing to obtain a prompt feature data set; the second feature data set and the prompt feature data set are input into a mask decoder for decoding processing to obtain a first decoded image set; and the target to be tracked is determined based on the first decoded image set. The solution of this application, by encoding the image sequence of the target to be tracked in a time series and combining the feature data of image frames already identified at historical time points during the tracking and recognition process, improves the temporal correlation of target tracking, thereby improving the accuracy of target tracking.
[0044] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0045] Example 1
[0046] Figure 1 An example is shown illustrating the flowchart of a target tracking method. Figure 1 ;like Figure 1 As shown, the method includes:
[0047] Step 101: Obtain the first image sequence; the first image sequence includes image frame data of the target to be tracked in a time series.
[0048] Step 102: Input the first image sequence into the image encoder for encoding processing to obtain the first feature data set;
[0049] Step 103: Input the first feature data set and the motion-aware memory library into the memory time series attention mechanism module to perform feature extraction and obtain the second feature data set; the motion-aware memory library includes feature information corresponding to image frame data at historical time points;
[0050] Step 104: Input the prompt information corresponding to the first image sequence into the prompt encoder for encoding processing to obtain the prompt feature data set;
[0051] Step 105: Input the second feature data set and the prompt feature data set into the mask decoder for decoding processing to obtain the first decoded image set; and determine the target to be tracked based on the first decoded image set.
[0052] Optionally, a first image sequence is first acquired; the first image sequence includes image frame data of the target to be tracked in a time series. The solution of this application can be applied to target recognition and tracking scenarios for bank branch robots and some service industry lobby robots; this application does not limit the application scenarios for target recognition and tracking. In one example, in a bank branch lobby scenario, at a certain time on a busy day, a customer arrives without an appointment and takes a number from a service terminal such as a ticket dispenser. The terminal uses facial recognition and bank card information to initially identify the customer's user information. Then, the branch service robot, combining the initially identified information with the busy branch environment, identifies and locks onto the customer taking the number, enabling the bank branch robot to navigate through crowds and obstacles to reach the target without losing sight of it. After further confirming the customer's information using dialogue or other means, the robot provides the necessary services or guides the customer to a VIP room, self-service counter, or other locations. For example, the first image sequence contains image frame data of the target to be tracked in a time series; multiple frames of image data can be dynamically acquired at each acquisition time. When dynamic tracking of a target is required, it is necessary to acquire the corresponding image frame data in real time and sort the image frame data in real time to improve the accuracy of dynamic target tracking. The target can be a moving person or object, etc.
[0053] The acquired first image sequence is input into an image encoder for encoding processing to obtain a first feature data set. For example, the image encoder can use a convolutional or attention mechanism framework. The image encoder uses one image frame at a time and cross-references the memory information of the target object from the previous frame on the input timeline. An input image sequence, after convolution, becomes a multi-dimensional image feature data stream, which is the first feature data set. For example, the encoding of image frame data follows a first-in, first-out (FIFO) principle, encoding each image frame in chronological order.
[0054] The obtained first feature data set and motion-aware memory library are input into the memory time-series attention mechanism module for feature extraction, resulting in a second feature data set. The motion-aware memory library includes feature information corresponding to image frame data at historical moments; the motion-aware memory library also includes feature information corresponding to image frame data that has been identified before the current image frame data; the current frame image data is associated with the feature information corresponding to image frame data at historical moments in the memory-aware memory library, thereby improving the accuracy of feature extraction by the memory time-series attention mechanism module and further improving the accuracy of target tracking.
[0055] For example, the in-memory time-series attention mechanism module may include a transformer module that acquires learnable attention features between frames in the time series. This module can associate the feature information of the current frame with the feature information of past frames and cue information. n transformer modules can be used, the number depending on the number of frames retained in memory; for example, one image data frame corresponds to at least one transformer module. Self-attention is performed on each transformer module, followed by cross-attention between them. Finally, a multilayer perceptron outputs the attention features.
[0056] Simultaneously, the prompt information corresponding to each image frame data in the first image sequence is input into the prompt encoder for encoding processing to obtain a prompt feature data set. For example, each frame of image data contains corresponding prompt information, which is used to indicate the target location information contained in each frame. The prompt information can be a hint about the location and region of the target in the image frame, defining the range of the target object in the current image frame. It can be a point, a bounding box, or a target mask. After encoding each type of prompt, the embedding is summed with the output of the image encoder. The mask information is embedded using convolution and summed using frame embedding to output a prompt token.
[0057] Furthermore, the second feature data set and the cue feature data set are input into the mask decoder for decoding to obtain the first decoded image set; and the target to be tracked is determined based on the first decoded image set. For example, the input to the mask decoder is the memory image embedding features output by the memory time-series attention mechanism, i.e., the second feature data set, and the cue token generated by the cue encoder. This decoder can also use a transformer module to predict multiple mask outputs to generate the first decoded image set, and the target to be tracked is determined based on the first decoded image set.
[0058] In this example, a first image sequence is acquired, comprising image frame data of the target to be tracked in a time series. The first image sequence is input into an image encoder for encoding to obtain a first feature data set. The first feature data set and a motion-aware memory library are input into a time-series attention mechanism module for feature extraction to obtain a second feature data set. The motion-aware memory library includes feature information corresponding to image frame data from historical time periods. The cue information corresponding to the first image sequence is input into a cue encoder for encoding to obtain a cue feature data set. The second feature data set and the cue feature data set are input into a mask decoder for decoding to obtain a first decoded image set. The target to be tracked is determined based on the first decoded image set. This solution, by encoding the image sequence of the target to be tracked in a time series and combining it with feature data from previously identified image frames during the tracking and recognition process, improves the temporal correlation of target tracking, thereby enhancing the accuracy of target tracking.
[0059] The method also includes:
[0060] Obtain the second image sequence;
[0061] The second image sequence is preprocessed to generate the first image sequence; the preprocessing includes at least one of the following: image resolution correction and image style transfer.
[0062] Optionally, lighting conditions vary under different weather conditions and at different times of day. For bank branches, the background of images captured by cameras under different lighting conditions will vary with the lighting, affecting target recognition. Therefore, image preprocessing is needed on the image frame data under different lighting conditions in the acquired second image sequence to generate the first image sequence. Image preprocessing includes at least one of the following: image resolution correction and image style transfer. The first image sequence consists of image frame data of relatively good quality and consistent lighting conditions after preprocessing. The second image sequence consists of image frame data with different lighting conditions and different qualities. In practical applications, image style transfer is a technique that uses algorithms to generate new images by using two input images as content and style references, respectively. Currently, image style transfer methods mainly include optimization-based methods, feedforward network-based methods, and generative adversarial network (GAN)-based methods. For example, a GAN network can be used for style transfer. The network can be a recurrent generative adversarial network (CycleGAN). Figure 2 An exemplary illustration shows a schematic diagram of image style transfer based on a recurrent generative adversarial network; such as Figure 2 As shown, CycleGAN's training does not rely on a fully paired image set except for style. To limit the distortion of native images caused by overtraining of the image generator, CycleGAN uses a recurrent generation structure and optimizes the consistency loss instead of the constraints of traditional style transfer methods.
[0063] The CycleGAN network model contains two image generators, G:Y→X and F:X→Y, and two adversarial discriminators. and . G is encouraged to transform X into an output indistinguishable from the Y domain, while distinguishing between the image {x} and the transformed image {F(y)}. of Work is the opposite.
[0064] The overall loss function of CycleGAN can be expressed as:
[0065]
[0066] in, and Let G be the loss function for image generators G and F. It uses the traditional loss calculation method for GAN networks. The specific formula for the loss function of image generator G is:
[0067]
[0068] in, For maximum likelihood estimation, It is the distribution of two datasets. and This is an adversarial discriminator. The loss function of image generator F is the opposite of that of image generator G.
[0069] The cycle consistency loss function ensures that the images generated by G and F not only satisfy their respective discriminators but can also be applied to other images. The specific formula is as follows:
[0070]
[0071] in, Representing the image matrix The norm of .
[0072] For example, style transfer can transform images captured by a camera under different lighting conditions within a dot matrix into images captured under lighting conditions within a fixed time period, while simultaneously optimizing the sharpness of the original camera images. Specifically, the CycleGAN model can be optimized and trained by selecting a time of day with good lighting conditions (1 PM to 3 PM) and using a high-resolution camera to take a large number of photos at different locations within the dot matrix, without labeling, as a dataset. (i.e., target image style). A large number of photos were taken at different locations within the network using a dot-matrix robot camera under random lighting conditions. No annotation was required, and these photos were compiled into a dataset. (i.e., the style of the input image). By training the CycleGAN network using these two datasets, a style transfer model can be obtained that can transform in-dot images under arbitrary lighting conditions into in-dot images under better lighting conditions, and can also optimize resolution.
[0073] In this example, image preprocessing is performed on the acquired image sequence to be identified, which further improves the accuracy and efficiency of target tracking.
[0074] Optional, Figure 3 An example is shown illustrating the flowchart of a target tracking method. Figure 2 ;like Figure 3 As shown, the method also includes:
[0075] Step 201: Calculate the mask affinity score and mask object score based on the first decoded image set;
[0076] Step 202: Based on the mask affinity score and the mask object score, filter the first decoded image set to obtain the second decoded image set;
[0077] Step 203: Based on the second decoded image set, perform encoding processing using the motion modeling method to obtain the third feature data set;
[0078] Step 204: Generate a frame memory sequence based on the third feature data set, and store the frame memory sequence into the motion-aware memory library.
[0079] Optionally, mask affinity scores and mask object scores can be further calculated based on the first decoded image set generated by the mask decoder. Based on the calculated mask affinity scores and mask object scores, the first decoded image set is filtered to obtain a second decoded image set. Then, based on the second decoded image set, encoding processing is performed using a motion modeling method to obtain a third feature data set. Based on the third feature data set, a frame memory sequence is generated and stored in a motion-aware memory library. This makes the frame sequence stored in the memory-aware library more accurate, and the feature data of historical image frames associated with it during image recognition are more accurate, thus improving the accuracy of image tracking.
[0080] For example, mask affinity scoring The probability that a masked region belongs to an object can be scored using IoU (Intersection over Union).
[0081] Interchange of Unit (IoU) is a metric for measuring the accuracy of object detection on a specific dataset, and is widely used in object detection tasks that output bounding boxes. The calculation process involves calculating the overlap between the predicted bounding box and the ground truth (GT) bounding box, and then dividing this overlap by the union of the two regions, as follows:
[0082]
[0083] in, Indicates mask affinity rating. For the predicted bounding box region, The GT box area. Indicates the area of the region. The value of is generally between 0 and 1. The closer it is to 1, the more the predicted bounding box overlaps with the actual bounding box, which means the detection result is more accurate.
[0084] Object score of the target object in each frame image This indicates whether the object appears in the frame or is obscured by other objects. It can be set to 0 or 1. 1 indicates appearance, and 0 indicates absence.
[0085] The final mask output by the mask decoder It is based on output. The highest affinity score among the masks is selected.
[0086]
[0087] (in )
[0088] in, The masked decoded image of the output at each acquisition time. This represents the N masked decoded images corresponding to any acquisition time. The formula indicates that the masked decoded image with the highest mask affinity score is selected, which requires that the mask object score be greater than 0, meaning that the masked decoded image contains the target to be tracked.
[0089] In this example, by generating a corresponding third feature dataset based on the mask affinity score and mask object score corresponding to the first decoded image set, and combining it with the motion modeling method, the third feature dataset that meets the preset requirements is then stored in the motion-aware memory library, thereby improving the accuracy of target tracking results.
[0090] Based on the second set of decoded images, encoding processing is performed using a motion modeling method to obtain the third feature data.
[0091] Sets, including:
[0092] Based on the second decoded image set, and using a motion modeling method, a predicted decoded image set of the target to be tracked corresponding to the second decoded image set is obtained;
[0093] Calculate the motion modeling score between the second decoded image set and the predicted decoded image set;
[0094] Based on mask affinity score, mask object score, and motion modeling score, determine the set of images to be encoded;
[0095] The set of images to be encoded is processed to obtain the third feature data set.
[0096] Optionally, based on the second decoded image set and a motion modeling method, a predicted decoded image set of the target to be tracked corresponding to the second decoded image set can be obtained. Then, a motion modeling score is calculated between the second decoded image set and the predicted decoded image set. Next, based on the mask affinity score, mask object score, and motion modeling score, the image set to be encoded is determined. This image set is then encoded to obtain the third feature data set.
[0097] For example, after the mask decoder outputs the masked image, in order to obtain more motion feature information for tracking the human body, motion modeling can be added to the encoder. Various motion modeling methods can be selected, such as Kalman filter (KF), trajectory box motion prediction, etc. An example is a method based on the Kalman filter:
[0098] The state vector Defined as:
[0099]
[0100] in, , Indicates the center coordinates of the bounding box. and These represent its width and height, respectively, and its corresponding velocity is represented by dots. For each mask... The corresponding bounding box By calculating the minimum and maximum of the non-zero pixels in the mask and The coordinates are derived. The Kalman filter operates in a prediction-correction cycle, where state prediction... It is given by the following formula:
[0101]
[0102] It is a linear state transition matrix.
[0103] Cross-Union Ratio (KF-IoU) Score Calculate the prediction mask The intersection-over-union (IoU) ratio between the bounding boxes predicted by the Kalman filter and the bounding boxes is obtained. Then we choose a mask that maximizes the weighted sum of the KF-IoU score and the original affinity score:
[0104]
[0105] in, Indicates weight, This indicates the mask affinity score.
[0106] Use the following formula to update:
[0107]
[0108] in These are the measured values, i.e., the bounding boxes derived from the mask we choose, used for updating. It is Kalman gain. It is the observation matrix. Furthermore, to ensure the robustness of motion modeling after the target object reappears or when the mask quality is poor for certain time periods, and to maintain a stable motion state, it is only in the past... Motion model updates are only considered when the tracked object is successfully updated within a frame.
[0109] Next, a frame-memory sequence (motion-aware memory library) is constructed using mask affinity scoring, object presence scoring, and motion scoring to store the motion features of moving objects. Motion features of moving objects are stored if and only if all three scores reach their corresponding thresholds (e.g., ...). The selected frame is placed into the memory sequence. The process iterates backward from the current frame and repeats the verification. We select based on the scoring function described above. One memory, and obtain a motion-aware memory library. :
[0110]
[0111] in, This is the maximum number of frames to be traced. Motion-aware memory library. It is then passed through the memory attention layer and then to the mask decoder. Perform mask decoding at the current timestamp.
[0112] In this example, by combining motion modeling methods to determine the generated third feature data, the accuracy of the generated third feature data is improved.
[0113] The method also includes:
[0114] The first image sequence is filtered to obtain the first image data set;
[0115] Based on the first image dataset, and using the inter-frame difference method and a collaborative training model, the first prediction result is output;
[0116] Based on the first prediction result and the first decoded image set, calculate the intersection-union ratio (IUGR) score; and determine the first image mask data set based on the IUGR score.
[0117] The first image dataset and the first image mask dataset are combined as training datasets to train the target tracking model. The target tracking model includes an image encoder, a memory time-series attention mechanism module, a cue encoder, and a mask decoder.
[0118] Optionally, the first image sequence can be filtered to obtain a first image dataset. Based on the first image dataset, a first prediction result is output using the inter-frame difference method and a co-trained model.
[0119] Based on the first prediction result and the first decoded image set, the cross-union ratio score is calculated; and based on the cross-union ratio score, the first image mask data set is determined.
[0120] The first image dataset and the first image mask dataset are combined as training datasets to train the target tracking model; the target tracking model includes an image encoder, a memory time-series attention mechanism module, a cue encoder, and a mask decoder.
[0121] For example, the prediction results of the traditional inter-frame difference method are combined with the results of the co-trained model for judgment. The core of the inter-frame difference method is to perform difference operations on two, three, or more temporally consecutive frames to obtain the moving region. First, the difference of pixel values (usually grayscale values) between adjacent frames is calculated. Then, similar to background subtraction, a reference threshold is set, and each pixel is binarized. The grayscale value of 255 is the foreground, and the grayscale value of 0 is the background. Since the position of the target in adjacent frames differs greatly, subtracting two frames does not yield a complete picture of the moving target. Therefore, based on the two-frame difference method, three-frame difference methods, five-frame difference methods, etc., have been proposed to improve the target bounding box. Figure 4 An exemplary schematic diagram of an inter-frame difference method is shown; as follows: Figure 4 As shown, Figure 4 Taking a person as the target to be identified as an example, the position of the human body in the first frame image is subtracted from the position where the human body moves in the next frame, and then binarized to obtain the position of the target box.
[0122] The collaborative training model can select any single-frame object detection model and output its target bounding box.
[0123] Finally, the IoU score is calculated by comparing the results of the inter-frame difference method, the output of the co-training model, and the prediction results corresponding to the scheme in this application. A threshold is set, and if the result exceeds a certain threshold, the prediction result is used as the mask label of the semi-supervised training dataset to train the target tracking model.
[0124] In this example, by filtering the image data and result data, a training dataset is obtained, and then the target tracking model is trained, which improves the performance of the model and thus improves the accuracy of the target tracking results.
[0125] Optionally, the first image sequence is filtered to obtain a first image data set, including:
[0126] Calculate the structural similarity index between the images in the first image sequence and the target image, and retain the set of images with a structural similarity index greater than a preset threshold as the first image data set.
[0127] Optionally, for judging the style transfer results: the structural similarity assessment coefficient (MSSIM) is used to determine whether the image is stylistically similar to the target image.
[0128] MSSIM is the mean of the Structural Similarity (SSIM) coefficients, which assess the structural similarity between two sets of images. Traditional MSE and PSNR evaluation methods only consider the two pixel values at the current location, ignoring pixels at any other location and neglecting the visual features of the image. When calculating the difference between two images at each location, SSIM takes pixels from a region in each image, considering the local structural information of the images.
[0129] SSIM uses brightness similarity Contrast similarity and structural similarity The similarity between the two images is comprehensively evaluated, and the calculation methods are as follows:
[0130]
[0131]
[0132]
[0133]
[0134]
[0135] It is an image The average value of the pixels, image The variance of the pixels, It is an image and The covariance of pixels, To maintain a stable constant, division by zero can be avoided. The range of pixel values represents the B-bit image. Value For uint8 data, the maximum pixel value is 255; for floating-point data, the maximum pixel value is 1. Generally speaking... If MSSIM is higher than a manually set threshold, it is determined that it can be used for training.
[0136] In this example, the training data is filtered by calculating the structural similarity index, which improves the quality of the training data and further enhances the performance of the target tracking model.
[0137] In one example Figure 5 This is a schematic diagram of a target tracking process as an example of this application; such as Figure 5 As shown, this application can be divided into a style transfer module. The original camera input image set under different lighting conditions in a time-series sequence is processed by a style transfer network composed of a generator network (Genc) and a decoder network (Gdenc) to obtain a style-transferred image set. The style-transferred image set is then sequentially input into an image encoder and a memory-based time-series attention module. Simultaneously, the memory-based time-series attention module also associates feature information from historical image data frames obtained after feature extraction in a motion-aware memory library, and outputs this information to a mask decoder. Simultaneously, the style-transferred image set and corresponding cue information are also input into an image sparse feature encoder, which outputs corresponding boxes, points, and masks. The mask decoder combines the outputs of the memory-based time-series attention module and the mask decoder to decode the images, obtaining a decoded image set. Based on this decoded image set, the target to be tracked can be determined. Furthermore, based on the decoded image set, mask affinity score and mask object score are calculated, and combined with the object motion mathematical model, the decoded images that meet the score requirements are finally selected. Then, they are encoded by the mask encoder, and the encoded feature data is stored in the frame memory sequence and stored in the motion-aware memory library. The sequence stored in the motion-aware memory library is a first-in-first-out sequence, where green represents high-quality data that meets the screening requirements, yellow is next, and red is low-quality data.
[0138] Furthermore, the entire target tracking model can also undergo a semi-supervised correction module, which automatically corrects the model during the target tracking and recognition process. The style-transferred image set is filtered by structural similarity index, and the filtered image set is saved to the semi-supervised training set library. Then, the inter-frame difference detection method and co-training model are performed based on the image set in the semi-supervised training set library to predict the results. The predicted results are compared with the mask image set output by the aforementioned mask decoder to perform IOU filtering to obtain the detection results. The target tracking model of this application is dynamically corrected based on the detection results and the corresponding image set in the semi-supervised training set library.
[0139] The target tracking method provided in this embodiment acquires a first image sequence, which includes image frame data of the target to be tracked in a time series. The first image sequence is input into an image encoder for encoding to obtain a first feature data set. The first feature data set and a motion-aware memory library are input into a time-series attention mechanism module for feature extraction to obtain a second feature data set. The motion-aware memory library includes feature information corresponding to image frame data from historical time periods. The cue information corresponding to the first image sequence is input into a cue encoder for encoding to obtain a cue feature data set. The second feature data set and the cue feature data set are input into a mask decoder for decoding to obtain a first decoded image set. The target to be tracked is determined based on the first decoded image set. This solution, by encoding the image sequence of the target to be tracked in a time series and combining the feature data of image frames already identified in historical time periods during the tracking and recognition process, improves the temporal correlation of target tracking, thereby improving the accuracy of target tracking.
[0140] Example 2
[0141] Figure 6 An exemplary schematic diagram of a target tracking device is shown; the device includes:
[0142] Acquisition module 21 is used to acquire a first image sequence; the first image sequence includes image frame data of the target to be tracked in a time series.
[0143] The first encoding module 22 is used to input the first image sequence into the image encoder for encoding processing to obtain the first feature data set;
[0144] Extraction module 23 is used to input the first feature data set and the motion-aware memory library into the memory time series attention mechanism module for feature extraction to obtain the second feature data set; the motion-aware memory library includes feature information corresponding to image frame data at historical time points;
[0145] The second encoding module 24 is used to input the prompt information corresponding to the first image sequence into the prompt encoder for encoding processing to obtain a prompt feature data set.
[0146] The decoding module 25 is used to input the second feature data set and the prompt feature data set into the mask decoder for decoding processing to obtain the first decoded image set; and to determine the target to be tracked based on the first decoded image set.
[0147] The target tracking device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0148] Example 3
[0149] Figure 7 The diagram above illustrates the structure of an electronic device, which includes:
[0150] The device includes a processor 291 and a memory 292; it may also include a communication interface 293 and a bus 294. The processor 291, memory 292, and communication interface 293 can communicate with each other via the bus 294. The communication interface 293 can be used for information transmission. The processor 291 can invoke logical instructions stored in the memory 292 to execute the methods described in the example above.
[0151] Furthermore, the logic instructions in the aforementioned memory 292 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0152] The memory 292, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this application. The processor 291 executes functional applications and data processing by running the software programs, instructions, and modules stored in the memory 292, that is, it implements the methods in the above method examples.
[0153] The memory 292 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 292 may include high-speed random access memory and may also include non-volatile memory.
[0154] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method in any of the embodiments.
[0155] This application also provides a computer program product, including a computer program that, when executed by a processor, is used to implement the method in any of the embodiments.
[0156] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0157] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0158] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.
[0159] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.
[0160] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.
[0161] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0162] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0163] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0164] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A target tracking method characterized by, The method comprises: obtaining a first image sequence; the first image sequence comprises image frame data of a target to be tracked in a time sequence; inputting the first image sequence into an image encoder for encoding processing to obtain a first feature data set; inputting the first feature data set and a motion-aware memory library into a memory time sequence attention mechanism module for feature extraction to obtain a second feature data set; the motion-aware memory library comprises feature information corresponding to image frame data at a historical time; inputting prompt information corresponding to the first image sequence into a prompt encoder for encoding processing to obtain a prompt feature data set; inputting the second feature data set and the prompt feature data set into a mask decoder for decoding processing to obtain a first decoded image set; and determining the target to be tracked according to the first decoded image set.
2. The method of claim 1, wherein, The method further comprises: obtaining a second image sequence; preprocessing the second image sequence to generate the first image sequence; the preprocessing comprises at least one of the following: image resolution correction and image style transfer.
3. The method of claim 1, wherein, The method further comprises: calculating a mask affinity score and a mask object score according to the first decoded image set; screening the first decoded image set according to the mask affinity score and the mask object score to obtain a second decoded image set; performing encoding processing on the second decoded image set based on a motion modeling method to obtain a third feature data set; generating a frame memory sequence based on the third feature data set and storing the frame memory sequence in the motion-aware memory library.
4. The method of claim 3, wherein, The encoding processing based on the motion modeling method on the second decoded image set to obtain the third feature data set comprises: obtaining a predicted decoded image set of the target to be tracked corresponding to the second decoded image set based on the motion modeling method; calculating a motion modeling score of the second decoded image set and the predicted decoded image set; determining a to-be-encoded image set according to the mask affinity score, the mask object score, and the motion modeling score; performing encoding processing on the to-be-encoded image set to obtain the third feature data set.
5. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: screening the first image sequence to obtain a first image data set; outputting a first prediction result based on an inter-frame difference method and a co-training model according to the first image data set; calculating a intersection over union score according to the first prediction result and the first decoded image set; and determining a first image mask data set according to the intersection over union score; training a target tracking model using the first image data set and the first image mask data set as training data sets; the target tracking model comprises the image encoder, the memory time sequence attention mechanism module, the prompt encoder, and the mask decoder.
6. The method of claim 5, wherein, The screening of the first image sequence to obtain the first image data set comprises: Calculate a structural similarity index between images contained in the first image sequence and a target image, retain a set of images with a structural similarity index greater than a preset threshold as the first image data set.
7. A target tracking device, characterized by, The method comprises the steps of: An acquisition module is configured to acquire a first image sequence. The first image sequence comprises image frame data of a target to be tracked in a time sequence. A first encoding module is configured to input the first image sequence into an image encoder to perform encoding processing and obtain a first feature data set. An extraction module is configured to input the first feature data set and a motion-aware memory library into a memory time sequence attention mechanism module to perform feature extraction and obtain a second feature data set. The motion-aware memory library comprises feature information corresponding to image frame data at a historical time. A second encoding module is configured to input prompt information corresponding to the first image sequence into a prompt encoder to perform encoding processing and obtain a prompt feature data set. A decoding module is configured to input the second feature data set and the prompt feature data set into a mask decoder to perform decoding processing and obtain a first decoded image set, and determine the target to be tracked according to the first decoded image set.
8. An electronic device, comprising: The method comprises the steps of: A processor and a memory connected to the processor in communication; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by the processor to implement the method according to any one of claims 1 to 6.
10. A computer program product, characterised in that, The computer program is executed by the processor to implement the method according to any one of claims 1 to 6.