Target tracking method and device

Through multi-channel feature extraction and object detection, bounding box and appearance features are extracted, and the problem of inaccurate target detection and tracking in the prior art is solved, and high-precision target tracking is achieved in complex scenarios.

CN119941795AActive Publication Date: 2025-05-06AEROSPACE INFORMATION RES INST CAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510143512.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-06
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The prior art in the non-fixed lens scene or high-speed motion scenes does not accurately detect and track the target, and does not fully utilize the timing information of the video.

Method used

Through multi-channel feature extraction, the feature maps corresponding to multiple channels are obtained, and the target detection is performed to extract the bounding box features and appearance features of the target. Based on these features, the matching result of the detection target and the tracking trajectory is determined.

Benefits of technology

Improve the accuracy of target tracking, especially in scenes where the lens is not fixed or high-speed motion scenarios, and enhance the detection and tracking capabilities of targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941795A_ABST
    Figure CN119941795A_ABST
Patent Text Reader

Abstract

The invention provides a target tracking method and device which can be applied to the field of computer vision. The method comprises the steps that multi-channel feature extraction is carried out on a to-be-detected image of a current frame, feature maps corresponding to multiple channels are obtained, and the to-be-detected image comprises a detection target; performing target detection on the feature maps corresponding to the plurality of channels to obtain a target bounding box feature of the detection target and a target appearance feature of the detection target; and based on the target bounding box feature of the detection target, the target appearance feature of the detection target and the trajectory feature of the tracking trajectory, determining the matching result trajectory bounding box feature and the trajectory appearance feature of the detection target and the tracking trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision, and in particular to a target tracking method, device, equipment, medium and program product. Background Art

[0002] Object tracking has always been an important issue in the field of computer vision. Object tracking is defined as continuously tracking each object of interest in a video sequence and generating a trajectory. The key is to accurately locate and track the object from consecutive image frames while maintaining the identity consistency of the object. In related technologies, the detection box information of the object is often extracted to associate the data between the object and the trajectory, thereby achieving object tracking.

[0003] In the process of realizing the concept of the present disclosure, the inventors found that there are at least the following problems in the related art: less information about the target is extracted during the tracking process, and in scenes such as scenes with unstable shots or high-speed motion, there is a problem of insufficient accuracy in target detection and tracking. Summary of the invention

[0004] In view of the above problems, the present disclosure provides a target tracking method, apparatus, device, medium and program product.

[0005] According to a first aspect of the present disclosure, there is provided a target tracking method, comprising: performing multi-channel feature extraction on an image to be detected in a current frame to obtain feature maps corresponding to each of the multiple channels, wherein the image to be detected includes a detection target; performing target detection on the feature maps corresponding to each of the multiple channels to obtain target bounding box features of the detection target and target appearance features of the detection target; determining a matching result between the detection target and the tracking trajectory based on the target bounding box features of the detection target, the target appearance features of the detection target and the trajectory features of the tracking trajectory, wherein the tracking trajectory is determined based on a historical detection target that is continuously tracked in at least one historical frame, and the trajectory features include trajectory bounding box features determined by historical bounding box features of the historical detection target and trajectory appearance features determined by historical appearance features.

[0006] According to an embodiment of the present disclosure, target detection is performed on feature maps corresponding to each of multiple channels to obtain target bounding box features of the detected target and target appearance features of the detected target, including: performing channel-unified processing on the feature maps corresponding to each of the multiple channels to obtain multiple target feature maps with the same number of channels; inputting the multiple target feature maps into a first feature extraction model to obtain target bounding box features of the detected target; and inputting the multiple target feature maps into a second feature extraction model to obtain target appearance features of the detected target.

[0007] According to an embodiment of the present disclosure, multiple target feature maps are input into a second feature extraction model to obtain target appearance features of the detection target, including: inputting multiple target feature maps into a first appearance extraction sub-model to obtain initial appearance features of the detection target; performing time series information analysis on the initial appearance features of the detection target and historical appearance features of at least one neighboring historical detection target to obtain target appearance features of the detection target, wherein the neighboring historical detection target is a historical detection target included in the image to be detected of M historical frames closest to the current frame, and M is an integer greater than or equal to 1.

[0008] According to an embodiment of the present disclosure, multiple target feature maps are input into a second feature extraction model to obtain target appearance features of the detection target, including: inputting multiple target feature maps into a second appearance extraction sub-model to obtain target appearance features of the detection target, wherein the second appearance extraction sub-model includes a feature extraction layer and a mapping layer, and the feature extraction layer includes at least six convolution sub-layers.

[0009] According to an embodiment of the present disclosure, the target bounding box feature includes: a target bounding box confidence and a target bounding box position; the track bounding box feature includes: a track bounding box position and a track speed and a track acceleration corresponding to the track bounding box position; wherein, based on the target bounding box feature of the detected target, the target appearance feature of the detected target and the track features of the tracking track, determining the matching result between the detected target and the tracking track includes: when it is determined that the target bounding box confidence is greater than a preset confidence, calculating the distance between the target appearance feature of the detected target and the track appearance feature to obtain the appearance distance; based on the track bounding box position, the track speed and the track acceleration corresponding to the track bounding box position, predicting the predicted bounding box position of the tracking track in the current frame; calculating the distance between the predicted bounding box position and the target bounding box position to obtain the bounding box distance; based on the appearance distance and the bounding box distance, obtaining a matching cost matrix; based on the matching cost matrix, determining the matching result between the detected target and the tracking track.

[0010] According to an embodiment of the present disclosure, the target bounding box feature includes: target bounding box confidence, target bounding box position and target bounding box size information; the track bounding box feature includes: track bounding box position, track speed corresponding to the track bounding box position, track acceleration and track bounding box size information; wherein, based on the target bounding box feature of the detected target, the target appearance feature of the detected target and the track features of the tracking track, determining the matching result between the detected target and the tracking track includes: predicting the predicted bounding box position and predicted bounding box size information of the tracking track in the current frame based on the track bounding box position, the track speed corresponding to the track bounding box position, the track acceleration and the track bounding box size information when determining that the target bounding box confidence is less than or equal to the preset confidence; determining a predicted sub-image based on the predicted bounding box position and the predicted bounding box size information; determining a current sub-image based on the target bounding box position and the target bounding box size information; determining a distance between the predicted bounding box feature and the target bounding box feature based on a pixel set included in the predicted sub-image and a pixel set included in the current sub-image; and determining the matching result between the detected target and the tracking track based on the distance.

[0011] According to an embodiment of the present disclosure, the above method also includes: when it is determined that the matching result represents that the detection target matches the tracking trajectory, replacing the trajectory bounding box feature in the trajectory feature with the target bounding box feature; obtaining an updated appearance feature based on the target appearance feature, the trajectory appearance feature and the preset appearance weight value; and replacing the trajectory appearance feature in the trajectory feature with the updated appearance feature.

[0012] According to an embodiment of the present disclosure, the above method also includes: when it is determined that the matching result indicates that the detected target does not match the tracking trajectory and the target bounding box confidence of the detected target is greater than a preset confidence, determining the trajectory features of the new tracking trajectory based on the target appearance features and the target bounding box features; using the current frame as the lost frame of the tracking trajectory, so that when the lost frames of the tracking trajectory are continuous frames and the number of lost frames is greater than a preset frame number threshold, the tracking trajectory is deleted.

[0013] According to an embodiment of the present disclosure, the first feature extraction model and the second feature extraction model are trained in the following manner: multiple sample target feature maps of the sample image are input into the initial first feature extraction model to obtain sample bounding box features of the sample targets included in the sample image; based on the sample bounding box features and the true bounding box features, a first loss value is obtained; multiple sample target feature maps are input into the initial second feature extraction model to obtain sample appearance features of the sample targets; based on the predicted identification and the true identification of the sample appearance features, a second loss value is obtained; based on the first loss value, a first weight value corresponding to the first loss value, the second loss value, and the second weight value corresponding to the second loss value, a target loss value is determined; based on the target loss value, the initial first feature extraction model and the initial second feature extraction model are trained to obtain the first feature extraction model and the second feature extraction model.

[0014] A second aspect of the present disclosure provides a target tracking device, characterized in that the device includes: an extraction module, used to perform multi-channel feature extraction on an image to be detected in a current frame to obtain feature maps corresponding to each of the multiple channels, wherein the image to be detected includes a detection target; a detection module, used to perform target detection on the feature maps corresponding to each of the multiple channels to obtain target bounding box features of the detection target and target appearance features of the detection target; a matching module, used to determine a matching result between the detection target and the tracking trajectory based on the target bounding box features of the detection target, the target appearance features of the detection target and the trajectory features of the tracking trajectory, wherein the tracking trajectory is determined based on a historical detection target included in the image to be detected of at least one historical frame, and the trajectory features include target boundary features determined by historical bounding box features of the historical detection targets and trajectory appearance features determined by historical appearance features.

[0015] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0016] The fourth aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the above computer program or instructions are executed by a processor.

[0017] The fifth aspect of the present disclosure further provides a computer program product, including a computer program or instructions, which implement the steps of the above method when the above computer program or instructions are executed by a processor.

[0018] According to the target tracking method disclosed in the present invention, by adopting a multi-channel feature extraction method, feature extraction is performed on the image to be detected from multiple scales, and feature maps corresponding to each of the multiple channels are obtained, thereby realizing the extraction of details and global features of the image to be detected. When performing target detection, not only the target bounding box features of the detected target are extracted, but also the target appearance features of the detected target are extracted, so that in subsequent matching, the tracking trajectory is matched with the detected target based on the information of the two angles of appearance features and detection frame features. Since the feature maps corresponding to each of the multiple channels are used for target detection, and the matching result of the tracking trajectory and the detected target is determined based on the two feature information of the target bounding box features and the target appearance features. Therefore, at least part of the technical problem of inaccurate target detection and tracking existing in the related art is solved, and the technical effect of improving the accuracy of target tracking is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0020] Figure 1 The application scenario diagram of the target tracking method, apparatus, device, medium and program product according to the embodiments of the present disclosure is schematically shown;

[0021] Figure 2 A flowchart of a target tracking method according to an embodiment of the present disclosure is schematically shown;

[0022] Figure 3 A model structure diagram for determining target bounding box features and target appearance features of a detected target according to an embodiment of the present disclosure is schematically shown;

[0023] Figure 4 A schematic diagram of a model structure for determining target appearance features of a detection target according to another embodiment of the present disclosure is shown;

[0024] Figure 5 A flowchart of a target tracking method according to another embodiment of the present disclosure is schematically shown;

[0025] Figure 6 A schematic diagram of a target tracking device according to an embodiment of the present disclosure is shown; and

[0026] Figure 7 A block diagram of an electronic device suitable for implementing a target tracking method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0027] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0028] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.

[0029] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0030] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0031] It should be noted that the target tracking method and device disclosed in the present invention can be used in the field of computer vision, and can also be used in any field other than the field of computer vision, such as the field of artificial intelligence technology. The present disclosure does not limit the application field of the target tracking method and device.

[0032] In the technical solution of the present disclosure, the user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0033] In the scenario of using personal information for automated decision-making, the methods, devices, and systems provided by the embodiments of the present disclosure provide users with corresponding operation portals for users to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision-making" here refers to the activity of automatically analyzing and evaluating a person's behavioral habits, interests and hobbies, or economic, health, credit status, etc. through computer programs, and making decisions. The expression "expert decision-making" here refers to the activity of making decisions by people who specialize in a certain field, have specialized experience, knowledge and skills, and have reached a certain level of professionalism.

[0034] During the research process, it was found that the target tracking algorithm in the related technology only uses the motion information prediction of the trajectory and the detection frame information of the target to associate the data between the target and the trajectory. In scenes with unstable shots and high-speed motion, the algorithm is not accurate enough in target detection and tracking, which will reduce the tracking performance to a certain extent. At the same time, only feature information is extracted from a single frame of the picture, and the timing information of the video is not fully utilized, resulting in poor tracking effect.

[0035] An embodiment of the present disclosure provides a target tracking method, comprising: performing multi-channel feature extraction on an image to be detected in a current frame to obtain feature maps corresponding to each of the multiple channels, wherein the image to be detected includes a detection target; performing target detection on the feature maps corresponding to each of the multiple channels to obtain target bounding box features of the detection target and target appearance features of the detection target; determining a matching result between the detection target and a tracking trajectory based on the target bounding box features of the detection target, the target appearance features of the detection target, and the trajectory features of the tracking trajectory, wherein the tracking trajectory is determined based on a historical detection target included in the image to be detected of at least one historical frame, and the trajectory features include target boundary features determined by historical bounding box features of the historical detection targets and trajectory appearance features determined by historical appearance features.

[0036] Figure 1 The application scenario diagram of the target tracking method, apparatus, device, medium and program product according to the embodiments of the present disclosure is schematically shown.

[0037] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.

[0038] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).

[0039] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0040] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0041] It should be noted that the target tracking method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the target tracking device provided in the embodiment of the present disclosure can generally be set in the server 105. The target tracking method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the target tracking device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0042] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0043] The following will be based on Figure 1 The scene described by Figure 2~Figure 5 The target tracking method of the disclosed embodiment is described in detail.

[0044] Figure 2 The flowchart of the target tracking method according to the embodiment of the present disclosure is schematically shown.

[0045] like Figure 2 As shown, the method includes operations S210 to S230.

[0046] In operation S210, multi-channel feature extraction is performed on the image to be detected in the current frame to obtain feature maps corresponding to each of the multiple channels, wherein the image to be detected includes a detection target.

[0047] In operation S220, object detection is performed on the feature maps corresponding to each of the multiple channels to obtain an object bounding box feature of the detected object and an object appearance feature of the detected object.

[0048] In operation S230, a matching result between the detection target and the tracking trajectory is determined based on the target bounding box features of the detection target, the target appearance features of the detection target, and the trajectory features of the tracking trajectory, wherein the tracking trajectory is determined based on a historical detection target that is continuously tracked in at least one historical frame, and the trajectory features include a trajectory bounding box feature determined by a historical bounding box feature of the historical detection target and a trajectory appearance feature, a trajectory bounding box feature, and a trajectory appearance feature determined by a historical appearance feature.

[0049] According to an embodiment of the present disclosure, the current frame may be an image frame to be detected in the video to be detected and being processed by the target tracking method. The image frame before the current frame is a historical frame relative to the current frame.

[0050] According to the embodiments of the present disclosure, before processing the video to be detected, the video to be detected can be split into continuous image frames to be detected, and the frame number of each image to be detected in the video can be marked. The original image frame resolution can be adjusted to 800 pixels × 1440 pixels through a preprocessing algorithm, and then the RGB mean of all image frames of the video can be subtracted from all image frames.

[0051] According to an embodiment of the present disclosure, each image to be detected in the video to be detected may be processed by the target tracking method in sequence according to the marked frame number.

[0052] According to the embodiments of the present disclosure, the method for performing multi-channel feature extraction on the image to be detected is not limited, and it can be implemented through the DarkNET53 network, or through FPN (Feature Pyramid Networks), ResNet (Residual Networks), etc.

[0053] According to an embodiment of the present disclosure, the feature maps corresponding to each of the multiple channels may have different resolutions in addition to the number of channels. For example, the above-mentioned feature maps are 256×100×180, 512×50×90 and 1024×25×45 high-resolution feature maps, respectively, where 256, 512, and 1024 are the number of channels, respectively.

[0054] According to an embodiment of the present disclosure, the detection target may be one or more, and the tracking track may also be one or more.

[0055] According to an embodiment of the present disclosure, when performing target detection on feature maps corresponding to multiple channels, target bounding box feature extraction and target appearance feature extraction can be performed on the multiple feature maps respectively, so as to obtain target bounding box features and target appearance features of all detection targets included in the image to be detected.

[0056] According to an embodiment of the present disclosure, the tracking trajectory may be a detection target that has appeared in an image to be detected in at least one historical frame.

[0057] According to an embodiment of the present disclosure, the trajectory features of the tracking trajectory may include trajectory bounding box features and trajectory appearance features, wherein, when there is a historical detection target that matches the tracking trajectory in the image to be detected of the historical frame closest to the current frame, the trajectory bounding box features are the historical bounding box features of the historical detection target; if there is no historical detection target that matches the tracking trajectory in the image to be detected of the historical frame closest to the current frame, but there is a detection target in the image to be detected of other historical frames that is closest to the current frame, the trajectory bounding box features are the historical bounding boxes of the target to be detected in the image to be detected of other historical frames that are closest to the current frame. The trajectory appearance features are similar, that is, the trajectory appearance features and the trajectory bounding box features are both obtained from the relevant features of the historical detection target of the historical frame that most recently appears in the tracking trajectory.

[0058] According to an embodiment of the present disclosure, by performing feature matching on the target bounding box features of the detection target, the target appearance features of the detection target and the trajectory features of the tracking trajectory determined in the image to be detected in the historical frame, it is determined whether the detection target matches the tracking trajectory, thereby achieving continuous tracking and identification of the tracking trajectory.

[0059] According to the target tracking method disclosed in the present invention, by adopting a multi-channel feature extraction method, feature extraction is performed on the image to be detected from multiple scales, and feature maps corresponding to each of the multiple channels are obtained, thereby realizing the extraction of details and global features of the image to be detected. When performing target detection, not only the target bounding box features of the detected target are extracted, but also the target appearance features of the detected target are extracted, so that in subsequent matching, the tracking trajectory is matched with the detected target based on the information of the two angles of appearance features and detection frame features. Since the feature maps corresponding to each of the multiple channels are used for target detection, and the matching results of the tracking trajectory and the detected target are determined based on the two feature information of bounding box features and appearance features. Therefore, the technical problem of insufficient precision in target detection and tracking existing in the related art is at least partially solved, and the technical effect of improving the accuracy of target tracking is achieved.

[0060] According to an embodiment of the present disclosure, performing target detection on feature maps corresponding to each of a plurality of channels to obtain target bounding box features of the detected target and target appearance features of the detected target may include the following steps.

[0061] The feature maps corresponding to each of the multiple channels are uniformly processed to obtain multiple target feature maps with the same number of channels; the multiple target feature maps are input into the first feature extraction model to obtain the target bounding box features of the detected target; the multiple target feature maps are input into the second feature extraction model to obtain the target appearance features of the detected target.

[0062] According to an embodiment of the present disclosure, the number of channels of feature maps corresponding to multiple channels can be unified. For example, the number of channels can be unified from 256, 512, and 1024 to 256 through 1×1 convolution.

[0063] According to the embodiments of the present disclosure, there is no limitation on the implementation methods of the first feature extraction model and the second feature extraction model, and both can be implemented through neural networks, such as convolutional neural networks.

[0064] According to an embodiment of the present disclosure, the target bounding box features may include: target bounding box size information, coordinates of the target bounding box center in the image, ie, the bounding box position, target bounding box confidence, and the like.

[0065] According to an embodiment of the present disclosure, inputting a plurality of target feature maps into a second feature extraction model to obtain target appearance features of the detection target may include the following steps.

[0066] Input multiple target feature maps into the first appearance extraction sub-model to obtain the initial appearance features of the detection target; perform time series information analysis on the initial appearance features of the detection target and the historical appearance features of at least one neighboring historical detection target to obtain the target appearance features of the detection target, wherein the neighboring historical detection target is a historical detection target included in the image to be detected of the M historical frames closest to the current frame, and M is an integer greater than or equal to 1.

[0067] According to an embodiment of the present disclosure, the first feature extraction model may include a first appearance extraction sub-model and a temporal processing sub-model.

[0068] According to an embodiment of the present disclosure, the timing processing sub-model is used to perform timing information analysis on the initial appearance features of the input detection target and the historical appearance features of at least one adjacent historical detection target to obtain the target appearance features of the final detection target, so that the target appearance features of the final detection target contain timing information.

[0069] According to an embodiment of the present disclosure, the time series processing sub-model can be implemented by ConvLSTM (Convolutional LongShort Term Memory) neurons, which are a variant of LSTM (Long Short Term Memory) neurons. LSTM is a special recursive neural network that can analyze input data using time series to introduce time series information. ConvLSTM optimizes the neuron structure on LSTM, achieving the same effect while reducing the number of parameters, and speeding up the training speed.

[0070] According to an embodiment of the present disclosure, if the current frame is the first frame of the input video, the initial appearance features of all detected targets in the current frame are used as target appearance features of the detected targets.

[0071] According to an embodiment of the present disclosure, the specific value of M can be determined according to the frame number of the current frame. For example, if the current frame is the third frame, that is, there are two historical frames before the current frame, then M can be 2 or 1.

[0072] According to the embodiments of the present disclosure, when extracting the target appearance features of the detection target, time series information is introduced, so that the target appearance features of the detection target are more accurate, the accuracy of the detection target is improved, and the tracking performance is further improved.

[0073] Figure 3 A model structure diagram for determining target bounding box features of a detected target and target appearance features of a detected target according to an embodiment of the present disclosure is schematically shown.

[0074] like Figure 3 As shown, the image to be detected 301 of the current frame is input into the image feature extraction layer 302 to obtain feature maps corresponding to each of the multiple channels, and the feature maps corresponding to each of the multiple channels are input into the stems layer 303, and the feature maps corresponding to each of the multiple channels are subjected to channel normalization processing to obtain multiple target feature maps with the same number of channels, and the multiple target feature maps are respectively input into the first feature extraction model 304, the second feature extraction model 305 and the third feature extraction model 306 to obtain the target bounding box features of the detection target and the target appearance features of the detection target, and then the target bounding box features of the detection target, the target appearance features of the detection target, and the trajectory features 307 of the tracking trajectory are input into the matching module 308 to determine the matching result 309 of the detection target and the tracking trajectory.

[0075] According to an embodiment of the present disclosure, the first feature extraction model 304 may include the following processing layers: Reg_conv, Reg_preds, and Obj_preds. Reg_conv is implemented by a convolutional neural network and is used to extract the target bounding box features of the detection target included in the multiple target feature maps; Reg_preds and Obj_preds may also be implemented by a convolutional neural network and are used to map the extracted target bounding box features of the detection target from the feature space to the prediction space.

[0076] According to an embodiment of the present disclosure, the second feature extraction model 305 may include the following processing layers: Redi_conv, ConvlLSTM, Reid_preds. Among them, Redi_conv is the first appearance extraction sub-model, which may include: two 3×3 convolution sub-layers, BN (Batch Normalization) sub-layers and SiLU (Sigmoid Gated Linear Unit) sub-layers. The two 3×3 convolution sub-layers are responsible for extracting relevant information about the initial appearance features of the detection target; the BN layer is responsible for batch data normalization. In this network structure, all data obtained through the convolution sub-layer in the same batch are normalized. This structure is intended to normalize the data to a uniform interval, thereby reducing the degree of data divergence, making the network model training process more stable, and reducing the complexity of the neural network learning process; the SiLU sub-layer is an activation function with the characteristics of sigmoid and ReLU (Rectified Linear Unit), which is used to appropriately map the feature value to the prediction interval. Compared with the ReLU sublayer, the activation function used in the SiLU sublayer has a smoother curve near 0, so it can retain more input information.

[0077] According to an embodiment of the present disclosure, the Reid_preds layer is composed of a 1×1 convolutional layer for mapping data from feature space to prediction space. ConvlLSTM is a temporal processing submodel for temporal processing of historical appearance features and initial appearance features of neighboring historical detection targets, thereby outputting target appearance features of the detection target of the current frame including temporal information.

[0078] According to an embodiment of the present disclosure, the third feature extraction model 306 may include: Cls_conv and Cls_preds, wherein Cls_conv and Cls_preds may be obtained by a convolutional neural network, Cls_conv is used to limit the tracking trajectory category and the upper limit of the tracking trajectory; Cls_preds is used for feature mapping.

[0079] According to the embodiments of the present disclosure, when performing target detection, in addition to the bounding box information of the detection target, the appearance information and timing information of the detection target are also introduced, thereby improving the detection accuracy of the detection target and further improving the overall tracking performance.

[0080] According to an embodiment of the present disclosure, a plurality of target feature maps are input into a second feature extraction model to obtain target appearance features of the detection target, and the following steps may also be included.

[0081] The plurality of target feature maps are input into the second appearance extraction sub-model to obtain target appearance features of the detection target, wherein the second appearance extraction sub-model includes a feature extraction layer and a mapping layer, and the feature extraction layer includes at least six convolutional sub-layers.

[0082] According to an embodiment of the present disclosure, the second appearance extraction sub-model may include at least six convolutional sub-layers, that is, the appearance feature extraction structure is deepened, so that more accurate appearance features can be extracted.

[0083] Figure 4 A model structure diagram for determining target appearance features of a detection target according to another embodiment of the present disclosure is schematically shown.

[0084] like Figure 4 As shown, the second appearance extraction sub-model may include a convs layer and a pred layer, wherein the convs layer may include six 3×3 convolutional sub-layers, namely, 3×3conv2d, a BN sub-layer, and a SiLU sub-layer; the pred layer may be composed of a 1×1 convolutional layer for mapping data from the feature space to the prediction space.

[0085] According to the embodiments of the present disclosure, during the research process, the present disclosure took into account the inference time factor of the target tracking algorithm, and found that an overly deep network structure would affect the speed of model inference and the immediacy of the tracking method. Under the condition that the inference speed was less affected, a comparative experiment was conducted on the second appearance extraction sub-models of different depths, i.e., with different numbers of convolutional sub-layers, and the obtained tracking results were compared using the commonly used indicators for evaluating target tracking results, such as IDF1 (Identification F1-Score), MOTA (Multiple Object Tracking Accuracy), and IDsw (Identification Switch). It was found that the second appearance extraction sub-model with six convolutional sub-layers had the best tracking effect for detecting targets.

[0086] According to an embodiment of the present disclosure,

[0087] The target bounding box features include: target bounding box confidence and target bounding box position; the track bounding box features include: track bounding box position and track speed and track acceleration corresponding to the track bounding box position; based on the target bounding box features of the detected target, the target appearance features of the detected target and the track features of the tracking track, determining the matching result between the detected target and the tracking track may include the following operations.

[0088] When it is determined that the target bounding box confidence is greater than the preset confidence, the distance between the target appearance feature of the detected target and the track appearance feature is calculated to obtain the appearance distance; based on the track bounding box position, the track speed corresponding to the track bounding box position, and the track acceleration, the predicted bounding box position of the tracking track in the current frame is predicted; the distance between the predicted bounding box position and the target bounding box position is calculated to obtain the bounding box distance; based on the appearance distance and the bounding box distance, a matching cost matrix is ​​obtained; based on the matching cost matrix, the matching result of the detected target and the tracking track is determined. According to the embodiments of the present disclosure, the detected target can be matched with the tracking track by using different matching methods through the target bounding box confidence of the detected target.

[0089] According to the embodiments of the present disclosure, there is no limitation on the preset confidence level, and it can be limited according to actual needs.

[0090] According to the embodiments of the present disclosure, there is no limitation on the calculation method of the appearance distance, which may be a cosine distance calculation or a Pearson correlation coefficient calculation, etc., wherein the cosine distance calculation method is shown in the following formula (1).

[0091] ; (1)

[0092] in, Represents the target appearance features of the detection target. is the trajectory appearance feature of the tracking trajectory.

[0093] According to an embodiment of the present disclosure, the prediction method for the predicted bounding box position of the current frame is not limited, and the prediction may be performed using a Kalman filter algorithm.

[0094] According to the embodiments of the present disclosure, there is no limitation on the calculation method of the bounding box distance, and the bounding box distance may be obtained by calculating the Mahalanobis distance, as specifically shown in formula (2).

[0095] ; (2)

[0096] in, represents the Mahalanobis distance between the i-th tracking trajectory and the j-th detection target, is the vector representing the bounding box position of the jth detected target, is the vector representing the predicted bounding box position of the i-th tracking trajectory, for and The covariance matrix between .

[0097] According to an embodiment of the present disclosure, the appearance distance between each detection target and each tracking trajectory can be used as a parameter of the corresponding position, thereby obtaining an initial matching cost matrix; and comparing whether the bounding box distance between the detection target and the tracking trajectory corresponding to each position in the initial matching cost matrix is ​​less than a preset bounding box threshold; thereby, when the bounding box distance is less than the preset bounding box distance threshold, the parameter of the position is set to ∞; and when the bounding box distance is greater than the preset bounding box distance threshold, the parameter of the position is not processed, and it is still set to the corresponding bounding box distance, thereby obtaining a final matching cost matrix.

[0098] According to an embodiment of the present disclosure, a tracking trajectory matching each detected target may be determined based on a matching cost matrix by using a KM algorithm (Kuhn-Munkres).

[0099] According to the embodiments of the present disclosure, a matching cost matrix is ​​determined by combining the appearance distance with the bounding box distance, and a final matching result of the detection target and the tracking trajectory is determined, so that the matching accuracy can be improved, that is, the tracking accuracy can be improved.

[0100] According to an embodiment of the present disclosure, the target bounding box features include: target bounding box confidence, target bounding box position and target bounding box size information; the track bounding box features include: track bounding box position, track speed corresponding to the track bounding box position, track acceleration and track bounding box size information; wherein, based on the target bounding box features of the detected target, the target appearance features of the detected target and the track features of the tracking track, determining the matching result between the detected target and the tracking track may include the following operations.

[0101] When it is determined that the target bounding box confidence is less than or equal to the preset confidence, the predicted bounding box position and predicted bounding box size information of the tracking trajectory in the current frame are predicted based on the trajectory bounding box position, the trajectory speed corresponding to the trajectory bounding box position, the trajectory acceleration and the trajectory bounding box size information; based on the predicted bounding box position and the predicted bounding box size information, a predicted sub-image is determined; based on the target bounding box position and the target bounding box size information, a current sub-image is determined; based on the pixel set included in the predicted sub-image and the pixel set included in the current sub-image, a distance between the predicted bounding box feature and the target bounding box feature is determined; based on the distance, a matching result of the detected target and the tracking trajectory is determined.

[0102] According to an embodiment of the present disclosure, when the target bounding box confidence is less than or equal to a preset confidence, it can be considered that the detected target at this time may be in a more extreme scene, such as a high occlusion scene, so that the appearance features of the inspection target are also unreliable.

[0103] According to an embodiment of the present disclosure, a Kalman filter algorithm may be used to predict the predicted bounding box position and predicted bounding box size information of the current frame.

[0104] According to an embodiment of the present disclosure, the target bounding box size information may include bounding box width, bounding box width and bounding box aspect ratio. The track bounding box size information may include track bounding box width, track bounding box width and track bounding box aspect ratio.

[0105] According to an embodiment of the present disclosure, the trajectory speed and trajectory acceleration corresponding to the trajectory bounding box position may be determined by the historical bounding box positions of the historical detection targets of a plurality of historical frames matched with the tracking trajectory.

[0106] According to an embodiment of the present disclosure, when there are multiple tracking trajectories and multiple detection targets, matching between unmatched tracking trajectories and unmatched detection targets may be performed.

[0107] According to an embodiment of the present disclosure, when matching is performed, a predicted sub-image corresponding to the predicted bounding box position and a current sub-image corresponding to the bounding box position can be determined from the image to be detected in the current frame. The distance between the predicted bounding box feature and the bounding box feature can be determined by the pixel set included in the predicted sub-image and the current sub-image. The specific calculation formula can be shown in the following formula (3).

[0108] ; (3)

[0109] Among them, IoU(A, B) is the distance between the predicted bounding box feature and the bounding box feature, A is the pixel set of the predicted sub-image, and B is the pixel set of the current sub-image.

[0110] According to an embodiment of the present disclosure, a matching cost matrix can be constructed by using the distance between the trajectory bounding box features of each tracking trajectory and the target bounding box features of each detection target, and the tracking trajectory matching each detection target can be determined using the KM algorithm, so as to obtain matching results between multiple detection targets and multiple tracking trajectories.

[0111] According to an embodiment of the present disclosure, by calculating the intersection between the pixel set of the predicted sub-image and the pixel set of the current sub-image, the distance between the predicted bounding box feature and the bounding box feature is determined, and then whether the detected target matches the tracking trajectory is determined, so that the matching accuracy can be improved. According to an embodiment of the present disclosure, the target tracking method also includes the steps of:

[0112] When it is determined that the matching result represents that the detection target matches the tracking trajectory, the target bounding box feature is used to replace the trajectory bounding box feature in the trajectory feature; based on the target appearance feature, the trajectory appearance feature and the preset appearance weight value, an updated appearance feature is obtained; and the trajectory appearance feature in the trajectory feature is replaced by the updated appearance feature.

[0113] According to an embodiment of the present disclosure, when it is determined that the detected target matches the tracking trajectory, the appearance feature, the trajectory appearance feature and the preset appearance weight value can be used to obtain the updated appearance feature for updating, and the calculation formula thereof can be shown as the following formula (4).

[0114] ; (4)

[0115] in, is the vector representing the updated appearance features, is a vector representing the appearance features of the trajectory, is the vector trajectory appearance feature that represents the appearance feature, and 0.1 and 0.9 are the preset appearance weight values.

[0116] According to an embodiment of the present disclosure, the updated appearance features calculated in the above manner can retain certain historical appearance features on the basis of the appearance features of the current frame, thereby improving the accuracy of matching the detection target and the tracking trajectory in subsequent frames.

[0117] According to an embodiment of the present disclosure, the target tracking method may further include the following steps.

[0118] When it is determined that the matching result indicates that the detected target does not match the tracking track and the target bounding box confidence of the detected target is greater than a preset confidence, the detected target is used as a new tracking target; the track features of the new tracking track are determined based on the target appearance features and the target bounding box features; the current frame is used as a lost frame of the tracking track, so that the tracking track is deleted when the lost frames of the tracking track are continuous frames and the number of lost frames is greater than a preset frame number threshold. According to an embodiment of the present disclosure, the preset frame number threshold is not limited and can be set according to actual conditions.

[0119] According to an embodiment of the present disclosure, a new tracking trajectory associated with a detected target may be established.

[0120] According to an embodiment of the present disclosure, when it is determined that the matching result indicates that the detection target does not match the tracking trajectory, the number of lost frames of the tracking trajectory can be counted for the tracking trajectory. When the lost frames are continuous frames and the number of frames is greater than a preset frame number threshold, the tracking trajectory is discarded and the relevant information of the tracking trajectory is deleted.

[0121] According to an embodiment of the present disclosure, the detected target may also be used as a new tracking track, and its target appearance features and target bounding box features may be determined as track features of the new tracking track.

[0122] According to an embodiment of the present disclosure, when the tracking trajectory is not deleted, for the image to be detected in the next frame of the current frame, its tracking trajectory includes a new tracking trajectory and a tracking trajectory.

[0123] According to an embodiment of the present disclosure, the target tracking method is executed for each frame of the image to be detected in the video to be detected, so as to obtain tracking information of all tracking tracks in the video to be detected.

[0124] According to an embodiment of the present disclosure, the first feature extraction model and the second feature extraction model are trained in the following manner.

[0125] Input multiple sample target feature maps of the sample image into the initial first feature extraction model to obtain sample bounding box features of the sample targets included in the sample image; obtain a first loss value based on the sample bounding box features and the true bounding box features; input multiple sample target feature maps into the initial second feature extraction model to obtain sample appearance features of the sample targets; obtain a second loss value based on the predicted identification and the true identification of the sample appearance features; determine a target loss value based on the first loss value, a first weight value corresponding to the first loss value, the second loss value, and a second weight value corresponding to the second loss value; train the initial first feature extraction model and the initial second feature extraction model based on the target loss value to obtain a first feature extraction model and a second feature extraction model.

[0126] According to an embodiment of the present disclosure, the formula for calculating the target loss function may be as shown in the following formula (5).

[0127] (5)

[0128] Among them, det_loss is the first loss value, loss_id is the second loss value, is the first weight value, is the second weight value.

[0129] According to the embodiments of the present disclosure, by jointly training the first feature extraction model and the second feature extraction model, the accuracy of model training can be improved, thereby making the final tracking result more accurate.

[0130] According to an embodiment of the present disclosure, at least one sample object included in a sample image may be pre-labeled with a bounding box and a logo, so as to determine a true bounding box feature and a true logo of each sample object.

[0131] According to an embodiment of the present disclosure, the predicted identification of the sample appearance feature may be output by the second feature extraction model. When a plurality of sample target feature maps are input into the initial second feature extraction model, the sample appearance features and the predicted identification of the sample target may be obtained.

[0132] According to the embodiments of the present disclosure, appearance information and timing information are introduced in the training process. These information are back-propagated through the loss function to positively constrain the detection box information prediction process, thereby improving the accuracy of target detection and further improving the tracking performance of the algorithm.

[0133] Figure 5 A flow chart of a target tracking method according to another embodiment of the present disclosure is schematically shown.

[0134] like Figure 5 As shown, the method includes operations S501 to S513.

[0135] In operation S501, the video to be detected is split into continuous image frames;

[0136] In operation S502, multi-channel feature extraction is performed on the image to be detected in the current frame to obtain feature maps corresponding to each of the multiple channels.

[0137] In operation S503, the feature maps corresponding to the multiple channels are subjected to channel-unified processing to obtain multiple target feature maps with the same number of channels.

[0138] In operation S504, a plurality of target feature maps are input into a first feature extraction model to obtain target bounding box features of the detected target.

[0139] In operation S505, a plurality of target feature maps are input into a first appearance extraction sub-model to obtain initial appearance features of the detection target.

[0140] In operation S506, a time series information analysis is performed on the initial appearance features of the detection target and the historical appearance features of at least one neighboring historical detection target to obtain the target appearance features of the detection target, wherein the neighboring historical detection target is a historical detection target included in the image to be detected of the M historical frames closest to the current frame, and M is an integer greater than or equal to 1.

[0141] In operation S507 , when there are multiple detection targets, the detection targets are classified according to the bounding box confidences included in the bounding box features of the multiple detection targets to obtain high-confidence targets and low-confidence targets.

[0142] In operation S508, based on the target bounding box features of the high-confidence target, the target appearance features of the high-confidence target, and the track features of the tracking track, the high-confidence target is matched with the tracking track to obtain a matching result. If the matching result indicates that there is a tracking track that matches the high-confidence target, operation S512 is performed. If the matching result indicates that there is no tracking track that matches the high-confidence target, operation S510 is performed.

[0143] According to an embodiment of the present disclosure, the tracking trajectory is at least one.

[0144] In operation S509, based on the target bounding box features of the low-confidence target, the target appearance features of the low-confidence target, and the track features of the remaining tracking tracks that do not match the high-confidence target, the low-confidence target is matched with the remaining tracking tracks to obtain a matching result. If the matching result indicates that there is a remaining tracking track that matches the low-confidence target, operation S512 is performed. If the matching result indicates that there is no remaining tracking track that matches the low-confidence target, operation S511 is performed.

[0145] In operation S510 , the detected target is used as a new tracking track.

[0146] In operation S511, low confidence targets are discarded.

[0147] In operation S512 , trajectory features of the tracking trajectory are updated based on the object appearance features and the object bounding box features of the detected object.

[0148] In operation S513 , the image to be detected in the next frame is read, and the image to be detected in the next frame is processed based on S502 to S509 .

[0149] Based on the above target tracking method, the present disclosure also provides a target tracking device. Figure 6 The device is described in detail.

[0150] Figure 6 The structural block diagram of a target tracking device according to an embodiment of the present disclosure is schematically shown.

[0151] like Figure 6 As shown, the target tracking device 600 of this embodiment includes an extraction module 610 , a detection module 620 and a matching module 630 .

[0152] The extraction module 610 is used to perform multi-channel feature extraction on the image to be detected in the current frame to obtain feature maps corresponding to each of the multiple channels, wherein the image to be detected includes a detection target.

[0153] The detection module 620 is used to perform target detection on the feature maps corresponding to each of the multiple channels to obtain target bounding box features of the detected target and target appearance features of the detected target.

[0154] The matching module 630 is used to determine a matching result between the detection target and the tracking trajectory based on the target bounding box feature of the detection target, the target appearance feature of the detection target and the trajectory feature of the tracking trajectory, wherein the tracking trajectory is determined based on the historical detection target continuously tracked in at least one historical frame, and the trajectory feature includes the trajectory bounding box feature determined by the historical bounding box feature of the historical detection target and the trajectory appearance feature determined by the historical appearance feature. Trajectory appearance feature.

[0155] According to an embodiment of the present disclosure, the detection module 620 includes: a processing submodule, a first extraction submodule and a second extraction submodule.

[0156] The processing submodule is used to uniformly process the feature maps corresponding to the multiple channels to obtain multiple target feature maps with the same number of channels.

[0157] The first extraction submodule is used to input multiple target feature maps into a first feature extraction model to obtain target bounding box features of the detected target.

[0158] The second extraction submodule is used to input the multiple target feature maps into the second feature extraction model to obtain the target appearance features of the detection target.

[0159] According to an embodiment of the present disclosure, the second extraction submodule includes: a first extraction unit and an information analysis unit.

[0160] The first extraction unit is used to input multiple target feature maps into the first appearance extraction sub-model to obtain initial appearance features of the detection target.

[0161] An information analysis unit is used to perform time series information analysis on the initial appearance features of the detection target and the historical appearance features of at least one adjacent historical detection target to obtain the target appearance features of the detection target, wherein the adjacent historical detection target is a historical detection target included in the image to be detected of the M historical frames closest to the current frame, and M is an integer greater than or equal to 1.

[0162] According to an embodiment of the present disclosure, appearance feature extraction is performed on a plurality of target feature maps to obtain the appearance features of each of the detection targets, including: a second extraction unit.

[0163] The second extraction unit is used to input multiple target feature maps into the second appearance extraction sub-model to obtain the target appearance features of the detection target, wherein the second appearance extraction sub-model includes a feature extraction layer and a mapping layer, and the feature extraction layer includes at least six convolution sub-layers.

[0164] According to an embodiment of the present disclosure, the target bounding box feature includes: target bounding box confidence and target bounding box position; the track bounding box feature includes: track bounding box position and track speed and track acceleration corresponding to the track bounding box position. The matching module 630 includes: a first calculation submodule, a first prediction submodule, a second calculation submodule, a matrix determination submodule and a first result determination submodule.

[0165] The first calculation submodule is used to calculate the distance between the target appearance feature and the track appearance feature of the detected target to obtain the appearance distance when it is determined that the target bounding box confidence is greater than a preset confidence.

[0166] The first prediction submodule is used to predict the predicted bounding box position of the tracking trajectory in the current frame based on the trajectory bounding box position, the trajectory speed and the trajectory acceleration corresponding to the trajectory bounding box position.

[0167] The second calculation submodule is used to calculate the distance between the predicted bounding box position and the target bounding box position to obtain the bounding box distance.

[0168] The matrix determination submodule is used to obtain the matching cost matrix based on the appearance distance and the bounding box distance.

[0169] The first result determination submodule is used to determine the matching result of the detection target and the tracking trajectory based on the matching cost matrix.

[0170] According to an embodiment of the present disclosure, the target bounding box feature includes: target bounding box confidence, target bounding box position and target bounding box size information; the track bounding box feature includes: track bounding box position, track speed corresponding to the track bounding box position, track acceleration and track bounding box size information. The matching module 630 includes: a second prediction submodule, a first image determination submodule, a second image determination submodule, a third calculation submodule and a first result determination submodule.

[0171] The second prediction submodule is used to predict the predicted bounding box position and predicted bounding box size information of the tracking trajectory in the current frame based on the trajectory bounding box position, the trajectory speed, the trajectory acceleration and the trajectory bounding box size information corresponding to the trajectory bounding box position when it is determined that the target bounding box confidence is less than or equal to the preset confidence.

[0172] The first image determination submodule is used to determine the predicted sub-image based on the predicted bounding box position and the predicted bounding box size information.

[0173] The second image determination submodule is used to determine the current sub-image based on the target bounding box position and the target bounding box size information.

[0174] The third calculation submodule is used to determine the distance between the predicted bounding box feature and the target bounding box feature based on the pixel set included in the predicted sub-image and the pixel set included in the current sub-image.

[0175] The first result determination submodule is used to determine the matching result between the detection target and the tracking trajectory based on the distance.

[0176] According to an embodiment of the present disclosure, the target tracking device 600 further includes: a first replacement module, an update feature determination module and a second replacement module.

[0177] The first replacement module is used to replace the track bounding box feature in the track feature with the target bounding box feature when it is determined that the matching result represents that the detected target matches the tracking track.

[0178] The update feature determination module is used to obtain the update appearance feature based on the target appearance feature, the track appearance feature and the preset appearance weight value.

[0179] The second replacement module is used to replace the track appearance feature in the track feature with the updated appearance feature.

[0180] According to an embodiment of the present disclosure, the target tracking device 600 further includes: a new target determination module, a feature determination module and a deletion module.

[0181] The new target determination module is used to use the current frame as the lost frame of the tracking trajectory, so that when the lost frames of the tracking trajectory are continuous frames and the number of lost frames is greater than a preset frame number threshold, the detected target is used as a new tracking target. The feature determination module is used to determine the trajectory features of the new tracking trajectory based on the target appearance features and the target bounding box features.

[0182] The deletion module is used to use the current frame as the lost frame of the tracking trajectory, so that the tracking trajectory is deleted when the lost frames of the tracking trajectory are continuous frames and the number of lost frames is greater than a preset frame number threshold.

[0183] According to an embodiment of the present disclosure, the target tracking device 600 further includes: a first input module, a first loss value determination module, a second input module, a second loss value determination module, a target loss value determination module and a training module.

[0184] The first input module is used to input multiple sample target feature maps of the sample image into the initial first feature extraction model to obtain sample bounding box features of the sample targets included in the sample image.

[0185] The first loss value determination module is used to obtain a first loss value based on the sample bounding box features and the real bounding box features.

[0186] The second input module is used to input a plurality of sample target feature maps into the initial second feature extraction model to obtain sample appearance features of the sample targets.

[0187] The second loss value determination module is used to obtain a second loss value based on the predicted identification and the true identification of the sample appearance feature.

[0188] The target loss value determination module is used to determine the target loss value based on the first loss value, the first weight value corresponding to the first loss value, the second loss value, and the second weight value corresponding to the second loss value.

[0189] The training module is used to train the initial first feature extraction model and the initial second feature extraction model based on the target loss value to obtain the first feature extraction model and the second feature extraction model.

[0190] According to an embodiment of the present disclosure, any multiple modules of the extraction module 610, the detection module 620 and the matching module 630 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the extraction module 610, the detection module 620 and the matching module 630 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation modes of software, hardware and firmware or in any appropriate combination of any of them. Alternatively, at least one of the extraction module 610, the detection module 620 and the matching module 630 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding function can be executed.

[0191] Figure 7 A block diagram of an electronic device suitable for implementing a target tracking method according to an embodiment of the present disclosure is schematically shown.

[0192] like Figure 7As shown, the electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage part 708 to a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (for example, an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include an onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0193] In RAM 703, various programs and data required for the operation of electronic device 700 are stored. Processor 701, ROM 702 and RAM 703 are connected to each other via bus 704. Processor 701 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM 702 and / or RAM 703. It should be noted that the program can also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in one or more memories.

[0194] According to an embodiment of the present disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 708 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed, so that a computer program read therefrom is installed into the storage portion 708 as needed.

[0195] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.

[0196] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 702 and / or RAM 703 described above and / or one or more memories other than ROM 702 and RAM 703.

[0197] The embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the target tracking method provided by the embodiment of the present disclosure.

[0198] The above functions defined in the system / device of the embodiment of the present disclosure are performed when the computer program is executed by the processor 701. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0199] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 709, and / or installed from the removable medium 711. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0200] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.

[0201] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).

[0202] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0203] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways. All of these combinations and / or combinations fall within the scope of the present disclosure.

[0204] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A target tracking method, characterized in that: The method comprises: Performing multi-channel feature extraction on the image to be detected of the current frame to obtain feature maps corresponding to each of the multiple channels, wherein the image to be detected includes a detection target; Performing target detection on the feature maps corresponding to each of the multiple channels to obtain target bounding box features of the detected target and target appearance features of the detected target; Based on the target bounding box features of the detection target, the target appearance features of the detection target and the trajectory features of the tracking trajectory, a matching result between the detection target and the tracking trajectory is determined, wherein the tracking trajectory is determined based on a historical detection target that is continuously tracked in at least one historical frame, and the trajectory features include trajectory bounding box features determined by historical bounding box features of the historical detection target and trajectory appearance features determined by historical appearance features.

2. The method according to claim 1, characterized in that The performing target detection on the feature maps corresponding to each of the multiple channels to obtain target bounding box features of the detected target and target appearance features of the detected target includes: Performing channel-unified processing on the feature maps corresponding to the multiple channels to obtain multiple target feature maps with the same number of channels; Inputting the plurality of target feature maps into a first feature extraction model to obtain target bounding box features of the detected target; The plurality of target feature maps are input into a second feature extraction model to obtain target appearance features of the detection target.

3. The method according to claim 2, characterized in that The step of inputting the plurality of target feature maps into a second feature extraction model to obtain target appearance features of the detection target includes: Inputting the plurality of target feature maps into a first appearance extraction sub-model to obtain initial appearance features of the detection target; Performing time series information analysis on the initial appearance features of the detection target and the historical appearance features of at least one neighboring historical detection target to obtain the target appearance features of the detection target, wherein the neighboring historical detection target is a historical detection target included in the image to be detected of M historical frames closest to the current frame, and M is an integer greater than or equal to 1.

4. The method according to claim 2, characterized in that: The step of inputting the plurality of target feature maps into a second feature extraction model to obtain target appearance features of the detection target includes: Inputting the plurality of target feature maps into a second appearance extraction sub-model to obtain target appearance features of the detection target, wherein the second appearance extraction sub-model comprises a feature extraction layer and a mapping layer, and the feature extraction layer comprises at least six convolutional sub-layers.

5. The method according to claim 1, characterized in that The target bounding box features include: target bounding box confidence and target bounding box position; the track bounding box features include: track bounding box position and track speed and track acceleration corresponding to the track bounding box position; Wherein, determining the matching result between the detection target and the tracking trajectory based on the target bounding box feature of the detection target, the target appearance feature of the detection target and the trajectory features of the tracking trajectory respectively includes: When it is determined that the confidence of the target bounding box is greater than a preset confidence, Calculating the distance between the target appearance feature of the detection target and the track appearance feature to obtain an appearance distance; Predicting a predicted bounding box position of the tracking trajectory in the current frame based on a trajectory bounding box position, a trajectory velocity and a trajectory acceleration corresponding to the trajectory bounding box position; Calculating the distance between the predicted bounding box position and the target bounding box position to obtain a bounding box distance; Obtaining a matching cost matrix based on the appearance distance and the bounding box distance; Based on the matching cost matrix, a matching result between the detection target and the tracking trajectory is determined.

6. The method according to claim 1, characterized in that The target bounding box features include: target bounding box confidence, target bounding box position and target bounding box size information; the track bounding box features include: track bounding box position, track speed corresponding to the track bounding box position, track acceleration and track bounding box size information; Wherein, determining the matching result between the detection target and the tracking trajectory based on the target bounding box feature of the detection target, the target appearance feature of the detection target and the trajectory features of the tracking trajectory respectively includes: When it is determined that the target bounding box confidence is less than or equal to a preset confidence, Predicting a predicted bounding box position and predicted bounding box size information of the tracking trajectory in the current frame based on the trajectory bounding box position, a trajectory speed corresponding to the trajectory bounding box position, the trajectory acceleration, and the trajectory bounding box size information; Determining a predicted sub-image based on the predicted bounding box position and the predicted bounding box size information; Determining a current sub-image based on the target bounding box position and the target bounding box size information; Determining a distance between the predicted bounding box feature and the target bounding box feature based on a pixel set included in the predicted sub-image and a pixel set included in the current sub-image; Based on the distance, a matching result between the detection target and the tracking trajectory is determined.

7. The method according to claim 1, characterized in that The method further comprises: In the case where it is determined that the matching result indicates that the detection target matches the tracking trajectory, Replacing the trajectory bounding box features in the trajectory features with the target bounding box features; Based on the target appearance feature, the trajectory appearance feature and a preset appearance weight value, an updated appearance feature is obtained; The track appearance features in the track features are replaced with the updated appearance features.

8. The method according to claim 1, characterized in that The method further comprises: In a case where it is determined that the matching result indicates that the detection target does not match the tracking trajectory and the target bounding box confidence of the detection target is greater than a preset confidence, taking the detection target as a new tracking target; Determining a trajectory feature of the new tracking trajectory based on the target appearance feature and the target bounding box feature; The current frame is used as a lost frame of the tracking trajectory, so that when the lost frames of the tracking trajectory are consecutive frames and the number of lost frames is greater than a preset frame number threshold, the tracking trajectory is deleted.

9. The method according to claim 2, characterized in that: The first feature extraction model and the second feature extraction model are trained in the following manner: Inputting a plurality of sample target feature maps of the sample image into an initial first feature extraction model to obtain sample bounding box features of the sample targets included in the sample image; Obtaining a first loss value based on the sample bounding box features and the true bounding box features; Inputting a plurality of the sample target feature maps into an initial second feature extraction model to obtain sample appearance features of the sample targets; Obtaining a second loss value based on the predicted identification and the true identification of the sample appearance feature; Determine a target loss value based on the first loss value, a first weight value corresponding to the first loss value, the second loss value, and a second weight value corresponding to the second loss value; The initial first feature extraction model and the initial second feature extraction model are trained based on the target loss value to obtain the first feature extraction model and the second feature extraction model.

10. A target tracking device, characterized in that: The device comprises: An extraction module, used for performing multi-channel feature extraction on the image to be detected of the current frame to obtain feature maps corresponding to each of the multiple channels, wherein the image to be detected includes a detection target; A detection module, configured to perform target detection on the feature maps corresponding to each of the multiple channels, and obtain target bounding box features of the detected target and target appearance features of the detected target; A matching module is used to determine a matching result between the detection target and the tracking trajectory based on a target bounding box feature of the detection target, a target appearance feature of the detection target, and a trajectory feature of the tracking trajectory, wherein the tracking trajectory is determined based on a historical detection target that is continuously tracked in at least one historical frame, and the trajectory feature includes a trajectory boundary feature determined by a historical bounding box feature of the historical detection target and a trajectory appearance feature determined by a historical appearance feature.

Citation Information

Patent Citations

  • Multi-target tracking method and device, computer equipment and storage medium

    CN118038341A