Image tracking method, target tracking model, readable medium and electronic equipment
By using convolutional calculation in the target tracking method to extract the first feature image, the problem of waste of computing resources in the prior art is solved, and more efficient image tracking is achieved.
Patent Information
- Application Number
- CN202311545999.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
Existing target tracking methods require a lot of computing resources, especially when the resolution of the images to be processed is low, resulting in wasted computing resources in the image tracking process.
Convolutional calculation is used to extract the first feature image to obtain a second feature image containing the correlation relationship between image feature values, reducing the computing resource consumption during image tracking.
Convolutional calculation reduces the computing resource consumption during the image tracking process and improves the efficiency of image tracking, especially when the resolution of the image to be processed is low.
Smart Images

Figure CN120020874A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an image tracking method, a target tracking model, a readable medium, and an electronic device. Background Art
[0002] Object tracking refers to determining a moving object (target object) in an image sequence based on the image sequence, identifying the target object in each frame image of the image sequence, and marking the identified target object, that is, determining the corresponding identifier (ID).
[0003] Object tracking includes single-object tracking and multi-object tracking (MOT), and is widely used in scenarios such as traffic analysis, security monitoring, pedestrian pose estimation, and autonomous driving assistance. For example, referring to Figure 1 the schematic diagram in the security monitoring scenario shown in Figure 1 the k-th frame image in (a) in Figure 1 after object tracking, the k-th frame image after marking is obtained, where the target objects are respectively marked as "Per1", "Per2", "Per3", and "Car1".
[0004]
[0005] Currently, object tracking methods mainly include: object tracking methods based on traditional technologies, such as those based on Kalman filtering technology, and object tracking methods based on deep learning technologies, such as those based on recursive neural networks (RNN). Summary of the Invention
[0006] The purpose of this application is to provide an image tracking method, a target tracking model, a readable medium, and an electronic device.
[0007] The first aspect of this application provides an image tracking method, including: obtaining the i-th frame image in a video to be detected, where i is greater than 1; obtaining a first feature image of the i-th frame image, where the first feature image includes a plurality of image feature values; performing convolution calculation on the first feature image to obtain N second feature images, where N is an integer greater than 0, and the second feature image includes: an image feature value representing the association relationship between a plurality of image feature values in the first feature image; determining first tracking data of the i-th frame image based on the N second feature images.
[0008] It can be understood that when the resolution of the image to be processed is relatively low, the accuracy of the first feature image obtained by extracting features from these images to be processed with relatively low resolution is usually low. Convolution calculation can be used to extract features from the first feature image to obtain a second feature image that includes the correlation relationship between the image feature values in the first feature image. The complexity of convolution calculation is relatively low, which can reduce the computing resources in the image tracking process.
[0009] In a possible implementation of the first aspect above, performing convolution calculation on the first feature image to obtain N second feature images includes: performing at least one convolution calculation on the first feature image based on the convolution operator in the feature fusion module to obtain N second feature images.
[0010] In a possible implementation of the first aspect above, performing convolution calculation on the first feature image to obtain N second feature images includes: corresponding to the resolution of the i-th frame image being lower than the resolution threshold, using multiple different sampling granularities to respectively downsample the first feature image to obtain multiple first feature sub-images with different resolutions; performing at least one convolution calculation on each of the first feature sub-images with different resolutions to obtain second feature images corresponding to the multiple first feature sub-images.
[0011] In a possible implementation of the first aspect above, performing at least one convolution calculation on each of the first feature sub-images with different resolutions to obtain second feature images corresponding to the multiple first feature sub-images includes: the calculation method of the second feature image P' of the first feature sub-image P among the multiple first feature sub-images is as follows: select the first feature sub-image P and the first feature sub-image Q with a resolution higher than that of the first feature sub-image P from the multiple first feature sub-images, and perform at least one convolution calculation on the first feature sub-image Q to obtain the corresponding second feature image Q'; perform feature merging on the second feature image Q' and the first feature sub-image P to obtain a fusion feature image K that includes the image features of the second feature image Q' and the first feature sub-image P; perform at least one convolution calculation on the fusion feature image K to obtain the second feature image P' corresponding to the first feature sub-image P.
[0012] In a possible implementation of the first aspect above, it further includes: performing a normalization operation on the second feature image; and determining the first tracking data of the i-th frame image based on the N second feature images, including: determining the tracking data of the i-th frame image based on the second feature image after the normalization operation.
[0013] In a possible implementation of the first aspect above, the first tracking data includes the tracking boxes of at least one first target object in the i-th frame image and the identification information corresponding to each first target object.
[0014] In a possible implementation of the above first aspect, determining the first tracking data of the i-th frame image based on N second feature images includes: determining the identification information of the second target object among at least one first target object in the following manner: determining the image features of the second target object based on the N second feature images, and obtaining the third tracking data corresponding to the (i - 1)-th frame image in the video to be detected, where the third tracking data includes: the image features corresponding to at least one third target object in the (i - 1)-th frame image and the identification information corresponding to each third target object; comparing the image features of the second target object with the image features of each third target object; if the matching degree between the image features of the fourth target object corresponding to at least one third target object and the image features of the second target object is greater than the first matching degree, determining the identification information of the fourth target object as the identification information of the second target object.
[0015] In the embodiment of the present application, if there is no target object in the third target object that matches the second target object, a new marking information is marked for the second target object, or the second target object is deleted.
[0016] In a possible implementation of the above first aspect, determining the image features of the second target object based on N second feature images includes: respectively performing bilinear sampling and nearest neighbor interpolation calculations on the image feature values in the regions corresponding to the second target object in the N second feature images to obtain the image features of the second target object.
[0017] The second aspect of the present application provides a target tracking model, which is characterized by including: a feature extraction module, configured to obtain the i-th frame image in the video to be detected and obtain the first feature image of the i-th frame image, where the first feature image includes multiple image feature values and i is greater than 1; a feature fusion module, configured to perform convolution calculations on the first feature image to obtain N second feature images, where N is an integer greater than 0, and the second feature image includes: image feature values representing the association relationship between multiple image feature values in the first feature image; a decoding module, configured to determine the first tracking data of the i-th frame image based on the N second feature images.
[0018] In a possible implementation of the above second aspect, it further includes: a normalization module, configured to perform a normalization operation on the second feature image or.
[0019] The third aspect of the present application provides a method for training a target tracking model, including: obtaining sample data, where the sample data includes sample images of multiple consecutive frames, sample marking information of target objects in each sample image, and sample tracking frames surrounding each target object; inputting the sample images of multiple consecutive frames into the target tracking model to obtain tracking data corresponding to each sample image, where the target tracking model includes a feature extraction module, a detection module, a feature fusion module, and a decoding module, and the tracking data includes detection tracking frames and detection marking information; adjusting the parameters of each module in the target tracking model according to the tracking data, the sample marking information, and the sample tracking frames in the sample data.
[0020] In the embodiments of the present application, based on parameters such as the shapes, positions, widths, and heights corresponding to the detection tracking frames and the sample tracking frames, a first loss value L1 is calculated. Based on the detection marking information and the sample marking information, a second loss value L2 is calculated. Weights α corresponding to the first loss value L1 and weights β corresponding to the second loss value L2 are determined, and further, the loss value Loss is the weighted average of the first loss value L1 and the second loss value L2 as the loss value Loss. Corresponding to the loss value being less than the loss threshold, the parameters currently used by each module are used as the parameters of each module of the trained model.
[0021] The fourth aspect of the present application provides a readable medium, on which instructions are stored, and when the instructions are executed on an electronic device, the electronic device is caused to execute any of the methods in the first aspect above.
[0022] The fifth aspect of the present application provides an electronic device, including a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the processors of the electronic device, for executing any of the methods in the first aspect above. Description of the Drawings
[0023] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 A schematic diagram in a security monitoring scenario is shown according to an embodiment of the present application;
[0025] Figure 2 A schematic diagram of the implementation process of a target tracking method is shown according to an embodiment of the present application;
[0026] Figure 3aAn embodiment according to the present application shows a schematic diagram of a target object and a detection box corresponding to the target object;
[0027] Figure 3b An embodiment according to the present application shows a schematic diagram of an image to be processed and a corresponding feature image;
[0028] Figure 3c An embodiment according to the present application shows a schematic diagram of a feature image and a corresponding feature image of the feature image;
[0029] Figure 4a An embodiment according to the present application shows a schematic structural diagram of a target tracking model based on the transformer algorithm;
[0030] Figure 4b An embodiment according to the present application shows a schematic diagram of module interaction of a target tracking model based on the transformer algorithm;
[0031] Figure 4c An embodiment according to the present application shows a schematic diagram of the process of an N-layer encoder processing a first feature image to obtain a second feature image containing semantic information;
[0032] Figure 5a An embodiment according to the present application shows a schematic diagram of a high-resolution image to be processed and a corresponding feature image;
[0033] Figure 5b An embodiment according to the present application shows a schematic diagram of a low-resolution image to be processed and a corresponding feature image;
[0034] Figure 5c An embodiment according to the present application shows a schematic diagram of a feature image and a corresponding feature image containing an association relationship;
[0035] Figure 6a An embodiment according to the present application shows a schematic structural diagram of a target tracking model, a schematic structural diagram of a target tracking model;
[0036] Figure 6b An embodiment according to the present application shows a schematic structural diagram of a feature fusion module;
[0037] Figure 6c An embodiment according to the present application shows a schematic diagram of feature fusion;
[0038] Figure 7 An embodiment according to the present application shows a flowchart of the training process of a target tracking model;
[0039] Figure 8A schematic flowchart of an image tracking method is shown according to an embodiment of the present application;
[0040] Figure 9 A schematic structural diagram of a single-layer decoder is shown according to an embodiment of the present application;
[0041] Figure 10 A schematic interaction flowchart of a variable attention model is shown according to an embodiment of the present application;
[0042] Figure 11 A schematic diagram of a processing process of an image tracking method is shown according to an embodiment of the present application;
[0043] Figure 12 A schematic diagram of a quantization process is shown according to an embodiment of the present application;
[0044] Figure 13 A schematic diagram of a flowchart of a target tracking method based on traditional technology is shown according to an embodiment of the present application;
[0045] Figure 14 A schematic diagram of identification interleaving is shown according to an embodiment of the present application;
[0046] Figure 15 A schematic diagram of the effect of adopting the image tracking method provided by the embodiment of the present application is shown according to an embodiment of the present application;
[0047] Figure 16 A schematic diagram of a flowchart of a target tracking method based on the transformer algorithm is shown according to an embodiment of the present application;
[0048] Figure 17a A schematic diagram of the effect of target tracking using a target tracking method based on traditional technology is shown according to an embodiment of the present application;
[0049] Figure 17b A schematic diagram of the effect of target tracking using a target tracking method based on the transformer algorithm is shown according to an embodiment of the present application;
[0050] Figure 17c A schematic diagram of the effect of target tracking using the image tracking method provided by the embodiment of the present application is shown according to an embodiment of the present application;
[0051] Figure 18 A schematic structural diagram of an electronic device is shown according to an embodiment of the present application. Detailed implementation manners
[0052] Illustrative embodiments of the present application include, but are not limited to, an image tracking method, a target tracking model, a readable medium, and an electronic device.
[0053] The image tracking method provided by the embodiments of the present application can be applied to terminal devices such as mobile phones, smart large screens, tablet computers, laptop computers, netbooks, personal digital assistants (PDAs), vehicle-mounted terminals, drones, virtual reality devices, etc. Exemplary embodiments of the terminal device include, but are not limited to, electronic devices equipped with or other operating systems. The embodiments of the present application do not impose any restrictions on the specific type of the terminal device.
[0054] The image tracking method provided by the embodiments of the present application can also be applied to a server connected to the terminal device.
[0055] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and elaborately described below with reference to the accompanying drawings.
[0056] As described above, object tracking includes single-object tracking and multi-object tracking, and is widely applied to scenarios such as traffic analysis, security monitoring, pedestrian pose estimation, autonomous driving assistance, etc.
[0057] Exemplarily, Figure 1 An application schematic diagram of a security monitoring scenario is shown according to an embodiment of the present application. As Figure 1 shown in (a) therein, the k-th frame image in the surveillance video is subjected to object tracking to obtain the marked k-th frame image, where the target objects are respectively marked as "Per1", "Per2", "Per3", and "Car1".
[0058] Figure 1 The second frame image in (b) therein is subjected to object tracking to obtain the marked (k + 1)-th frame image, where the target objects are respectively marked as "Per1", "Per2", "Per4", and "Car1". It can be understood that the target object marked as "Per1" in the (k + 1)-th frame image and the target object marked as "Per1" in the k-th frame image are the same target object, the target object marked as "Per2" in the (k + 1)-th frame image and the target object marked as "Per2" in the k-th frame image are the same target object, and the target object marked as "Car1" in the (k + 1)-th frame image and the target object marked as "Car1" in the k-th frame image are the same target object. Also, there is no target object in the (k + 1)-th frame that is the same as the target object marked as "Per3" in the k-th frame image, and there is a new target object (marked as "Per4").
[0059] In this way, by marking multiple frames of images in the surveillance video, trajectory prediction can be performed on the target object corresponding to the target mark, for example, determining the movement trajectory of the target object.
[0060] For example, Figure 2 An embodiment of the target tracking method according to the present application is shown in a schematic diagram of an implementation process. It can be understood that, Figure 2 The execution subject of the target tracking method shown is the electronic device 100. The electronic device 100 can be a terminal device or a server connected to the terminal device. As Figure 2 shown, the process includes:
[0061] S201: Obtain the i-th frame of the image to be processed in the video to be processed.
[0062] It can be understood that in various embodiments of the present application, the video to be processed may include a surveillance video in a surveillance scenario. The i-th frame of the image to be processed can be a frame of the video to be processed, and can be determined, for example, by frame decomposition and frame extraction of the video to be processed, such as extracting frames at an interval of n frames. It can be understood that for the real-time image obtained by the camera of the electronic device, real-time frame decomposition and frame extraction can be performed while obtaining.
[0063] In some embodiments, the electronic device 100 obtains the i-th frame of the image to be processed in the sequence of images to be processed. Where i is a positive integer.
[0064] S202: Perform object detection on the i-th frame of the image to be processed to obtain detection frames (bounding boxes) corresponding to at least one target object.
[0065] In some embodiments, the electronic device 100 can perform object detection on the i-th frame of the image to be processed based on an object detection algorithm (such as the YOLO neural network model) to determine at least one target object and its corresponding detection frame. For example, as Figure 3a shown, the target objects are 30a, 31a, and 32a, the detection frame corresponding to the target object 30a is 30b, the detection frame corresponding to the target object 31a is 31b, and the detection frame corresponding to the target object 32a is 32b. The target object refers to the object to be tracked by the target tracking method, usually a moving object in a video image, such as a vehicle, a pedestrian, etc. The detection frame is used to identify the position information of the target object, and the area where the detection frame is located corresponds to the area where the target object is located.
[0066] It can be understood that the detection frame is usually rectangular, and in other embodiments, the detection frame can also be of other shapes.
[0067] It can be understood that one target object can correspond to one detection frame or multiple detection frames. If the target object corresponds to multiple detection frames, the detection frame that best matches the target object can be determined by comparing the overlapping degree (such as the intersection over union IoU) of the target object and its corresponding multiple detection frames in the image as the detection frame of the target object.
[0068] It can be understood that if the i-th frame of the image to be processed is the first frame of the image to be processed, the detection box of the first frame of the image to be processed is directly marked as the tracking box of the target object. Among them, the tracking box of the target object refers to the detection box of the marked target object, which can be associated with the target object in the images of adjacent frames.
[0069] S203: Determine the tracking data of at least one target object in the i-th frame of the image to be processed based on the detection box and feature image of the i-th frame of the image to be processed, and the tracking data of the (i - 1)-th frame of the image to be processed.
[0070] It can be understood that the target tracking methods mainly include: the target tracking method based on the target tracking model of traditional technologies, such as the Kalman filtering technology, and the target tracking method based on the target tracking model of deep learning technologies, such as the Transformer algorithm.
[0071] It can be understood that in some other embodiments, the tracking data of at least one target object in the i-th frame of the image to be processed can also be determined based on the detection box and feature image of the i-th frame of the image to be processed, and the tracking data of the (i - j)-th frame of the image to be processed. Among them, j is usually an integer greater than 50 and less than 100.
[0072] Among them, the target tracking model based on the Transformer algorithm mainly associates the target object by matching (or comparing) the image features of the front and back two frames of images in the video to be processed. For example, the same mark is added to the target objects with matching image features to achieve the tracking of the target object. For example, when the user views the surveillance video and needs to view the movement trajectory of a certain target object (such as Figure 1 the pedestrian marked as "Per1" in the figure), the tracking data corresponding to the mark (such as "Per1") of the target object can be viewed.
[0073] Specifically, in some embodiments, the feature image of the i-th frame of the image to be processed can be obtained first. It can be understood that the feature image includes image features such as the color information, edge information, and brightness information of the image to be processed. For example, Figure 3bAs shown, the content of the image A to be processed includes a white car and a pedestrian, with a size of 16x16, that is, the image A to be processed includes 256 pixel points. The corresponding feature image of the image A to be processed can be the feature image F, and the size of the feature image F is 8×8, that is, the feature image F includes 64 feature values. Among them, the feature value B' in the feature image F corresponds to the image feature of the region B in the image A to be processed. The image content of the region B is a part of the car, so the feature value B' can be used to represent that the color information is white, and the edge information is the edge of the roof and the windshield, etc. The feature value B0' in the feature image F corresponds to the image feature of the region B0 in the image A to be processed. The image content of the region B0 is a part of the car, so the feature value B0' can be used to represent that the color information is white, and the edge information is the edge of the car door and the window, etc.
[0074] It can be understood that the correlation relationship between the image features (feature values) in the feature image can improve the accuracy of image recognition and matching. Therefore, in some embodiments, the feature image of the image to be processed can be extracted multiple times to extract a high-precision feature image with a correlation relationship between the image features. For example, for the above-mentioned image A to be processed, after extracting the feature image F, the feature image F can be further feature-extracted to extract the correlation relationship between the image features in the feature image F, such as the semantic relationship between the image features. For example, whether two feature values belong to the same object. For example Figure 3c As shown, further feature extraction is performed on the feature image F to obtain the feature image F'. The feature value B12 in the feature image F' corresponds to the image feature of the region B11 in the feature image F, and the feature value B12 can represent that in the feature image F, the feature value B0' and the feature value belong to the feature values of the same object. It can be understood that the feature image F and the feature image F' can be the feature images of a certain channel in the multi-channel feature images.
[0075] It can be understood that in the embodiments of the present application, the accuracy of the feature image refers to the number of image features included in the feature image. The more the number of image features included in the feature image, the higher the accuracy of the feature image. And, the smaller the size of the feature image, the less detailed information it contains and the more global information it contains. For example, Figure 3c the feature values in the feature image F' shown Figure 3b correspond to a larger region of the image to be processed than the feature values in the feature image F shown Figure 3b That is, the feature value B' in the feature image F shown corresponds to the region B of the image A to be processed, and the feature value B0' corresponds to the region B0 of the image A to be processed; Figure 3c the region corresponding to the feature value B12 in the feature image F' shown in the image A to be processed includes the region B and the region B0.
[0076] Then, obtain the tracking data of the (i - 1)-th to-be-processed image in the to-be-processed image sequence and the detection box of the i-th to-be-processed data. Based on the matching degree between the feature image of the target object corresponding to the detection box of the i-th to-be-processed image and the image feature of the target object corresponding to the tracking box of the (i - 1)-th to-be-processed image, determine the identification information of the target object corresponding to the detection box of the i-th to-be-processed image. Determine the prediction box of the i-th to-be-processed image based on the tracking box of the (i - 1)-th to-be-processed image. Among them, the tracking data includes: the tracking box, the image feature of the target object corresponding to the tracking box, and the identification information.
[0077] In some embodiments, if the feature image of the target object corresponding to the detection box of the i-th to-be-processed image matches the image feature of the target object corresponding to the tracking box of the (i - 1)-th to-be-processed image, for example, the similarity is greater than the similarity threshold, the marking information of the corresponding target object in the (i - 1)-th to-be-processed image can be used as the marking information of the corresponding target object in the i-th to-be-processed image.
[0078] In some embodiments, if the feature image of the target object corresponding to the detection box of the i-th to-be-processed image does not match the image feature of the target object corresponding to the tracking box of the (i - 1)-th to-be-processed image, new marking information can be added to the corresponding target object in the i-th to-be-processed image, or the corresponding target object in the i-th to-be-processed image can be deleted.
[0079] It can be understood that after determining the identification information of the detection box in the i-th to-be-processed image, the identified detection box can be used as the tracking box of the i-th to-be-processed image; or, the prediction box of the i-th to-be-processed image can be used as the tracking box of the i-th to-be-processed image.
[0080] The following combines Figure 4a and Figure 4b to introduce in detail the object tracking model based on the transformer algorithm and its basic principle.
[0081] Specifically, as Figure 4a and Figure 4b shown, the object tracking model based on the transformer includes a feature extraction module, a detection module (detector), an encoding module (encoder), and a decoding module (decoder).
[0082] The following takes the i-th image in the to-be-processed video as an example to illustrate the functions of each module in the object tracking model based on the transformer.
[0083] Specifically, the detection module is used to perform object detection on the i-th to-be-processed image to obtain at least one detection box corresponding to the target object.
[0084] The feature extraction module is used to extract features from the i-th frame of the image to be processed, and obtain the first feature image of the i-th frame of the image to be processed. For example, for Figure 3b the feature image F shown.
[0085] The encoding module is used to further process the feature image output by the feature extraction module to obtain a second feature image with higher precision.
[0086] It can be understood that, as described above, the first feature image generally includes image features such as color information, edge information, and brightness information. Since the correlation relationship between image features can improve the accuracy of image recognition and matching, among which, the correlation relationship between image features is used to represent the semantic relationship between feature values. For example, whether two feature values belong to the same object. Therefore, it is necessary to perform an association process on each feature value in the first feature image through the encoding module to obtain a feature image containing semantic relationships, that is, to obtain the second feature image.
[0087] It can be understood that the encoding module can be an N-layer encoder (for example, a 6-layer encoder), which will be described in detail below.
[0088] The decoding module is used to determine the tracking data of at least one target object in the i-th frame of the image to be processed based on the detection frame and feature image of the i-th frame of the image to be processed, and the historical tracking data (the tracking data of the (i - 1)-th frame of the image to be processed). And, the obtained tracking data is stored in the historical tracking data for the processing of the (i + 1)-th frame of the image to be processed.
[0089] As described above, the encoding module can adopt a 6-layer encoder. When the 6-layer encoder extracts the second feature image of the image to be processed, that is, when determining the correlation relationship between the image features in the feature image, regardless of the accuracy of the first feature image, that is, regardless of whether the image features included in the first feature image are many or few, a series of specific and relatively complex feature processing processes need to be performed. In the surveillance video, there are often a large number of images to be processed with low resolution; or, in the scenario of real-time acquisition of images to be processed (such as real-time shooting), in order to immediately display the tracking data, that is, the tracking frame and identification information, it is necessary to reduce the resolution of the acquired images to be processed, and then process the images to be processed with lower resolution. The accuracy of the first feature image obtained by feature extraction of these images to be processed with lower resolution is usually low. The number of feature values included in these first feature images with lower accuracy is small, and the correlation relationship is relatively simple. If these first feature images with lower accuracy are subjected to this series of complex feature processing, it will cause waste of computing resources and prolong the image tracking time. For example, Figure 4c According to an embodiment of the present application, a schematic diagram of the process of a 6-layer encoder processing the first feature image to obtain a second feature image containing semantic information is shown. As Figure 4cAs shown, the process includes:
[0090] First, the first-layer encoder first converts the input first feature image into a one-dimensional sequence, and performs convolution processing on the one-dimensional sequence to obtain K (key), V (value), and Q (query); where V is used to represent the first feature image, and K and Q are used to represent correlation coefficients.
[0091] Then, perform positional encoding on K and Q: After processing K and Q with sine and cosine functions of different frequencies and then adding them pixel by pixel, two K' and Q' with positional information are obtained. Input K', Q', and V into the multi-head attention model, and after being processed by the multi-head attention model, a feature image T including semantic information is obtained. This feature image T can represent the association relationship between the feature values in the i-th frame of the image to be processed. Among them, the multi-head attention model is implemented based on the bilinear interpolation algorithm.
[0092] Finally, perform post-processing on the feature image T to obtain the feature image T'. For example, fuse the feature image T with K, Q, and V and perform layer normalization (LayerNorm) to avoid information loss. Input the result of layer normalization into the feedforward neural network model (FNN) for processing, and output through residual connection to obtain the feature image T'.
[0093] Subsequently, input the feature image T' into the second-layer encoder. The second-layer encoder uses the feature image T' as the first feature image and performs the above operations until the last-layer encoder outputs a feature image. The feature image output by the last-layer encoder is used as the second feature image.
[0094] It can be understood that the second feature image contains the association relationship of the feature values, that is, the global information of the image to be processed.
[0095] It can be understood that for the transformer-based object tracking model, since the processing process corresponding to the encoding module is preset by the developer, the processing process for any-sized first feature image input is the same, that is, it is necessary to complete Figure 4c the processing process shown.
[0096] That is to say, for the same image to be processed, when the resolution of the image to be processed is relatively high, for example Figure 5aThe to-be-processed image A1 shown, the to-be-processed image A1 includes 16x16 pixel points (i.e., 256 pixel points), and the first feature image output by the feature extraction module can be the feature image F1. The feature image F1 includes 8×8 feature values (i.e., 64 feature values), and the correlation relationship between the feature values can be obtained by using the encoding module.
[0097] When the resolution of the to-be-processed image is low, for example Figure 5b The to-be-processed image A2 shown, the to-be-processed image A2 includes 4×4 pixel points (i.e., 16 pixel points), and the first feature image output by the feature extraction module can be the feature image F2. The feature image F2 includes 2×2 feature values (i.e., 4 feature values). At this time, if the encoding module is used to obtain the correlation relationship between the feature values, the above-mentioned cumbersome calculation process still needs to be carried out.
[0098] However, since when the resolution of the to-be-processed image is low, the size of the corresponding first feature image is also small, for example Figure 5b The feature image F2 shown, directly performing a convolution calculation on the feature image F2 can also obtain the correlation relationship (global information of the image) between the feature values. For example, referring to Figure 5c , performing another convolution on the feature image F2 to obtain the feature image F2', and the feature image F2' contains the correlation relationship between the feature values in the feature image F2.
[0099] In summary, for the object tracking model based on transformer, since the processing process corresponding to the encoding module is preset by the developer, the processing process for the first feature image of any size input is the same. When the resolution of the to-be-processed image is low, it is also necessary to obtain the second feature image containing semantic information through the cumbersome operation process of the encoding module, wasting a large amount of computing resources.
[0100] In view of this, in the case where the resolution of the to-be-processed image is low, in order to reduce the computing resources in the object tracking process, the embodiment of the present application provides an image tracking method. This method uses a feature fusion module based on a convolutional neural network architecture with a convolutional function to replace the above-mentioned encoding module. This feature fusion module can obtain the correlation relationship between the feature values in the feature image without performing complex operations on the input low-resolution feature image.
[0101] Specifically, in some embodiments of the present application, the feature fusion module can perform a convolution process on the first feature image obtained after the to-be-processed image is subjected to feature extraction through a convolution operator to obtain a second feature image including the correlation relationship between the image features in the first feature image.
[0102] It can be understood that in some embodiments of the present application, the functions of the above convolution operator can be obtained by training a target tracking model including a feature fusion module.
[0103] It can be understood that in the encoding module of the target tracking model based on transformers, there are complex operations such as position encoding and bilinear interpolation, which usually need to run on a central processing unit (CPU) and will occupy a large amount of CPU computing resources. And in the feature fusion module based on the convolutional neural network architecture, there are simple operations such as convolution and normalization, and circuits for processing simple operations such as convolution operations are usually set in a neural-network processing unit (NPU). Therefore, the feature fusion module based on the convolutional neural network architecture can be fully deployed on the NPU to reduce CPU computing resources.
[0104] The following will introduce the structure and training process of the above target tracking model including the feature fusion module in combination with Figures 6a to 7 ,
[0105] Exemplarily, Figure 6a According to an embodiment of the present application, a schematic structural diagram of a target tracking model including the above feature fusion module is shown. As Figure 6a shown, the target tracking model includes: a feature extraction module, a detection module, a feature fusion module, and a decoding module.
[0106] Among them, the functions of the feature extraction module and the detection module are the same as those of the corresponding modules in the above Figure 4a and Figure 4b , and will not be elaborated here.
[0107] The feature fusion module performs a convolution operation on the first feature image output by the feature extraction module based on a convolution operator to obtain N second feature images including the correlation relationships between the image features in the first feature image, where N is an integer greater than 0. Among them, the convolution operator has the function of extracting the correlation relationships between the feature values in the feature image, and the parameters of the convolution kernel in the convolution operator can be determined during the training process of the target tracking model.
[0108] Particularly, for the case where the resolution of the image to be processed is low (the resolution of the image to be processed is lower than the resolution threshold), that is, the accuracy of the first feature image is low, the process of the feature fusion module performing a convolution operation on the first feature image based on the convolution operator to obtain the second feature image may include:
[0109] The first feature image output by the feature extraction module is sampled at different granularities to obtain multiple first feature sub-images of different sizes. Then, based on the convolution operator, convolution operations are performed on the obtained first feature sub-images to obtain a second feature image with higher accuracy (including semantic information). Moreover, before performing convolution operations on the first feature sub-images with smaller sizes, feature fusion (or feature merging) can be performed between the second feature image with a larger size and the first feature sub-images with smaller sizes. It can be understood that the accuracy of the fused first feature sub-images is improved, thereby improving the accuracy of the second feature image obtained based on the fused first feature sub-images.
[0110] For example, the first feature image output by the feature extraction module is sampled at different granularities to obtain multiple first feature sub-images of different sizes. The first feature sub-image P and the first feature sub-image Q with a resolution higher than that of the first feature sub-image P are selected from the multiple first feature sub-images, and at least one convolution calculation is performed on the first feature sub-image Q to obtain the corresponding second feature image Q'. The second feature image Q' is merged with the first feature sub-image P to obtain a fused feature image K including the image features of the second feature image Q' and the first feature sub-image P. At least one convolution calculation is performed on the fused feature image K to obtain the second feature image P' corresponding to the first feature sub-image P.
[0111] The decoding module is used to determine multiple tracking sub-data of at least one target object in the corresponding i-th frame to-be-processed image respectively based on multiple second feature images, as well as historical tracking data (tracking data of the (i - 1)-th frame to-be-processed image) and the detection box of the i-th frame to-be-processed image. And the tracking sub-data with a matching degree greater than the first matching degree (or the highest coincidence degree) between the image features of the target object in the multiple tracking sub-data and the image features of the corresponding target object in the tracking data of the (i - 1)-th frame to-be-processed image is determined as the tracking data of the i-th frame to-be-processed image. Moreover, the obtained tracking data is stored in the historical tracking data for the processing of the (i + 1)-th frame to-be-processed image.
[0112] Specifically, refer to Figure 6b the structural schematic diagram of the feature fusion module shown. The feature fusion module may include a downsampling sub-module, a convolution sub-module, a fusion sub-module, and a normalization sub-module.
[0113] Among them, the downsampling sub-module is used to downsample the first feature image output by the feature extraction module to obtain a plurality of first feature sub-images with sizes smaller than the first feature image. For example, the downsampling sub-module can scale the first feature image through a pooling operation to obtain a plurality of first feature sub-images with different sizes. For example, the size of the first feature sub-image can be 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64, etc. of the size of the first feature image. The present application does not make specific limitations on the size of the first feature sub-image.
[0114] The convolution sub-module is used to perform at least one convolution operation on the first feature sub-image to obtain a feature image containing semantic information. Among them, the convolution sub-module contains a convolution operator, and the convolution operator has the function of extracting the correlation relationship between the feature values in the feature image. The parameters of the convolution kernel in the convolution operator can be determined during the model training process.
[0115] It can be understood that the sizes of the convolution kernels corresponding to the convolution operations on each first feature sub-image can be the same or different. The number of convolution operations performed on each first feature sub-image can be the same or different.
[0116] The normalization sub-module is used to perform a LayerNorm regularization operation on the feature image containing semantic information output by the convolution sub-module to obtain a corresponding second feature image.
[0117] The fusion sub-module is used to perform a fusion operation on the second feature image of the first size and the first feature sub-image of the second size to obtain a fused first feature sub-image of the second size, and input the fused first feature sub-image of the second size into the convolution sub-module. The convolution sub-module performs at least one convolution operation on the fused first feature sub-image of the second size to obtain a second feature image of the second size containing semantic information. Among them, the first size is greater than the second size.
[0118] It can be understood that the method of feature fusion can be serial feature fusion (concat), directly connecting two features. For example, if the dimensions of two input features x and y are p and q, respectively, then the dimension of the fused feature z is p + q. The method of feature fusion can also be additive fusion (add), adding / subtracting two features to obtain a fused feature. For example, for input features x and y, the fused feature z = x ± y. The method of feature fusion can also be weighted fusion, summing two features with weights to obtain a fused feature. For example, for input features x and y, the fused feature z = ax + by, where a and b correspond to the weights of the two feature values respectively. The present application does not make specific limitations on the way of feature fusion.
[0119] Exemplarily, the process of the feature fusion module further processing the feature image output by the feature extraction module to obtain a second feature image with higher precision may include:
[0120] First, the first feature image can be sampled by adjusting the sampling granularity (i.e., the size of the sampling area) of image feature extraction to obtain first feature sub-images of different sizes.
[0121] It can be understood that each sampling area of the first feature image corresponds to each feature value in the first feature sub-image.
[0122] To improve the precision of the feature image, further, a convolution operation is performed on the larger-sized feature sub-image based on a convolution kernel having the function of extracting the correlation relationship between the feature values in the feature image to obtain a corresponding larger-sized second feature image containing semantic information; then the larger-sized second feature image is fused with the smaller-sized first feature sub-image so that the smaller-sized first feature sub-image after fusion contains semantic information; and then a convolution operation is performed on the fused smaller-sized first feature sub-image based on a convolution kernel having the function of extracting the correlation relationship between the feature values in the feature image to obtain a corresponding smaller-sized second feature image. It can be understood that the smaller-sized second feature image is obtained by performing a convolution operation on the smaller-sized first feature sub-image, contains more semantic information, and can be achieved only through convolution operations and fusion operations without complex calculations.
[0123] Specifically, it includes:
[0124] The first feature image output by the feature extraction module is downsampled to obtain L first feature sub-images of different sub-sizes. Among them, L is an integer greater than 2. The size relationship from the first size to the L-th size is from large to small.
[0125] After performing a convolution operation on the first feature sub-image of the first size using a convolution kernel of a preset size (for example, 3×3), a second feature image of the first size is obtained.
[0126] After downsampling the second feature image of the first size and performing feature fusion (concat) with the first feature sub-image of the second size, a convolution operation is performed on the fused first feature sub-image of the second size using a convolution kernel of a preset size to obtain a second feature image of the first size.
[0127] By analogy, after downsampling the second feature image of the (L - 1)-th size and performing feature fusion with the first feature sub-image of the L-th size, a convolution operation is performed on the fused first feature sub-image of the L-th size using a convolution kernel of a preset size to obtain a second feature image of the L-th size.
[0128] It can be understood that, in order to avoid information loss, post-processing can also be performed on the first feature sub-image after the convolution operation, such as regularization processing (LayerNorm).
[0129] Exemplarily, refer to Figure 6c the schematic diagram of feature fusion shown.
[0130] As Figure 6c shown, after performing convolution calculation on the feature image with a size of 16x (i.e., reducing the size of the first feature image by 16 times), a regularization operation is performed to obtain a feature image 1 with a corresponding size of 16x. After downsampling the feature image with a size of 16x and fusing it with the feature image with a size of 32x (i.e., reducing the size of the first feature image by 32 times), a convolution operation and a regularization operation are performed on the fused feature image to obtain a feature image 2 with a corresponding size of 32x. After downsampling the feature image with a size of 32x and fusing it with the feature image with a size of 64x (i.e., reducing the size of the first feature image by 64 times), a convolution operation and a regularization operation are performed on the fused feature image to obtain a feature image 3 with a corresponding size of 64x.
[0131] To better understand the technical solution of the embodiments of the present application, the training process of the target tracking model shown below is introduced. Figure 6a the training process of the target tracking model shown.
[0132] Figure 7 According to the embodiments of the present application, a flowchart of the training process of a target tracking model is shown. As Figure 7 shown, the process includes:
[0133] S701: Obtain sample images of multiple consecutive frames.
[0134] In some embodiments, sample images of multiple consecutive frames are obtained, the marking information (or sample marking information) of each real target object is determined, and the real box (or sample tracking box) surrounding each real target object is determined. Among them, the sample images can be from a public dataset for video object tracking, such as the VID dataset, the DET dataset, or can also be from a video captured by a camera of an electronic device.
[0135] It can be understood that the sample images of consecutive frames can contain the same target object, so as to facilitate the subsequent training of the target tracking model based on the sample object.
[0136] It can be understood that the real boxes and marking information of each real target object can be obtained from the dataset or manually marked by developers, and the present application does not limit this.
[0137] S702: Input multiple sample images into the target tracking model to obtain the tracking data corresponding to each sample image.
[0138] In some embodiments, the target tracking model obtains at least one predicted target object and the corresponding prediction box for each sample image, as well as the first feature image corresponding to each sample image. Then, based on the first feature image, the second feature image is determined, and the tracking data corresponding to each sample image is determined. Among them, the tracking data includes the predicted tracking box (or detection tracking box) corresponding to the predicted target object, the predicted identification information (or detection marking information), and the predicted image features.
[0139] S703: Calculate the loss value.
[0140] In some embodiments, based on the shape, position, width, height and other parameters of the predicted tracking box corresponding to at least one predicted target object in each sample image and the ground truth box, the first loss value L1 is calculated. Based on the marking information of at least one predicted target object corresponding to each sample image and the marking information of the ground truth target, the second loss value L2 is calculated.
[0141] It can be understood that in some other embodiments, the first loss value L1 can also be calculated based on the shape, position, width, height and other parameters of the predicted tracking box corresponding to each second feature image and the ground truth box.
[0142] It can be understood that loss functions such as the cross-entropy loss function, Focal Loss function, and IOU LOSS function can be used to calculate the first loss value L1 and the second loss value L2.
[0143] Determine the weight α corresponding to the first loss value L1 and the weight β corresponding to the second loss value L2, and further determine the loss value Loss. It can be understood that the loss value Loss is the weighted average of the first loss value L1 and the second loss value L2.
[0144] S704: Determine whether the loss value is less than the loss threshold.
[0145] In some embodiments, if the judgment result is no, then execute step S705, and adjust the parameters of each module in the target tracking model based on the loss value and perform the next training.
[0146] In some other embodiments, if the judgment result is yes, then execute step S706, and use the parameters currently used by each module as the parameters of each module of the trained model.
[0147] S705: Adjust the parameters of each module in the target tracking model based on the loss value and perform the next training.
[0148] In some embodiments, the parameters in each module are adjusted based on the loss value. For example, the overlapping ratio threshold between the ground truth box and the predicted box in the prediction module, the parameters of the convolution kernel in the convolution operator in the feature fusion module, and so on. Then, the next training is performed based on the adjusted parameters, that is, step S702 is executed: multiple sample images are input into the target tracking model to obtain the tracking data corresponding to each sample image.
[0149] S706: Use the parameters currently used in each module as the parameters of each module of the trained model.
[0150] It can be understood that in some other embodiments, before training the target tracking model, based on the parameters in the pre-trained detection module, feature extraction module, and decoding module, the parameters in the corresponding modules of the target tracking model can be fixed, that is, there is no need to adjust them during the training process, and only the parameters in the feature fusion module need to be adjusted during the training process.
[0151] To better understand the technical solutions of the embodiments of the present application, some technical solutions of the present application will be introduced in detail below.
[0152] Figure 8 According to an embodiment of the present application, a flowchart of an image tracking method is shown. It can be understood that Figure 8 The execution subject of each step of the shown process is the electronic device 100. The electronic device 100 can be a terminal device or a server connected to the terminal device. For the sake of simplicity of description, the execution subject of each step will not be repeatedly described below when introducing Figure 8 each step of the shown process. As Figure 8 shown, the process includes but is not limited to the following steps:
[0153] S801: Obtain the i-th to-be-processed image in the to-be-processed video.
[0154] It can be understood that in each embodiment of the present application, the to-be-processed video may include a surveillance video in a surveillance scenario. The i-th to-be-processed image may be a frame of the to-be-processed video, which can be determined, for example, by frame decomposition and frame extraction of the to-be-processed video, such as extracting frames at an interval of n frames. It can be understood that for the real-time video captured by the camera of the electronic device, frame decomposition and frame extraction can be performed in real time while capturing.
[0155] In some embodiments, the electronic device 100 obtains the i-th to-be-processed image in the to-be-processed image sequence. Where i is a positive integer.
[0156] S802: Perform target detection on the i-th to-be-processed image to obtain detection boxes corresponding to at least one target object.
[0157] In some embodiments, the electronic device 100 may perform object detection on the i-th frame of the image to be processed based on an object detection algorithm (such as the YOLO neural network model), and determine at least one target object and its corresponding detection box. For example, as Figure 3a shown, the target objects are 30a, 31a, and 32a. The detection box corresponding to the target object 30a is 30b, the detection box corresponding to the target object 31a is 31b, and the detection box corresponding to the target object 32a is 32b. A target object refers to an object to be tracked by a target tracking method, usually a moving object in a video image, such as a vehicle, a pedestrian, etc. The detection box is used to identify the position information of the target object, and the area where the detection box is located corresponds to the area where the target object is located.
[0158] It can be understood that the detection box is usually rectangular. In other embodiments, the detection box may also be of other shapes.
[0159] It can be understood that one target object may correspond to one detection box or multiple detection boxes. If a target object corresponds to multiple detection boxes, the detection box that best matches the target object can be determined by comparing the overlapping degree (such as the intersection over union IoU) of the target object and its corresponding multiple detection boxes in the image as the detection box of the target object.
[0160] It can be understood that if the i-th frame of the image to be processed is the first frame of the image to be processed, the detection box of the first frame of the image to be processed is directly marked as the tracking box of the target object. The tracking box of the target object refers to the marked detection box of the target object, which can be associated with the target object in the adjacent frame of the image.
[0161] S803: Determine the feature images of multiple sizes corresponding to the i-th frame of the image to be processed.
[0162] In some embodiments, the electronic device 100 performs feature extraction on the i-th frame of the image to be processed based on a feature extraction model (such as the Residual Neural Network model ResNet) to obtain the first feature image of the i-th frame of the image to be processed. The first feature image is downsampled to obtain L first feature sub-images of different sizes. Where L is an integer greater than 2.
[0163] For example, Figure 6b the feature images with sizes of 16x (i.e., reducing the size of the first feature image by 16 times), 32x (i.e., reducing the size of the first feature image by 32 times), and 64x (i.e., reducing the size of the first feature image by 64 times) shown.
[0164] It can be understood that the downsampling method can be to perform a pooling operation on the first feature image, or to perform a convolution operation with a convolutional kernel of size m×m on the first feature image with a stride of m - 1. This application does not make specific limitations on the downsampling method.
[0165] S804: Convolve the feature images of multiple sub - sizes to obtain the feature images containing semantic information corresponding to each size.
[0166] In some embodiments, after the electronic device 100 performs a convolution operation on the first feature sub - image of the first size with a convolutional kernel of a preset size (for example, 3×3), a regularization operation (LayerNorm) is performed to obtain the second feature image of the first size.
[0167] After downsampling the second feature image of the first size and performing feature fusion with the first feature sub - image of the second size, a convolution operation is performed on the fused first feature sub - image of the second size with a convolutional kernel of a preset size, and then a regularization operation is performed to obtain the second feature image of the first size.
[0168] By analogy, after downsampling the second feature image of the (L - 1) - th size and performing feature fusion with the first feature sub - image of the L - th size, a convolution operation is performed on the fused first feature sub - image of the L - th size with a convolutional kernel of a preset size, and then a regularization operation is performed to obtain the second feature image of the L - th size.
[0169] As Figure 6c shown, after performing convolution calculation on the feature image of size 16x and then performing a regularization operation, the feature image 1 of size 16x is obtained. After downsampling the feature image of size 16x and fusing it with the feature image of size 32x, a convolution operation and a regularization operation are performed on the fused feature image to obtain the feature image 2 of size 32x. After downsampling the feature image of size 32x and fusing it with the feature image of size 64x, a convolution operation and a regularization operation are performed on the fused feature image to obtain the feature image 3 of size 64x.
[0170] It can be understood that by fusing feature images of different sizes, more refined image features corresponding to the low - resolution image to be processed can be obtained. Moreover, compared with the transformer algorithm, the convolution operation consumes less computing resources.
[0171] Furthermore, by performing feature fusion on multiple first feature sub - images obtained by downsampling the first feature image, there is no need to perform operations such as feature fusion and convolution on the first feature image, further reducing the memory space that needs to be accessed in subsequent processing, and thus reducing the bandwidth consumption.
[0172] S805: Determine the tracking frame and identification information of at least one target object in the image to be processed based on the detection frame and feature image of the i-th frame of the image to be processed, and the tracking data of the (i - 1)-th frame of the image to be processed.
[0173] In some embodiments, the electronic device 100 first determines multiple pieces of tracking sub-data of at least one target object in the corresponding i-th frame of the image to be processed respectively based on the detection frame and each second feature image of the i-th frame of the image to be processed, and the tracking data of the (i - 1)-th frame of the image to be processed.
[0174] Specifically, the electronic device 100 determines the image features corresponding to the target object of the i-th frame of the image to be processed respectively based on the second feature image and the detection frame of the i-th frame of the image to be processed, matches the image features of the target object corresponding to the detection frame of the i-th frame of the image to be processed with the image features of the target object corresponding to the tracking frame of the (i - 1)-th frame of the image to be processed, and determines the identification information of the target object corresponding to the detection frame of the i-th frame of the image to be processed. Determine the prediction frame of the i-th frame of the image to be processed based on the tracking frame of the (i - 1)-th frame of the image to be processed. Among them, the tracking sub-data includes: the tracking frame, the image features of the target object corresponding to the tracking frame, and the identification information.
[0175] In some embodiments, if the feature image of the target object corresponding to the detection frame of the i-th frame of the image to be processed matches the image features of the target object corresponding to the tracking frame of the (i - 1)-th frame of the image to be processed, for example, the similarity is greater than the similarity threshold, then the marking information of the corresponding target object in the (i - 1)-th frame of the image to be processed can be used as the marking information of the corresponding target object in the i-th frame of the image to be processed. For example, Figure 1 the target object in the (k + 1)-th frame shown matches the target object in the k-th frame in terms of features, and the target object in the (k + 1)-th frame is marked, for example, "Per1".
[0176] In some embodiments, if the feature image of the target object corresponding to the detection frame of the i-th frame of the image to be processed does not match the image features of the target object corresponding to the tracking frame of the (i - 1)-th frame of the image to be processed, then new marking information can be added to the corresponding target object in the i-th frame of the image to be processed, or the target object corresponding to the detection frame of the i-th frame of the image to be processed can be deleted. For example, Figure 1 the target object in the (k + 1)-th frame shown does not match the target object in the k-th frame in terms of features, and new marking, for example, "Per4", is added to the target object in the (k + 1)-th frame.
[0177] Then, determine the tracking sub-data with the highest degree of matching (for example, the highest degree of coincidence) between the image features of the target object in the multiple pieces of tracking sub-data and the image features of the corresponding target object in the tracking data of the (i - 1)-th frame of the image to be processed as the tracking data of the i-th frame of the image to be processed.
[0178] It can be understood that after determining the identification information of the detection box in the i-th frame of the image to be processed, the identified detection box can be used as the tracking box of the i-th frame of the image to be processed, or the prediction box of the i-th frame of the image to be processed can be used as the tracking box of the i-th frame of the image to be processed.
[0179] Specifically, the above step S806 can be implemented by a decoder module based on a transformer. The decoder module can be an N-layer decoder (for example, a 4-layer decoder). Figure 9 An embodiment according to the present application shows a schematic structural diagram of a single-layer decoder. As Figure 9 shown, the single-layer decoder includes a self-attention model, a cross-attention model, and a feed-forward neural network model. Among them, the cross-attention model is a variable attention model.
[0180] In some embodiments, the specific operation process of the variable attention model can refer to Figure 10 the schematic diagram of the interaction process of the variable attention model shown. As Figure 10 shown, the variable attention model first performs bilinear upsampling on the feature image, and then inputs the result of the bilinear upsampling, the offset value, and the reference position information into the nearest neighbor interpolation model for nearest neighbor interpolation to obtain the output (feature image). Among them, the offset value represents the position offset of the detection box in the feature image. The reference position information can be the position information of the detection box of the i-th frame of the image to be processed, or the position information of the tracking box of the (i - 1)-th frame of the image to be processed.
[0181] It can be understood that compared with bilinear interpolation, the complexity of the nearest neighbor interpolation algorithm is lower. Compared with the method of directly using the nearest neighbor interpolation algorithm to interpolate the feature image, the method of first performing bilinear upsampling on the feature image and then performing the nearest neighbor interpolation algorithm has higher accuracy of the obtained result.
[0182] In some other embodiments, the variable attention model uses a constrained sigmoid function to control the value range of the offset value of the feature image, for example, to control the offset of the feature image to be between 0 and 20. By controlling the range of the offset value, while ensuring the accuracy of the algorithm, the data reading range is reduced, and the memory occupancy is reduced.
[0183] Figure 11 An embodiment according to the present application shows a schematic diagram of the processing process of an image tracking method.
[0184] As Figure 11As shown, for the i-th frame of the image to be processed, it includes:
[0185] The detection vector (Q_yolo) obtained by the detection module (such as an object detection network) performing an embedding operation on the i-th frame of the image to be processed. The detection vector can be the pixel information of the target object in the i-th frame of the image to be processed.
[0186] The feature extraction module (such as a residual neural network) extracts the image features of the i-th frame of the image to be processed (such as generating a feature image), and inputs the image features of the i-th frame of the image to be processed into the feature fusion module. After being processed by the feature fusion module, the feature image of the i-th frame of the image to be processed is obtained and converted into a one-dimensional vector (embedding vector).
[0187] The 4-layer decoder (decoder*4) based on the tracking data of the (i - 1)-th frame image, the embedding vector output by the encoder, the prior vector (Q_det), the detection vector (Q_yolo), the placeholder vector (Q_null), and the detection box position information (xywh), obtains the features corresponding to the i-th frame image, and further obtains the result (labeling result) of the i-th frame image. It can be understood that the prior vector can be important feature information, such as image features corresponding to a face, a license plate number, etc. The detection box position information xywh represents the position and size information of the detection box corresponding to the target object. Among them, x and y are used to represent the position information of the detection box (such as the center point of the detection box) in the image, and w and h are used to represent the width information and height information of the detection box.
[0188] In some embodiments, the electronic device 100 can also control the length of the vector composed of multiple vectors input to the decoder to be the first length. Specifically, a null vector (Q_null) is added to the vectors input to the decoder as a placeholder vector, so that the vectors input to the decoder are of a fixed length, ensuring that the model does not frequently allocate / release memory space due to the change in the length of the input vectors, thereby ensuring the stability of the model.
[0189] After the post-processing module (QIM) processes the features of the i-th frame image, the tracking data of the i-th frame image is obtained. Among them, the tracking data includes: a tracking box, the image features of the target object corresponding to the tracking box, and identification information.
[0190] In some embodiments, the electronic device 100 can adopt a quantization mechanism to quantize each data in the image tracking method.
[0191] It can be understood that the quantization precision can be 8 bits, or 16 bits, or 4 bits. This application does not limit this.
[0192] It can be understood that the quantization precision of data in different processing processes can be the same or different. For example, the data in the nearest neighbor interpolation algorithm is quantized with 16 bits, and other data is quantized with 8 bits.
[0193] Specifically, as Figure 12 shown, the quantization process includes: performing 8-bit quantization on the feature image, then performing a convolution operation with a convolution kernel of 1×1 layer by layer, performing 8-bit weight quantization on each layer respectively, sequentially performing matrix multiplication on the results of the weight quantization, and finally, performing dequantization on the results after the matrix multiplication.
[0194] It can be understood that compared with a linear layer, the convolution operation with a convolution kernel of 1×1 can achieve layer-by-layer quantization, thereby improving the quantization precision.
[0195] It can be understood that in some other embodiments, according to actual requirements, the above Figure 8 shown steps can be combined, deleted, or replaced with other steps that are beneficial to achieving the purpose of this application. For example, the above step S801 and step S802 can be combined into one step, and this application does not make any restrictions here.
[0196] In summary, the image tracking method provided by the embodiments of this application can obtain more refined image features corresponding to the low-resolution image to be processed through the fusion of feature images of different sizes, and, compared with the transformer algorithm, the convolution operation consumes less computing resources.
[0197] Exemplarily, Figure 13 According to the embodiments of this application, a schematic diagram of a flowchart of a target tracking method based on traditional technology is shown.
[0198] As Figure 13 shown, the target tracking model based on traditional technology includes: a matching module, an intersection over union matching module (IoU matching), and a Kalman filter update module. Among them, the matching module includes a detector, a re-identification model (ReID model), a feature matching module, and a Kalman filter prediction module.
[0199] Specifically, as Figure 13As shown, the detector in the matching module performs object detection on the input frame image (the i-th image to be processed), obtaining detection boxes. The re-identification model in the matching module determines the identification feature (ReID feature) corresponding to the input frame image, i.e., the image feature corresponding to the target object, based on the input frame image and the detection boxes output by the detector. The feature matching module in the matching module performs feature matching between the historical tracking data (the (i - 1)-th image to be processed) and the identification feature output by the re-identification model, obtaining the minimum cosine similarity. The Kalman filter prediction module in the matching module obtains the Mahalanobis distance based on the historical tracking data and the detection boxes output by the detector. The matching module can further determine the weighted average of the target object based on the minimum cosine similarity and the Mahalanobis distance, and then determine the final score, so as to facilitate subsequent comparison and matching based on the final score and the historical tracking data (the (i - 1)-th image to be processed).
[0200] The intersection over union (IoU) matching module compares based on the historical tracking data and the final score output by the matching module, determines the tracking data corresponding to the input frame image as new tracking data, and then updates this new tracking data to the historical tracking data through the Kalman filter update module. For example, if the intersection over union is greater than the preset intersection over union threshold, it is confirmed that the comparison result is a match, and the label in the historical tracking data is used as the label of the detection box in the input frame image. If the comparison result is a non-match, a new label is determined for the detection box in the input frame image.
[0201] It can be understood that the above object tracking method based on traditional technologies requires a large number of post-processing processes, such as calculating the Mahalanobis distance, calculating the minimum cosine similarity, intersection over union matching, etc. Moreover, the intersection over union threshold often needs to be adjusted by developers according to experience, consuming a large amount of computing resources. Also, traditional technologies, such as the Kalman filter, have poor accuracy, which may lead to unstable tracking effects, such as identity switches (ID switch).
[0202] Exemplarily, Figure 14 shows a schematic diagram of an identity switch. As Figure 14As shown, the first frame image includes target object 301 and target object 302, where target object 301 corresponds to the label "01" and target object 302 corresponds to the label "02". In the second frame image, target object 301 and target object 302 blend (or intersect, overlap with each other, etc.). At this time, target object 301 corresponds to the label "01" and target object 302 corresponds to the label "02". Since target object 301 and target object 302 interacted in the second frame image, the stability of the tracking result was affected. As a result, in the third frame image, the image features of target object 301 were recognized as the image features of target object 302, and then target object 302 was labeled as "01". The image features of target object 302 were recognized as the image features of target object 301, and then target object 301 was labeled as "02". That is, compared with the first frame image, the identifiers corresponding to target object 301 and target object 302 were exchanged. And in the subsequent fourth frame image, target object 302 was still labeled as "01" and target object 301 was still labeled as "02". Moreover, in the subsequent images of the image sequence, target object 302 will continue to be labeled as "01" and target object 301 will continue to be labeled as "02", which affects the tracking of the target object.
[0203] Compared with the above Figure 13 target tracking model based on traditional technology shown, the model adopted in the embodiments of the present application does not require complex post-processing, and moreover, the stability of the tracking effect is relatively high. For example, the probability of label interlacing is relatively low.
[0204] Exemplarily, Figure 15 According to the embodiments of the present application, a schematic diagram of the effect of adopting the image tracking method provided by the embodiments of the present application is shown. As Figure 15 shown, the first frame image includes target object 301 and target object 302, where target object 301 corresponds to the label "01" and target object 302 corresponds to the label "02". In the second frame image, target object 301 and target object 302 blend (or intersect, overlap with each other, etc.). At this time, target object 301 corresponds to the label "01" and target object 302 corresponds to the label "02". Although target object 301 and target object 302 interacted in the second frame image, the stability of the tracking result was not affected. In the third frame image, target object 301 was still labeled as "01" and target object 302 was still labeled as "02". And in the subsequent fourth frame image, target object 301 was still labeled as "01" and target object 302 was still labeled as "02". There was no situation of label interlacing between target object 301 and target object 302.
[0205] Exemplarily, Figure 16The schematic diagram of the flowchart of a target tracking method based on the Transformer algorithm is shown according to an embodiment of the present application.
[0206] As Figure 16 shown, for the i-th frame of the image to be processed, it includes:
[0207] The detection module (such as an object detection network) performs an embedding operation (embed) on the i-th frame of the image to be processed, embeds the position information of each pixel point in the image to be processed into the i-th frame of the image to be processed, and obtains a detection vector (Q_yolo). The detection vector can be the image to be processed with position information.
[0208] The feature extraction module (such as a residual neural network) extracts the image features of the i-th frame of the image to be processed (such as generating a feature image), and inputs the image features of the i-th frame of the image to be processed into a 6-layer encoder (encoder * 6, encoder * 6). After being processed by the encoder, embedding values (embeddings) are obtained as the feature image of the i-th frame of the image to be processed.
[0209] Based on the tracking data of the (i - 1)-th frame image, the feature information output by the encoder, the prior vector (Q_det), the detection vector (Q_yolo), and the detection box position information (xywh), the 6-layer decoder (decoder * 6, decoder * 6) obtains the image features corresponding to the i-th frame image that contain position information and are associated with the previous frame of the image to be processed, and further obtains the result (label result) of the i-th frame image.
[0210] After the post-processing module (QIM) processes the features of the i-th frame image, the tracking data of the i-th frame image is obtained. Among them, the tracking data includes: a tracking box, the image features of the target object corresponding to the tracking box, and identification information.
[0211] It can be understood that the process of the encoder based on the Transformer algorithm to confirm the feature information is relatively complex and requires a large amount of computing resources. The model adopted in the embodiment of the present application uses a feature fusion module based on a neural network model to replace Figure 16 the 6-layer encoder shown, saving computing power.
[0212] In other embodiments, since a feature fusion module based on a convolutional neural network is used to determine the feature information of the image to be processed, the dependence on the depth of the neural network is reduced, and the levels of the decoder part can be further reduced, for example, reduced to a 4-level decoder, further saving the computing overhead.
[0213] Exemplarily, Figure 17a the schematic diagram of the effect of target tracking using a target tracking method based on traditional technology is shown.
[0214] As Figure 17a shown, the target objects 001, 005, 006, 007, 008, and 009 in the first frame image are all marked. However, since the target objects 002, 003, and 004 are occluded, they are not marked. The target objects 001, 002, 004, 005, 006, 007, and 009 in the second frame image are all marked. However, since the target object 003 is occluded, it is not marked. Also, since the target object 008 has undergone a morphological change and cannot be recognized, a new ID, i.e., 010, is assigned to this target object.
[0215] Exemplarily, Figure 17b FIG. shows a schematic diagram of the effect of target tracking using a target tracking method based on the transformer algorithm.
[0216] As Figure 17b shown, the target objects 001, 002, 003, 004, 005, 006, 007, 008, and 009 in the first frame image are all marked. The target objects 001, 002, 003, 004, 005, 006, 007, and 009 in the second frame image are all marked. However, since the target object 008 has undergone a morphological change and cannot be recognized, a new ID, i.e., 010, is assigned to this target object.
[0217] It can be understood that compared with the Figure 17a effect diagram shown, using the target tracking method based on the transformer algorithm for target tracking can accurately mark occluded target objects.
[0218] Exemplarily, Figure 17c FIG. shows a schematic diagram of the effect of target tracking using the image tracking method provided in the embodiments of the present application.
[0219] As Figure 17c shown, the target objects 001, 002, 003, 004, 005, 006, 007, 008, and 009 in the first frame image are all marked. The target objects 001, 002, 003, 004, 005, 006, 007, 008, and 009 in the second frame image are all marked.
[0220] It can be understood that compared with the Figure 17b rendering effect shown, when using the image tracking method provided by the embodiments of the present application for target tracking, it is possible to accurately mark the target object with morphological changes.
[0221] Table 1 shows a schematic table of the effects of tracking the DanceTrack dataset based on various existing tracking methods according to the embodiments of the present application.
[0222] As shown in Table 1, the tracking accuracy (HOTA) of the tracking result based on the first tracking method is 67.536902, ranking first; the detection accuracy (DetA) is 81.598418, ranking fourth; the association accuracy (AssA) is 51.909623, ranking second.
[0223] The HOTA of the tracking result based on the second tracking method is 65.001069, ranking second; the DetA is 81.974485, ranking second; the AssA is 55.775719, ranking first.
[0224] The HOTA of the tracking result based on the third tracking method is 64.584515, ranking third; the DetA is 83.757089, ranking first; the AssA is 49.908850, ranking third.
[0225] The HOTA of the tracking result based on the fourth tracking method is 62.373225, ranking fourth; the DetA is 81.179970, ranking fifth; the AssA is 48.034447, ranking fourth.
[0226] The HOTA of the tracking result based on the fifth tracking method is 55.559424, ranking fifth; the DetA is 81.686765, ranking third; the AssA is 37.871264, ranking fifth.
[0227] The HOTA of the tracking result based on the sixth tracking method is 54.576257, ranking sixth; the DetA is 80.172400, ranking seventh; the AssA is 37.240550, ranking sixth.
[0228] The HOTA of the tracking result based on the seventh tracking method is 54.438421, ranking seventh; the DetA is 80.109442, ranking eighth; the AssA is 37.082848, ranking seventh.
[0229] The HOTA of the tracking result based on the eighth tracking method is 52.573008, ranking eighth; the DetA is 80.823679, ranking sixth; the AssA is 34.280085, ranking ninth.
[0230] The HOTA of the tracking result based on the 9th tracking method is 51.479865, ranking 9th; DetA is 75.395501, ranking 11th; AssA is 34.719143, ranking 8th.
[0231] The HOTA of the tracking result based on the 10th tracking method is 50.081193, ranking 10th; DetA is 76.831998, ranking 10th; AssA is 33.530118, ranking 10th.
[0232] The HOTA of the tracking result based on the 11th tracking method is 47.016470, ranking 11th; DetA is 70.819804, ranking 12th; AssA is 31.316355, ranking 11th.
[0233] The HOTA of the tracking result based on the 12th tracking method is 41.832869, ranking 12th; DetA is 78.080086, ranking 9th; AssA is 22.599878, ranking 12th.
[0234] The HOTA of the tracking result based on the 13th tracking method is 36.479512, ranking 13th; DetA is 7860.170472, ranking 13th; AssA is 22.315846, ranking 13th.
[0235] Table 1 Schematic table of the effects of tracking the dance tracking set based on existing tracking methods
[0236] Tracking method HOTA DetA AssA 1 67.536902(1) 81.974485(2) 55.775719(1) 2 65.001069(2) 81.598418(4) 51.909623(2) 3 64.584515(3) 83.757089(1) 49.908850(3) 4 62.373225(4) 81.179970(5) 48.034447(4) 5 55.559424(5) 81.686765(3) 37.871264(5) 6 54.576257(6) 80.172400(7) 37.240550(6) 7 54.438421(7) 80.109442(8) 37.082848(7) 8 52.573008(8) 80.823679(6) 34.280085(9) 9 51.479865(9) 76.831998(10) 34.719143(8) 10 50.081193(10) 75.395501(11) 33.530118(10) 11 47.016470(11) 70.819804(12) 31.316355(11) 12 41.832869(12) 78.080086(9) 22.599878(12) 13 36.479512(13) 60.170472(13) 22.315846(13)
[0237] By tracking the dance tracking set using the image tracking method provided in this application, the obtained HOTA is 64.97, Det is 76.67, and AssA is 55.26. Compared with the tracking results of various tracking methods in Table 1, when tracking the dance tracking set using the image tracking method provided in this application, the HOTA ranks third. That is, the image tracking method provided in this application can improve the accuracy of the tracking result on the basis of reducing computing resources.
[0238] If only the feature fusion module is used to replace the encoding module in the object tracking model based on the Transformer algorithm, and then the dance tracking set is tracked, the obtained HOTA is 62.88, Det is 77.18, and AssA is 51.42. Compared with the tracking results of various tracking methods in Table 1, when the image tracking method provided by this application is used to track the dance tracking set, the HOTA ranks fourth. That is, if only the feature fusion module is used to replace the encoding module in the object tracking model based on the Transformer algorithm, the accuracy of the tracking result can also be improved on the basis of reducing computing resources.
[0239] To better understand the technical solution of the embodiments of this application, the structure of the electronic device involved in this application is introduced below with reference to the accompanying drawings.
[0240] Exemplarily, Figure 18 According to the embodiments of this application, a schematic structural diagram of an electronic device is shown.
[0241] As Figure 18 shown, the electronic device may include a processor 110, a memory 120, an interface module 130, a power module 140, a wireless communication module 150, a mobile communication module 160, an audio module 170, a sensor module 180, a button 190, a camera 191, and a display screen 192, etc.
[0242] It can be understood that the structure schematically shown in the embodiments of this application does not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0243] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a GPU, an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. The processor 110 may be used to execute the image tracking method provided by the embodiments of this application.
[0244] The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of instruction fetching and execution.
[0245] A memory can also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the said memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0246] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0247] The power module 140 is used to receive a charging input from a charger to supply power to the processor 110, the memory 120, the display screen 192, the wireless communication module 150, etc. In some other embodiments, the power module 140 may also be provided in the processor 110.
[0248] The wireless communication module 150 can provide solutions for wireless communications applied to the terminal device 100, including wireless local area networks (WLANs) (such as Wi-Fi networks), Bluetooth, global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 150 can be one or more devices integrating at least one communication processing module.
[0249] The mobile communication module 160 can provide solutions for wireless communications applied to the terminal device 100, including 2G / 3G / 4G / 5G, etc. The mobile communication module 160 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc.
[0250] The electronic device implements the display function through the GPU, the display screen 192, and the application processor, etc. The GPU is a microprocessor for image tracking, connected to the display screen 192 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change the display information.
[0251] The display screen 192 is used to display images. The display screen 192 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 192, where N is a positive integer greater than 1. In some embodiments, the display screen 192 can display the first reconstructed image, the second reconstructed image, etc.
[0252] The electronic device can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0253] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element, where the optical signal is converted into an electrical signal. The camera's photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin tone of the image through algorithms. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 193.
[0254] The camera 193 is used to capture still images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1. In some embodiments, the electronic device obtains the image to be processed through the camera 193.
[0255] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals.
[0256] The video codec is used to compress or decompress digital videos. The electronic device can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0257] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process the input information and can also continuously learn on its own. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc.
[0258] The memory 120 can be used to store computer-executable program codes, and the executable program codes include instructions. The memory 120 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.). In addition, the memory 120 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the terminal device 100 by running the instructions stored in the memory 120, and / or the instructions stored in the memory provided in the processor. In some implementation examples, the processor 110 executes the image tracking method provided in the embodiments of the present application by running the instructions stored in the memory 120.
[0259] The electronic device 100 can implement an audio function through the audio module 170, and an application processor, etc.
[0260] The sensor module 180 can include a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.
[0261] The keys 190 include a power-on key, a volume key, etc. The keys 190 can be mechanical keys or touch keys. The terminal device 100 can receive key inputs and generate key signal inputs related to the user settings and function controls of the terminal device 100.
[0262] It can be understood that Figure 18 The structure of the illustrated electronic device 100 is only an example. In some other embodiments, the electronic device 100 can include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.
[0263] It can be understood that in some embodiments, the encoding end and the decoding end can be two electronic devices. Figure 11 The illustrated electronic device can be the encoding end or the decoding end. In some other embodiments, the encoding end and the decoding end can also be in the same electronic device, such as Figure 11 the illustrated electronic device. The present application does not limit this.
[0264] The embodiments of the present application also provide a program product, which includes instructions that can cause an electronic device to implement the image tracking method provided by the embodiments of the present application when the instructions are executed by the electronic device.
[0265] The embodiments of the present application also provide a chip. The chip system includes a processing circuit and a storage medium, and computer program code is stored in the storage medium; when the computer program code is executed by the processing circuit, the image tracking method provided by the embodiments of the present application is implemented.
[0266] The embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.
[0267] The program code can be applied to the input instructions to execute the various functions described in the present application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of the present application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0268] The program code can be implemented in a high-level procedural language or an object-oriented programming language to communicate with the processing system. When needed, the program code can also be implemented in assembly language or machine language. In fact, the mechanism described in the present application is not limited to the scope of any specific programming language. In any case, the language can be a compiled language or an interpreted language.
[0269] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more transient or non-transient machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions may be distributed via a network or via other computer-readable media. Thus, machine-readable media can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, floppy disks, optical disks, optical discs, magneto-optical discs, read only memory (ROM), random access memory (RAM), erasable programmable read only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in electrical, optical, acoustic, or other forms using the Internet. Thus, machine-readable media includes any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0270] In the drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0271] It should be noted that the various units / modules mentioned in the device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation manner of these logical units / modules themselves is not the most important. The combination of the functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, to highlight the innovative part of the present application, the above device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application, which does not mean that there are no other units / modules in the above device embodiments.
[0272] It should be noted that in the examples and the description of this patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0273] Although this application has been illustrated and described by reference to certain preferred embodiments thereof, those of ordinary skill in the art should understand that various changes in form and detail may be made therein without departing from the spirit and scope of this application.
Claims
1. An image tracking method, characterized in that: include: Get the i-th frame image in the video to be detected, where i is greater than 1; Acquire a first feature image of the i-th frame image, wherein the first feature image includes a plurality of image feature values; Performing convolution calculation on the first feature image to obtain N second feature images, where N is an integer greater than 0, and the second feature image includes: image feature values representing association relationships between multiple image feature values in the first feature image; Based on the N second feature images, first tracking data of the i-th frame image is determined.
2. The method according to claim 1, characterized in that The performing convolution calculation on the first feature image to obtain N second feature images includes: Based on the convolution operator in the feature fusion module, at least one convolution calculation is performed on the first feature image to obtain N second feature images.
3. The method according to claim 1, characterized in that The performing convolution calculation on the first feature image to obtain N second feature images includes: Corresponding to the resolution of the i-th frame image being lower than the resolution threshold, the first feature image is downsampled respectively using a plurality of different sampling granularities to obtain a plurality of first feature sub-images with different resolutions; At least one convolution calculation is performed on each of the first characteristic sub-images with different resolutions to obtain the second characteristic images corresponding to the multiple first characteristic sub-images.
4. The method according to claim 3, characterized in that: The step of performing at least one convolution calculation on each of the first feature sub-images with different resolutions to obtain the second feature images corresponding to the plurality of first feature sub-images includes: The second characteristic image P' of the first characteristic sub-image P in the plurality of first characteristic sub-images is calculated as follows: Selecting a first feature sub-image P and a first feature sub-image Q having a higher resolution than the first feature sub-image P from the plurality of first feature sub-images, and performing at least one convolution calculation on the first feature sub-image Q to obtain a corresponding second feature image Q'; Merging the second feature image Q' and the first feature sub-image P to obtain a fused feature image K including the image features of the second feature image Q' and the first feature sub-image P; Perform at least one convolution calculation on the fused feature image K to obtain a second feature image P' corresponding to the first feature sub-image P.
5. The method according to claim 2 or 3, characterized in that: Also includes: performing a normalization operation on the second feature image; and Determining first tracking data of the i-th frame image based on the N second feature images includes: Based on the second feature image after the normalization operation, the tracking data of the i-th frame image is determined.
6. The method according to claim 1, characterized in that The first tracking data includes a tracking frame of at least one first target object in the i-th frame image and identification information corresponding to each first target object.
7. The method according to claim 6, characterized in that The determining the first tracking data of the i-th frame image based on the N second feature images includes: Determine the identification information of the second target object in the at least one first target object by: Determine the image features of the second target object based on the N second feature images, and obtain third tracking data corresponding to the i-1th frame image in the video to be detected, wherein the third tracking data includes: image features corresponding to at least one third target object in the i-1th frame image and identification information corresponding to each of the third target objects; comparing the image features of the second target object with the image features of each of the third target objects; The matching degree between the image feature of a fourth target object among the at least one third target object and the image feature of the second target object is greater than the first matching degree, and the identification information of the fourth target object is determined as the identification information of the second target object.
8. The method according to claim 7, characterized in that The determining the image feature of the second target object based on the N second feature images includes: Bilinear sampling and nearest neighbor interpolation calculation are performed on image feature values in areas corresponding to the second target object in the N second feature images to obtain image features of the second target object.
9. A target tracking model, characterized in that: include: A feature extraction module, used to obtain an i-th frame image in a video to be detected, and obtain a first feature image of the i-th frame image, wherein the first feature image includes a plurality of image feature values, and i is greater than 1; a feature fusion module, performing convolution calculation on the first feature image to obtain N second feature images, wherein N is an integer greater than 0, and the second feature image includes: an image feature value representing a correlation relationship between multiple image feature values in the first feature image; A decoding module is used to determine the first tracking data of the i-th frame image based on the N second feature images.
10. The model according to claim 9, characterized in that Also includes: A normalization module is used to perform a normalization operation on the second feature image.
11. A target tracking model training method, characterized in that: include: Acquire sample data, where the sample data includes sample images of a plurality of consecutive frames, sample marking information of a target object in each sample image, and a sample tracking frame surrounding each target object; Inputting the sample images of the plurality of continuous frames into a target tracking model to obtain tracking data corresponding to each sample image, wherein the target tracking model includes a feature extraction module, a detection module, a feature fusion module and a decoding module, and the tracking data includes a detection tracking frame and detection mark information; According to the tracking data and the sample marking information and the sample tracking frame in the sample data, the parameters of each module in the target tracking model are adjusted.
12. A readable medium, characterized in that The readable medium stores instructions, which, when executed on an electronic device, enable the electronic device to execute the method according to any one of claims 1 to 8.
13. An electronic device, characterized in that: include: a memory for storing instructions to be executed by one or more processors of the electronic device, and The processor is one of the processors of the electronic device, and is used to execute the method according to any one of claims 1 to 8.