Target Tracking Method, Device, Equipment and Computer Readable Storage Medium

By using the target's attribute information to determine the extraction model and extracting more reasonable and accurate feature points, the problem of degradation of target tracking accuracy in complex backgrounds is solved, and higher tracking accuracy and stability are achieved.

CN114170267BActive Publication Date: 2025-06-13YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010949864.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-10
Publication Date
2025-06-13
Estimated Expiration
2040-09-10

AI Technical Summary

Technical Problem

The existing target tracking algorithms easily select invalid pixel points as feature points in complex backgrounds, resulting in a decrease in tracking accuracy.

Method used

By analyzing the target's attribute information (such as occlusion information, heading angle and category information), the extraction model is determined, and the model is used to extract feature points within the detection area of ​​the target to ensure that the feature points are distributed in the center of the target, and the characteristics are obvious and unobstructed areas.

Benefits of technology

Improve the accuracy and accuracy of target tracking, reduce errors due to background complexity, and enhance tracking stability in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114170267B_ABST
    Figure CN114170267B_ABST
Patent Text Reader

Abstract

The present application provides a target tracking method, apparatus, device, and computer-readable storage medium, which can be used for assisted driving and autonomous driving. The method includes: detecting a target in a first image and determining a first feature point set according to an extraction model; the first feature point set is distributed within a detection region, the detection region includes a first region and a second region other than the first region, and the number of feature points of the target in the first region is greater than the number of feature points of the target in the second region; determining a second feature point set of the target in a second image according to the first feature point set, the second image being the next frame image of the first image; and determining a position region of the target in the second image according to the second feature point set. This method can improve the accuracy of target tracking, enhance the advanced driver assistance system (ADAS) capabilities of the terminal in autonomous or assisted driving, and can be applied to vehicle networking, such as vehicle-to-everything (V2X), long-term evolution for vehicle-to-vehicle communication (LTE-V), vehicle-to-vehicle (V2V), etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, specifically to the field of computer vision, and particularly to a method, apparatus, device, and computer-readable storage medium for object tracking. Background Art

[0002] With the development of computer vision theory and technology, object tracking has become one of the current hot topics in computer vision research and has been widely applied in fields such as video surveillance, virtual reality, human-computer interaction, and unmanned driving. Its purpose is to analyze the video or image sequence obtained by the sensor, identify the target of interest from the background of each image, extract the information of the target, and thus achieve continuous and accurate tracking of the target.

[0003] Currently, the optical flow method is one of the most commonly used object tracking algorithms. Optical flow is the instantaneous velocity of the pixels of a moving object in space on the imaging plane. The optical flow method uses the change of pixel intensity in the time domain and the correlation between adjacent frames in the image sequence to calculate the motion information of the object between adjacent frames. However, in practical applications, when selecting the feature points for object tracking, due to the complexity and variability of the background in the video or image, to a certain extent, invalid pixel points are also regarded as feature points, resulting in a decrease in tracking accuracy.

[0004] Therefore, how to effectively improve the accuracy of object tracking is an urgent problem to be solved. Summary of the Invention

[0005] The present application provides a method, apparatus, device, and computer-readable storage medium for object tracking, which can improve the accuracy of object tracking.

[0006] In a first aspect, the present application provides a method for object tracking, which may include:

[0007] Detect the target in the first image to obtain the detection area of the target and the attribute information of the target; determine the extraction model according to the attribute information of the target; determine the first feature point set according to the extraction model; determine the second feature point set of the target in the second image according to the first feature point set; determine the position area of the target in the second image according to the second feature point set; wherein, the first feature point set is distributed in the detection area of the target, the detection area includes a first area and a second area other than the first area, the number of feature points of the target in the first area is greater than the number of feature points of the target in the second area; the second image is the next frame image of the first image.

[0008] In the method provided by the first aspect, when selecting target tracking feature points, an extraction model is determined according to the attribute information of the target, and this extraction model is used to extract feature points, which can extract more reasonable and accurate feature points. The feature points are distributed in the central area, the area with obvious features, and the unoccluded area, making the distribution of feature points more reasonable and accurate, thereby improving the accuracy of target tracking.

[0009] In a possible implementation manner, the above-mentioned attribute information of the target includes one or more of the occlusion information of the target, the heading angle of the target, and the category information of the target. In this manner, through the occlusion information of the target, the feature points extracted by the extraction model can be more distributed in the unoccluded area of the target to improve the rationality of the feature point distribution; through the heading angle of the target, it can be determined that the feature points extracted by the extraction model are more distributed in the area with more obvious features to improve the accuracy of the feature point distribution; through the category information of the target, the feature points extracted by the extraction model can be specifically distributed in the area of the target. Furthermore, the accuracy of target tracking can be improved.

[0010] In a possible implementation manner, determining the extraction model according to the above-mentioned attribute information of the target includes: determining the central position of the extracted feature points according to the above-mentioned attribute information of the target. In this manner, by changing the parameters of the central position of the feature points extracted by the extraction model, the rationality of the feature point distribution is improved, thereby improving the accuracy of target tracking. And by using the extraction model to extract feature points, the extraction of feature points is independent of the size and the number of pixels of the target, which can improve the running efficiency of the method so as to be applied to the scenario of real-time target tracking with strong real-time performance.

[0011] In a possible implementation manner, the occlusion information of the target includes the occlusion ratio of the target and the occlusion area information of the target; in this case, determining the central position of the extracted feature points according to the attribute information of the target includes: obtaining the initial central position of the extracted feature points, and determining the adjustment distance corresponding to the occlusion ratio of the target according to the corresponding relationship between the occlusion ratio and the adjustment distance, and determining the central position of the extracted feature points according to the initial central position of the extracted feature points, the adjustment distance corresponding to the occlusion ratio of this target, and the opposite direction of the direction corresponding to the occlusion area information. In this manner, the initial central position of the extracted feature points is adjusted to a position closer to the center of the unoccluded part, so that the extracted feature points are more distributed in the unoccluded area, which is beneficial to improving the accuracy of target tracking.

[0012] In a possible implementation, determining the central position of the feature points to be extracted according to the attribute information of the target includes: obtaining the initial central position of the feature points to be extracted; determining the feature region of the target according to the heading angle of the target; and determining the central position of the feature points to be extracted according to the initial central position of the feature points to be extracted, a preset adjustment distance, and the direction of the center of the feature region. Wherein, the feature region is a region that matches the preset feature information except for the background region. In this way, the orientation of the target can be determined through the heading angle, the region where the target features are obvious can be determined according to the orientation of the target, and then the central position of the extracted feature points can be adjusted to be closer to the position of the center of the region where the target features are more obvious, so that the feature points can be distributed in the region where the target features are more obvious, which is beneficial to improving the accuracy of target tracking.

[0013] In a possible implementation, determining the density of the extracted feature points distributed in each region of the target detection region according to the attribute information of the target includes: determining the density of the extracted feature points distributed in each region of the target detection region according to the attribute information of the target; and determining an extraction model according to the density of the extracted feature points distributed in each region of the target detection region. In this way, the parameters of the model are determined according to the density of each region in the target detection region, so as to determine an extraction model including the determined parameters. Thus, it can be ensured that when the total number of feature points is certain, the extracted feature points are more densely distributed in the region where the center of the target is located and more dispersed and sparse in other regions, thereby ensuring the effect of target tracking. Moreover, the efficiency of extracting feature points through the model is high and the real-time performance is strong.

[0014] In a possible implementation, determining the density of the extracted feature points distributed in each region of the target detection region according to the attribute information of the target includes: obtaining the initial density of the extracted feature points distributed in each region of the target detection region, determining the first adjustment parameter corresponding to the category information of the target according to the correspondence between the category information and the adjustment parameter, and then determining the density of the extracted feature points distributed in each region of the target detection region according to the initial density of the extracted feature points distributed in each region of the target detection region and the first adjustment parameter corresponding to the category information of the target. In this way, different categories of targets have different distributions of feature points to meet the requirements of targets with different category information; the adjustment parameters determined by different category information are different, so the obtained extraction models are different, ensuring the tracking accuracy of targets with different category information.

[0015] In a possible implementation manner, according to the attribute information of the target, determine the number of feature points extracted in each region of the detection region of the target; according to the number of feature points extracted in each region of the detection region of the target, determine the extraction model. In this manner, when the total number of feature points is fixed, the number of feature points in each region can be set according to different types of targets, so as to meet the tracking requirements of different types of targets; the feature points are extracted by the model, with high running efficiency and strong real-time performance.

[0016] In a possible implementation manner, according to the attribute information of the target, determining the number of feature points extracted in each region of the detection region of the target includes: obtaining the initial number of feature points extracted in each region of the detection region of the target; determining the second adjustment parameter corresponding to the category information of the target according to the corresponding relationship between the category information and the adjustment parameter; determining the number of feature points extracted in each region of the detection region of the target according to the initial number of feature points extracted in each region of the detection region of the target and the second adjustment parameter corresponding to the category information of the target. In this manner, by using the second adjustment parameter corresponding to the category information to determine the number of feature points in each region of the detection region of the target, the method of the present application can be adapted to the target tracking scenarios of different types of information, with wide applicability; and it can make the number of feature points distributed in the region where the center of the target is located more, and the number of feature points distributed in other regions less, making the distribution of feature points more accurate and improving the accuracy of target tracking.

[0017] In a possible implementation manner, determining the second feature point set of the target in the second image according to the feature point set may be to perform two optical flow estimations, which are forward estimation and backward estimation respectively: perform forward estimation on all feature points in the first feature point set according to the pyramid optical flow algorithm, and determine the feature points corresponding to each feature point in the first feature point set in the second image to obtain the third feature point set; perform backward estimation on all feature points in the third feature point set according to the pyramid optical flow algorithm, and determine the feature points corresponding to each feature point in the third feature point set in the first image to obtain the fourth feature point set; determine the second feature point set according to the first feature point set, the third feature point set and the fourth feature point set. In this manner, using the feature point set estimated backward to screen out the feature points with large optical flow estimation errors in the forward estimation can obtain a feature point set with higher estimated accuracy, achieving the purpose of improving the accuracy of target tracking.

[0018] In a possible implementation, obtaining a second set of feature points according to the first set of feature points, the third set of feature points, and the fourth set of feature points includes: determining the distance between each feature point in the first set of feature points and the corresponding feature point in the fourth set of feature points; removing the feature points in the third set of feature points corresponding to distances greater than a preset threshold to obtain the second set of feature points. In this method, since the greater the distance, the greater the error estimated by the optical flow. Therefore, removing the feature points with large errors to obtain the second set of feature points can reduce the errors generated during the forward estimation and backward estimation of the feature points, which is beneficial to improving the accuracy of tracking.

[0019] In a possible implementation, determining the position area of the target in the second image according to the second set of feature points includes: determining the weight of the distance mapping corresponding to each feature point in the second set of feature points according to the mapping relationship between the distance and the weight; determining the position area of the target in the second image according to the weight of the distance mapping corresponding to each feature point in the second set of feature points and the detection area of the target in the first image. In this method, when the distance between the corresponding feature points in the fourth set of feature points obtained by the backward estimation and the first set of feature points in the first image is larger, the error is larger. Therefore, it can be set that the smaller the weight of the mapping when the distance is larger, so as to improve the accuracy of determining the position area of the target in the second image and reduce the errors generated during the forward estimation and backward estimation processes.

[0020] In a possible implementation, determining the position area of the target in the second image according to the weight corresponding to each feature point in the second set of feature points and the detection area of the target in the first image includes: determining the central position of the target in the second image according to the position and the corresponding weight of each feature point in the second set of feature points; determining the size and shape of the detection area; determining the position area of the target in the second image according to the central position of the target in the second image, the size and shape of the detection area. In this method, the central position of the target in the second image is determined by the weighted average method, and it is assumed that the size and shape of the detection area of the target do not change in adjacent two frames. The position area of the target is determined according to the central position of the target in the second image, the size and shape of the detection area of the target. It can use a more accurate position area to track the target when performing the target tracking of the next frame image of the second image, so as to continuously optimize the accuracy and stability of the target tracking.

[0021] In a possible implementation, inputting the first image into a preset target detection model to obtain a detection result, where the detection result includes the detection area of the target and the attribute information of the target. In this method, using the preset target detection model can improve the operation efficiency of the entire algorithm and can meet the real-time requirements as much as possible.

[0022] Second aspect, an object tracking device according to an embodiment of the present application, the device may include:

[0023] A detection unit, configured to detect an object in a first image to obtain a detection area of the object and attribute information of the object; a first determination unit, configured to determine an extraction model according to the attribute information of the object; a second determination unit, configured to determine a first feature point set according to the extraction model; the first feature point set is distributed in the detection area of the object, the detection area includes a first area and a second area other than the first area, and the number of feature points of the object in the first area is greater than the number of feature points of the object in the second area; a third determination unit, configured to determine a second feature point set of the object in a second image according to the first feature point set, where the second image is the next frame image of the first image; a fourth determination unit, configured to determine a position area of the object in the second image according to the second feature point set.

[0024] In a possible implementation manner, the attribute information of the object includes one or more of occlusion information of the object, a heading angle of the object, and category information of the object.

[0025] In a possible implementation manner, the first determination unit is specifically configured to: determine a central position of the extraction feature points according to the attribute information of the object; and determine an extraction model according to the central position of the extraction feature points.

[0026] In a possible implementation manner, the occlusion information of the object includes an occlusion ratio of the object and occlusion area information of the object; the first determination unit is specifically configured to: obtain an initial central position of the extraction feature points; determine an adjustment distance corresponding to the occlusion ratio of the object according to a correspondence between the occlusion ratio and the adjustment distance; and determine the central position of the extraction feature points according to the initial central position of the extraction feature points, the adjustment distance corresponding to the occlusion ratio of the object, and a direction opposite to a direction corresponding to the occlusion area information.

[0027] In a possible implementation manner, the first determination unit is specifically configured to: obtain an initial central position of the extraction feature points; determine a feature area of the object according to the heading angle of the object, where the feature area is an area that matches preset feature information except for a background area; and determine the central position of the extraction feature points according to the initial central position of the extraction feature points, a preset adjustment distance, and a direction of a center of the feature area.

[0028] In a possible implementation, the first determining unit is specifically configured to: determine the density of the extracted feature points distributed in each region of the detection region of the target according to the attribute information of the target; determine an extraction model according to the density of the extracted feature points distributed in each region of the detection region of the target.

[0029] In a possible implementation, the first determining unit is specifically configured to: obtain the initial density of the extracted feature points distributed in each region of the detection region of the target; determine a first adjustment parameter corresponding to the category information of the target according to the correspondence between the category information and the adjustment parameter; determine the density of the extracted feature points distributed in each region of the detection region of the target according to the initial density of the extracted feature points distributed in each region of the detection region of the target and the first adjustment parameter corresponding to the category information of the target.

[0030] In a possible implementation, the first determining unit is specifically configured to: determine the number of the extracted feature points distributed in each region of the detection region of the target according to the attribute information of the target; determine an extraction model according to the number of the extracted feature points distributed in each region of the detection region of the target.

[0031] In a possible implementation, the first determining unit is specifically configured to: obtain the initial number of the extracted feature points distributed in each region of the detection region of the target; determine a second adjustment parameter corresponding to the category information of the target according to the correspondence between the category information and the adjustment parameter; determine the number of the extracted feature points distributed in each region of the detection region of the target according to the initial number of the extracted feature points distributed in each region of the detection region of the target and the second adjustment parameter corresponding to the category information of the target.

[0032] In a possible implementation, the third determining unit is specifically configured to: perform forward estimation on all the feature points in the first feature point set according to the pyramid optical flow algorithm, and determine the corresponding feature points of each feature point in the first feature point set in the second image to obtain a third feature point set; perform backward estimation on all the feature points in the third feature point set according to the pyramid optical flow algorithm, and determine the corresponding feature points of each feature point in the third feature point set in the first image to obtain a fourth feature point set; determine the second feature point set according to the first feature point set, the third feature point set, and the fourth feature point set.

[0033] In a possible implementation manner, the third determination unit is specifically configured to: determine the distance between each feature point in the first feature point set and the corresponding feature point in the fourth feature point set; remove the feature points in the third feature point set corresponding to the distances greater than a preset threshold to obtain the second feature point set.

[0034] In a possible implementation manner, the fourth determination unit is specifically configured to: determine the weight of the distance mapping corresponding to each feature point in the second feature point set according to the mapping relationship between the distance and the weight; determine the position area of the target in the second image according to the weight of the distance mapping corresponding to each feature point in the second feature point set and the detection area of the target in the first image.

[0035] In a possible implementation manner, the fourth determination unit is specifically configured to: determine the central position of the target in the second image according to the position and the corresponding weight of each feature point in the second feature point set; determine the size and shape of the detection area; determine the position area of the target in the second image according to the central position of the target in the second image, the size and shape of the detection area.

[0036] In a possible implementation manner, the detection unit is specifically configured to: input the first image into a preset target detection model to obtain a detection result, where the detection result includes the detection area of the target and the attribute information of the target.

[0037] In a third aspect, an embodiment of the present application further provides a computer device, which may include a memory and a processor. The memory is used to store a computer program that supports the device to execute the above method. The computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of the first aspects above.

[0038] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium. The computer storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the method according to any one of the first aspects above.

[0039] In a fifth aspect, an embodiment of the present application further provides a computer program, and the computer program includes computer software instructions. When the computer software instructions are executed by a computer, the computer is caused to execute any one of the optical flow estimation methods according to any one of the first aspect or the second aspect.

[0040] In a sixth aspect, the present application further provides a chip, and the chip is used to implement the method according to the first aspect or any possible design of the first aspect. Brief Description of the Drawings

[0041] Figure 1 It is a schematic structural diagram of an image pyramid.

[0042] Figure 2 It is a schematic diagram of the architecture of an object tracking system provided by an embodiment of the present application.

[0043] Figure 3 It is a schematic flowchart of a method for tracking a graph object provided by an embodiment of the present application.

[0044] Figure 4a It is a schematic diagram of an application scenario of a method for tracking an object provided by an embodiment of the present application.

[0045] Figure 4b It is a schematic diagram of a first image provided by an embodiment of the present application.

[0046] Figure 5 It is a schematic diagram of a model of a two-dimensional Gaussian distribution provided by an embodiment of the present application.

[0047] Figure 6a It is a schematic diagram of forward estimation and backward estimation provided by an embodiment of the present application.

[0048] Figure 6b It is a schematic diagram of the determination of a location area provided by an embodiment of the present application.

[0049] Figure 7a It is a schematic diagram of adjusting the center position of feature points according to an expectation provided by an embodiment of the present application.

[0050] Figure 7b It is a schematic diagram of adjusting the center position of feature points according to the vehicle heading angle provided by an embodiment of the present application.

[0051] Figure 8a It is a schematic diagram of determining an extraction model according to the density of the distribution of feature points of an object in different regions provided by an embodiment of the present application.

[0052] Figure 8b It is another schematic diagram of determining an extraction model according to the density of the distribution of feature points of an object in different regions provided by an embodiment of the present application.

[0053] Figure 8c It is a schematic diagram of determining an extraction model according to the number of feature points in different regions to determine the extraction model provided by an embodiment of the present application.

[0054] Figure 8d It is yet another schematic diagram of determining an extraction model according to the number of feature points in different regions to determine the extraction model provided by an embodiment of the present application.

[0055] Figure 9 This is a schematic structural diagram of an object tracking device provided by an embodiment of the present application.

[0056] Figure 10 This is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0057] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0058] In the description of the present application, the terms "first" and "second" and the like in the specification and drawings of the present application are used to distinguish different objects or to distinguish different treatments of the same object, rather than to describe a specific order of the objects. In addition, the terms "including" and "having" and any variations thereof mentioned in the description of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. It should be noted that in the embodiments of the present application, words such as "exemplarily" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design method described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being superior or more advantageous than other embodiments or design solutions. Rather, the use of words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner. In the embodiments of the present application, "A and / or B" means both A and B, and either A or B. "A, and / or B, and / or C" means any one of A, B, and C, or means any two of A, B, and C, or means A and B and C. Next, the technical solutions in the present application will be described with reference to the accompanying drawings.

[0059] First, some terms in the present application will be explained to facilitate the understanding of those skilled in the art.

[0060] 1. Feature point set

[0061] The feature point set can also be described as a set of feature points. It is a set of pixel points screened from an image for object tracking. Among them, a certain feature point can also represent the gray value at the image position where the pixel point is located.

[0062] In the present application, a first set of feature points is determined according to an extraction model. Exemplarily, taking the extraction model as a truncated Gaussian model, the set of feature points is a set of multiple pixel points that satisfy the Gaussian distribution extracted from an image. Among them, each feature point carries the position of the image where the feature point is located and the gray value of the pixel point corresponding to the feature point.

[0063] 2. Pyramid optical flow algorithm

[0064] Optical flow is the instantaneous velocity of the pixels of a spatially moving object on the observed imaging plane. It is a method that uses the change of pixels in the time domain in an image sequence and the correlation between adjacent frame images to find the corresponding relationship between the previous frame image and the current frame image, so as to calculate the motion information of the object between adjacent frame images. Generally speaking, optical flow is generated by the movement of foreground objects in the scene itself, the movement of the camera, or the combined movement of both. It should be noted that the pyramid optical flow algorithm in the embodiments of the present application realizes object tracking by estimating the optical flow of the object between two images, and can be applied to fields such as unmanned driving, motion detection, action capture and recognition, and augmented reality (AR).

[0065] For the optical flow estimation algorithm, the optical flow estimation algorithm used in the present application is the sparse optical flow estimation algorithm (lucas-kanade method, LK), which constructs an optical flow equation based on three basic assumptions. Among them, the three basic assumptions are:

[0066] (1) Brightness constancy: The pixel brightness of an object in an image does not change between consecutive frames, that is, the pixels of the image do not change when they appear to move from frame to frame. For a grayscale image, this means that the gray value of the pixel does not change with the tracking of the frame.

[0067] (2) Spatial consistency: Adjacent pixels have similar motions, that is, adjacent pixels in the current frame should also be adjacent in the next frame.

[0068] (3) Short-distance (short-time) motion: The time between adjacent frames is short enough that the object motion is small. That is, the change in time does not cause a drastic change in the target position, and the displacement between adjacent frames should be relatively small.

[0069] Thus, an optical flow equation can be constructed as shown in Formula 1.1:

[0070] I(x, y, t) = I(x + dx, y + dy, t + dt) Formula 1.1

[0071] Among them, I(x, y, t) is the gray value of the pixel at the position (x, y) in the image at time t. Expand I(x + dx, y + dy, t + dt) in formula 1.1 at (x, y, t) by the first-order Taylor series as shown in formula 1.2:

[0072]

[0073] Thus, we can obtain as shown in formula 1.3:

[0074]

[0075] Denote u and v are the optical flows of the pixel in the x-direction and y-direction to be solved. Among them, is the differential of the pixel gray space, is the temporal gray differential of the pixel coordinate point. Denote Then formula 1.3 can be abbreviated into matrix form as shown in formula 1.4:

[0076]

[0077] According to the aforementioned assumption of spatial consistency, for multiple points around this point, optical flow equations can be constructed. Abbreviated into matrix form as shown in formula 1.5:

[0078]

[0079] Since the above formula 1.4 contains two unknowns (u and v), therefore, through the linear equations of multiple points, the least squares method can be used to solve It can also be solved iteratively to obtain u and v.

[0080] Since the optical flow algorithm requires that the time between adjacent frames is short enough and the object motion is small. When the object motion is large, that is, when the object moves fast, the basic assumption does not hold, resulting in a large error in the finally solved optical flow value. Therefore, the image pyramid can be combined to reduce the original large motion to a small motion that satisfies the basic assumption of optical flow, and then optical flow estimation is performed.

[0081] An image pyramid is a collection of images. All the images in the collection are derived from the same original image and are obtained by continuously downsampling the original image. Image downsampling can also be described as downsampling, image subsampling, subsampling, or shrinking the image. By downsampling an image, the size of the image can be reduced. For example, assuming the size of a certain frame of image is M*N, by downsampling this image by a factor of S, the image a can be shrunk by S times, that is: the length and width of image a are both shrunk by S times. Then, the size of the image after downsampling is (M / S)*(N / S). When the image size is reduced, the object motion can be reduced, so the optical flow estimation algorithm can be used for small enough images. It should be noted that the image pyramid is used to process grayscale images. If the first image and the second image obtained are color images, the first image and the second image can be converted into grayscale images, and an image pyramid can be constructed based on the converted images.

[0082] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the structure of an image pyramid. As Figure 1 shown, there are two image pyramids. Among them, the left image pyramid is the image pyramid of the first image, and the right image pyramid is the image pyramid of the second image. The second image is the next frame image of the first image. At the bottom layer of the two image pyramids is the image of the original size, that is, I 0 is the image of the original size. Each image in the upper layer is obtained by downsampling the image in the lower layer by a factor of 2, that is, the width and height of image I 1 are both half of those of I 0 , and so on. Then, the image sequence of I 0 , I 1 , I 2 ,..., I L ,..., I n-1 can be obtained. In the image sequence, the relationship between the image in the L-th layer and the image in the upper layer (L-1 layer) can be shown as in Formula 1.6:

[0083]

[0084] where, for the feature point p(x, y) in I 0 , there is p L (x, y) corresponding to p in the image I L in the L-th layer, where x L =x / 2 L , y L =y / 2 L . In order to avoid a large amount of information loss in the image, the number of layers of the image pyramid generally does not exceed 4 layers. As Figure 1 shown by the 4-layer image pyramid (I 0, I 1 , I 2 , I 3 )。

[0085] In the pyramid optical flow algorithm, denote the initial optical flow of the L-th layer as the initial optical flow g of the top layer I n-1 as n -1 = [0 0] T . Then the optical flow calculation between adjacent layers can be shown as in Formula 1.7:

[0086] g L-1 = 2(g L + d L ) Formula 1.7

[0087] where d L is the optical flow increment calculated by the optical flow estimation algorithm for two adjacent frames of images (the first image and the second image) I t-1 and I t at the L-th layer. And so on, starting from the top layer and iterating down to the bottom layer, then the optical flow of the bottom layer can be as shown in Formula 1.8:

[0088]

[0089] Add the calculated optical flow of the bottom layer to the position of the feature point in the first image I t-1 to obtain the optical flow estimation position of the feature point in the second image I t .

[0090] To facilitate the understanding of the embodiments of the present application, first, one of the target tracking system architectures on which the embodiments of the present application are based will be described below. The target tracking system can be deployed in any computer device involved in target tracking. For example, it can be deployed on one or more computing devices (such as a central server) in a cloud environment, or on one or more computing devices (edge computing devices) in an edge environment, and the edge computing device can be a server. Among them, the cloud environment refers to a central computing device cluster owned by a cloud service provider for providing computing, storage, and communication resources, and the cloud environment has relatively more storage resources and computing resources. The edge environment refers to an edge computing device cluster that is geographically close to the original data acquisition device and is used to provide computing, storage, and communication resources. The target tracking system in the present application can also be deployed on one or more terminal devices. For example, the target tracking system can be deployed on a terminal device, and the terminal device has certain computing, storage, and communication resources, and can be a computer, a vehicle-mounted terminal, etc.

[0091] Please refer to Figure 2 , Figure 2It is a schematic diagram of a target tracking system architecture provided by an embodiment of the present application. As Figure 2 shown, the target tracking system includes a cloud and an intelligent vehicle 20. A target tracking device 201 and a collection device 202 are deployed on the intelligent vehicle 20. The target tracking device 201 can obtain images on the traffic road through the collection device 202, and store the collected images locally or in the cloud. Further, the target tracking device 201 tracks targets such as pedestrians and vehicles on the traffic road in two adjacent frames of images, determines the position area of the target in each frame of image, and obtains the target tracking result of each frame of image. And the target tracking device 201 can also store the target tracking result of each frame of image locally and in the cloud.

[0092] Among them, the collection device 202 is used to collect images of the traffic road from its own perspective. The collection device 202 can include, but is not limited to, cameras, infrared cameras, lidar, etc.

[0093] It can be understood that Figure 2 the target tracking system architecture in

[0094] The target tracking method provided by the embodiment of the present application can track the change of the position area of the target between two images. The target tracking method described in the present application can be applied to many fields such as driverless, motion detection, action capture and recognition, augmented reality (AR), etc. to achieve specific functions.

[0095] Exemplarily, the application scenario is a driverless scenario.

[0096] The driverless scenario means that the intelligent vehicle can rely on the in-vehicle intelligent driver mainly composed of a computer system to achieve driverless. During the driving process of the intelligent vehicle, the target tracking device deployed in the intelligent vehicle can obtain the images on the traffic road collected by the collection device in real time. For example, it can obtain the image of the vehicle driving in front of the current intelligent vehicle. Then, two adjacent frames of images are selected from the obtained continuous images, and the targets in the selected images are tracked to obtain the change result of the position area of the target between the two images. However, currently, after obtaining two consecutive frames of images, the target tracking will select feature points in a uniformly distributed manner in the previous frame of image and perform target tracking. However, the feature points selected in a uniform distribution will, to a certain extent, select feature points in the background area, resulting in low accuracy of target tracking.

[0097] At this time, the target in the previous frame image can be detected to obtain the attribute information of the target, and then the parameters of the extraction model can be determined according to the attribute information of the target, so as to determine the extraction model; and the extraction model is used to extract feature points in the first image, so that the feature points are more distributed in the area of the target, so as to improve the tracking accuracy during target tracking. Furthermore, the intelligent vehicle can generate corresponding driving instructions according to the result of target tracking.

[0098] Exemplarily, the target tracking device in the intelligent vehicle can detect the target in front from the first image. For example, pedestrians, cars, bicycles, etc. in front of the intelligent vehicle. According to the attribute information in the detection result, the central position, density, number of feature points, etc. for extracting feature points can be determined. Thus, the parameters of the extraction model can be determined according to the central position, density, and number of feature points, and an extraction model with specific parameters is obtained. Then, the extraction model with specific parameters is used to extract feature points for target tracking to obtain the position area where the target is located in the second image. Further, the position area where the target is located in the second image is analyzed and processed to generate unmanned driving instructions, such as accelerating, decelerating, braking, turning, etc., to ensure the smooth and safe driving of the intelligent vehicle.

[0099] It should be noted that the above-described application scenarios are for illustration purposes and do not constitute a limitation on the embodiments of the present application.

[0100] Based on Figure 2 the architecture of the provided target tracking system, combined with the target tracking method provided in the present application, the technical problems proposed in the present application are specifically analyzed and solved. Please refer to the attached Figure 3 , Figure 3 FIG. is a schematic flowchart of a target tracking method provided by an embodiment of the present application. The method may include the following steps S301-S305.

[0101] Step S301: The target tracking device detects the target in the first image to obtain the detection area of the target and the attribute information of the target.

[0102] The target, which can also be described as an object, is a visual target in an image. Exemplarily, the target can be a vehicle, a person, an animal, etc. in the image. In the embodiments of the present application, the target can be a vehicle in the first image or a pedestrian in the first image, etc.

[0103] Specifically, the target tracking device can obtain a first image, which can be a captured image or an image obtained from the cloud, and detect the target in the first image. Among them, detecting the target in the first image is to use a preset target detection model for detection to obtain a detection result. The preset target detection model can be a region-based convolutional neural network model (RCNN), or a single shot multiBox detector model (SSD), or a youonly look once (YOLO) model. There is no limitation here. After detection by the preset detection model, the obtained detection result includes the detection area of the target and the attribute information of the target. The detection area of the target can be a rectangular box including the target in the first image, and the attribute information of the target can include the occlusion information of the target, the heading angle of the target, and the category information of the target, etc.

[0104] Step S302: The target tracking device determines an extraction model according to the above-mentioned attribute information of the target.

[0105] In a possible implementation manner, the attribute information of the target may include the occlusion information of the target, and the occlusion information of the target includes the occlusion ratio of the target and the occlusion area information of the target. Please refer to Figure 4a , Figure 4a which is a schematic diagram of an application scenario of target tracking provided by an embodiment of the present application. As Figure 4a shown, the target tracking device in the intelligent vehicle 401 can obtain the image on the traffic road collected by the collection device and detect the obtained image. For example, the intelligent vehicle 401 can obtain the image of the vehicle 402 traveling in front of the intelligent vehicle 401 and detect the vehicle 402 in the obtained image. Please refer to Figure 4b , Figure 4b which is a schematic diagram of a first image provided by an embodiment of the present application. As Figure 4b shown, the first image may include the vehicle 402 traveling in front of the intelligent vehicle 401. Optionally, the first image may further include the occluded area of the vehicle 402 and the detection area of the vehicle 402. For example, Figure 4b the shaded part in [reference numeral] is the occluded part of the vehicle 402, 20% indicates the detected occlusion ratio of the vehicle 402, and the black rectangular box surrounding the vehicle 402 is the detection area of the vehicle 402.

[0106] Among them, a preset extraction model can be stored in the target tracking device. The extraction model can be a Gaussian model, a Weber model, a triangular model, a Rice model, a Laplace model, etc. The extraction model contains preset initial parameters, and the initial parameters can determine the initial center position of the extracted feature points, the density of the feature points in each area of the target detection area, and the number of feature points.

[0107] Exemplarily, for the convenience of description, the present application takes a one-dimensional Gaussian model as an example for explanation. It should be noted that in actual use, a two-dimensional Gaussian model is used to extract feature points, as shown in the model of the two-dimensional Gaussian distribution in Figure 5 As shown. Among them, the function of the one-dimensional Gaussian model can be as shown in Formula 4.1:

[0108]

[0109] Among them, Φ(·) in Formula 4.1 is the cumulative distribution function of the standard normal distribution; a and b are range parameters; μ is the expectation; σ 2 is the variance. It should be noted that the range parameters a and b are equivalent to the truncation values of the Gaussian model and are used to determine the extracted feature points in the target detection area.

[0110] Next, two parameters of the extraction model will be introduced by taking the one-dimensional Gaussian distribution as an example.

[0111] Parameter 1: Expectation μ

[0112] The expectation μ in the Gaussian model is used to determine the center position of the model. For example, in the one-dimensional Gaussian distribution model, the center of the normal distribution curve is determined by the expectation μ, and the normal distribution curve is symmetric about the expectation μ. Similarly, the expectation in the two-dimensional Gaussian model is also used to determine the position where the center of the model is located, that is, the center position of the extracted feature points. In the present application, the center position of the extracted feature points can be determined by the occlusion information and the heading angle of the target in the attribute information of the target. Furthermore, the position of the center of the model is determined as the center position of the feature points, and the expectation of the model is obtained, so as to obtain an extraction model containing the determined expectation.

[0113] In a possible implementation, the extraction model includes an initial expectation, that is, the initial center position of the extracted feature points. Exemplarily, the initial center position of the extracted feature points can be the center of the detection area of the target. The attribute information of the target includes the occlusion information of the target, and the occlusion information of the target can include the occlusion ratio of the target and the information of the occluded area of the target. To improve the accuracy of target tracking, the center position of the extracted feature points can be adjusted from the initial center position towards the direction of the unoccluded area to obtain a new center position of the extracted feature points, thereby determining the expectation of the extraction model and obtaining an extraction model containing the determined expectation. Using the extraction model containing the determined expectation can make more feature points distributed in the unoccluded area and improve the accuracy of target tracking. For the determination of the extraction model according to the occlusion information, refer to the specific description of Embodiment 1.

[0114] In a possible implementation, the attribute information of the target includes the heading angle of the target. The heading angle of the target can be used to represent the orientation of the target. It can be understood that the target has different feature information in different orientations, and the feature information can be the information with obvious features of the target in the image. For example, when the target is a vehicle, the feature information can be the front of the vehicle, the body, and the rear of the vehicle. If the front of the vehicle is shown in the first image, the area of the front can be used as the feature area; when a part of the rear and a part of the body of the vehicle are shown in the first image, the shown part of the body or the shown part of the rear can be used as the feature area. Taking the area where the features of the target are more obvious as the feature area, the center position of the extraction model can be adjusted from the initial center position towards the center of the feature area of the target to obtain a determined expectation, thereby obtaining an extraction model containing the determined expectation. For the determination of the extraction model according to the heading angle, refer to the specific description of Embodiment 2.

[0115] Parameter Two: Variance σ 2

[0116] Variance σ in the Gaussian model 2 Used to determine the shape of the model. For example, in a one-dimensional Gaussian distribution model, the degree of flatness / sharpness of the waveform of the normal distribution curve is determined by the value of the variance σ 2 The larger the value of the variance σ 2 The flatter the normal distribution curve, and the smaller the value of the variance σ 2 The sharper the normal distribution curve. Similarly, in a two-dimensional Gaussian model, the variance is also used to determine, such as Figure 5The sharpness or stoutness of the model. In the embodiments of the present application, by using the corresponding first adjustment parameter and second adjustment parameter in the attribute information of the target, the density of the extracted feature points distributed in each region and the number of feature points can be determined. Furthermore, the variance of the extraction model can be determined by using the first adjustment parameter and / or the second adjustment parameter, so that the extraction model including the determined variance can be determined. For determining the extraction model according to the density of the feature points distributed in different regions, refer to the specific description of Embodiment 3. For determining the extraction model according to the number of feature points distributed in different regions, refer to the specific description of Embodiment 4.

[0117] Step S303: The target tracking device determines a first set of feature points according to the above extraction model. The first set of feature points is distributed in the detection region of the target. The detection region includes a first region and a second region other than the first region. The number of feature points of the target in the first region is greater than the number of feature points of the target in the second region.

[0118] Specifically, the target tracking device determines the first set of feature points according to the extraction model. Taking the Gaussian model as an example, it may be to randomly obtain the coordinates of a pixel point in the first image, and determine whether the coordinates of the pixel point satisfy the function corresponding to the Gaussian model. If so, the point is determined as a feature point, and then randomly obtain the coordinates of a pixel point again and judge again; if the coordinates of the pixel point do not satisfy the function corresponding to the Gaussian model, randomly obtain the coordinates of a pixel point again for judgment until the number of feature points that satisfy the function corresponding to the Gaussian model reaches the set total number. Among them, the feature points extracted by the above extraction model including a specific expectation and a specific variance are distributed in the first region and the second region in the detection region of the target, and the number of feature points of the target in the first region is greater than the number of feature points of the target in the second region.

[0119] Step S304: The target tracking device determines a second set of feature points of the target in the second image according to the first set of feature points. The second image is the next frame image of the first image.

[0120] In a possible implementation manner, the target tracking device estimates the optical flow of each feature point in the first set of feature points. Optical flow estimation includes forward estimation and backward estimation. Forward estimation is to estimate the target in the next frame image according to the previous frame image by using the pyramid optical flow algorithm. For example, taking the first image as the previous frame image of the second image, the forward estimation is to estimate the position area of the target in the first image in the second image. Backward estimation is to estimate the target in the previous frame image according to the next frame image by using the pyramid optical flow algorithm. For example, estimate the position area of the target in the second image in the first image. Reference can be made together Figure 6a , Figure 6aSchematic diagram of forward estimation and reverse estimation provided by an embodiment of the present application. Figure 6a As shown, the bold line on the left represents the first image, and the bold line on the right represents the second image, and the first image is the previous frame of the second image. Each feature point in the first feature point set extracted from the first image is forward estimated using the pyramid optical flow algorithm to obtain a third feature point set in which each feature point in the first feature point set in the first image corresponds to a position in the second image. Then, each feature point in the third feature point set extracted from the second image is reversely estimated to obtain a fourth feature point set in which each feature point in the third feature point set in the second image corresponds to a position in the first image.

[0121] Furthermore, in order to reduce the error generated by the pyramid optical flow algorithm in the optical flow estimation process, feature points with large errors can be removed, and the target tracking device can respectively calculate the distance between feature points in the fourth feature point set corresponding to each feature point in the first feature point set. The distance calculation can be performed by calculating the distance between two feature points by using calculation methods such as Euclidean distance, Manhattan distance, and cosine distance.

[0122] Taking the Euclidean distance as an example, the feature point A(x i,1 ,y i,1 ), corresponding to the feature point B(x i,2 ,y i,2 ), then the distance d between feature point A and feature point B i The calculation of can be shown as formula 4.2:

[0123]

[0124] In the third feature point set, that is, the feature point set in the second image, remove d i Specifically, the target tracking device can record the indexes of the feature points greater than the preset threshold, put them into the index set, and remove the feature points with the indexes recorded in the third feature point set to obtain the second feature point set.

[0125] It can be understood that, before and after the elimination, each feature point in the second feature point set corresponds one-to-one with each feature point in the first feature point set and the fourth feature point set.

[0126] Step S305: The target tracking device determines the location area of ​​the target in the second image according to the second feature point set.

[0127] In a possible implementation, to reduce the error of optical flow estimation, the target tracking device may obtain the mapping relationship between distance and weight. It can be understood that the larger the calculated distance is, the larger the error is. Therefore, a decreasing function can be constructed to calculate the weight, which can meet the requirement that the smaller the distance is, the larger the weight is. In the embodiments of the present application, the mapping relationship between distance and weight may satisfy the mapping relationship of the decreasing function shown in Formula 4.3:

[0128]

[0129] where w i is the weight of the i-th feature point, and d i is the distance of the i-th feature point.

[0130] Further, reference may be made to Figure 6b , Figure 6b which is a schematic diagram for determining a position area provided in the embodiments of the present application. As shown in Figure 6b , the seven circles represent seven feature points. Each feature point includes the coordinates of the feature point and the corresponding weight. The central position of the target in the second image is calculated by weighted average, that is, the position of the five-pointed star in Figure 6b . Further, the target tracking device determines the position area of the target in the second image according to the size and shape of the detection area of the target in the first image. It may be a region centered on the central position of the target in the second image and having the size and shape of the detection area determined in the first image as the position area of the target in the second image.

[0131] The following provides a detailed description of Embodiment 1 - Embodiment 4.

[0132] Embodiment 1 - Determining an extraction model according to occlusion information

[0133] As shown in Figure 4b , the occlusion information may include an occlusion ratio (20%) and occlusion area information. The occlusion area information may be that the left side is occluded or the right side is occluded. It should be noted that in the scenario provided in the embodiments of the present application, the occlusion area information may include information in two directions: the left direction and the right direction. As shown in Figure 4b , a part of the left side of the vehicle is occluded. To make more feature points distributed in the unoccluded area of the target, the target tracking device may determine the central position of the extracted feature points as the center of the unoccluded area, or adjust the central position of the initial feature points in the direction towards the center of the unoccluded area. Then, according to the adjusted central position of the extracted feature points, the expectation of the model is determined, and an extraction model including the determined expectation is obtained.

[0134] Specifically, the target tracking device can determine the adjustment distance corresponding to the occlusion ratio of the target according to the corresponding relationship between the occlusion ratio and the adjustment distance; and determine the adjustment direction of the center position of the feature points to be extracted according to the opposite direction corresponding to the occlusion area information of the target, and then determine the center position of the feature points to be extracted by the extraction model according to the adjustment distance and the adjustment direction. Among them, the adjustment distance can be a value set artificially corresponding to the occlusion ratio, or a value of the adjustment distance corresponding to adjusting the center position of the feature points to be extracted to the center of the unoccluded part of the target within the detection area.

[0135] Further, there is a one-to-one correspondence between the occlusion ratio of the target and the expected adjustment value, and there is a corresponding relationship between the adjustment direction of the center position of the feature points to be extracted and the expected adjustment sign. Then, the target tracking device can determine the expectation according to the initial expectation, the expected adjustment value, and the expected adjustment sign. For example, when the direction corresponding to the occlusion area information of the target is the left direction, the sign corresponding to the expected adjustment value can be a negative sign or a positive sign. When the sign is a negative sign, expectation = initial expectation - expected adjustment value; conversely, when the sign is a positive sign, expectation = initial expectation + expected adjustment value. Then an extraction model including the determined expectation is obtained.

[0136] Please refer to Figure 7a , Figure 7a which is a schematic diagram of adjusting the center position of the feature points according to the expectation provided by the embodiment of the present application. As Figure 7a shown, the target is a vehicle, and the initial center position is the position where the gray circle is located, that is, the center of the target detection area; according to the occlusion ratio of 20%, the adjustment distance, that is, the expected adjustment value, can be d as in Figure 7a . The adjustment direction is the opposite direction of the direction corresponding to the occlusion area information, that is, the right direction shown by the arrow in Figure 7b , and the expected adjustment sign is obtained, so as to obtain the expected value, and the center position of the feature points to be extracted is also obtained, that is, the position of the black circle. Thus, an extraction model including the determined expectation can be obtained. By adopting the method of Embodiment 1 of the present application, the center position of the feature points extracted by the extraction model can be adjusted to the unoccluded area, so that the feature points are distributed in the unoccluded area, making the distribution of the feature points more accurate and improving the accuracy of target tracking.

[0137] Embodiment 2 - Determining the Extraction Model According to the Heading Angle

[0138] The heading angle is the angle between the target movement direction and the horizontal direction of the image, and the range of the heading angle is (0°, 360°). When the target is at different heading angles, different feature regions can be displayed in the image. Among them, the feature region can be a region containing obvious features of the target, that is, a region that matches the feature information except for the background region in the first image.

[0139] When the target is a vehicle, the characteristic area of the vehicle can be the area where the vehicle head is located, the area where the side of the vehicle is located, or the area where the vehicle tail is located. Exemplarily, when the target is a vehicle, when the heading angle of the vehicle is 0° or 180°, the side of the vehicle can be the characteristic area; when the heading angle of the vehicle is within the range of (0°, 180°), the vehicle tail can be the characteristic area; when the heading angle of the vehicle is within the range of (180°, 360°), the vehicle head can be the characteristic area. Optionally, the characteristic information of the target can be determined according to the heading angle, and then the target can be recognized and detected to obtain the characteristic area. Please refer to Figure 7b , Figure 7b which is a schematic diagram for adjusting the center position of the feature points according to the vehicle heading angle provided by an embodiment of the present application. As Figure 7b shown, in the attribute information of the target, the heading angle can be 120°, then the vehicle tail is a relatively obvious area of the vehicle, that is, the characteristic area. Furthermore, the center position of the extracted feature points can be adjusted by a preset distance in the direction from the center of the target detection area, that is, the position of the gray circle, to the center of the characteristic area, which is Figure 7b the position of the black circle in

[0140] Among them, the expected adjustment value can be determined according to the preset adjustment distance, and the expected adjustment symbol can be determined according to the adjustment direction of the center of the feature point extraction. When the symbol is a negative sign, expected = initial expected - expected adjustment value; conversely, when the symbol is a positive sign, expected = initial expected + expected adjustment value. Furthermore, an extraction model including the determined expected value is obtained. By using the method of Embodiment 2 of the present application, the center position of the feature points extracted by the extraction model can be adjusted to an area where the target features are more obvious, making the distribution of the feature points more reasonable and improving the accuracy of target tracking.

[0141] Embodiment 3 - Determining the extraction model according to the density of the distribution of feature points in different regions

[0142] After the target tracking device detects the first image, the category information of the target is obtained. For example, the category information can include the category of the target, such as vehicle, pedestrian, bicycle, etc. Among them, there is a corresponding relationship between the category information of the target and the first adjustment parameter, and the first adjustment parameter can be used to determine the density of the distribution of the extracted feature points in each area of the target detection area. And the initial variance can be adjusted according to the value of the first adjustment parameter to obtain the variance of the extraction model. It should be noted that adjusting the variance of the Gaussian model with the first adjustment parameter can be to modify the original variance to the first adjustment parameter, or to adjust the variance of the Gaussian model according to the first adjustment parameter to obtain the variance of the extraction model.

[0143] Among them, the detection area of the target may include a first area and a second area. The first area may be an area centered on the center of the detection area of the target and with a length and width of 10% or 20% of the rectangular frame of the detection area. The second area is the area other than the first area in the detection area of the target.

[0144] Exemplarily, please refer to Figure 8a and Figure 8b , Figure 8a and Figure 8b are all schematic diagrams of an extraction model provided by an embodiment of the present application determined according to the density of the distribution of feature points of the target in different regions. When the category information of the target is a pedestrian, the target tracking device determines the first adjustment parameter corresponding to the pedestrian, and then the variance in the Gaussian model can be determined through the first adjustment parameter, and then an extraction model including the determined variance is obtained, such as Figure 8a the Gaussian model shown on the left; using Figure 8a the Gaussian model shown on the left, the density of the feature points extracted in the first area and the second area in the detection area of the pedestrian can be as Figure 8a shown on the right.

[0145] When the category information of the target is a vehicle, the target tracking device determines the variance in the Gaussian model through the first adjustment parameter corresponding to the vehicle, and then an extraction model including the determined variance is obtained, such as Figure 8b the Gaussian model shown on the left; using Figure 8b the Gaussian model shown on the left, the density of the feature points in the first area and the second area in the detection area of the vehicle is obtained, and the feature points of the gray circles shown in Figure 8b on the right are obtained. Among them, the dashed rectangular frame is the first area, and the second area is the area other than the first area in the detection area of the target, such as Figure 8b the area around the first area shown. It can be seen that the density of the feature points in the first area is greater than the density of the feature points in the second area. By adopting the method of Embodiment 3 of the present application, the feature points extracted by the extraction model can be distributed more densely in the central area and more dispersed and sparse in the surrounding area, using the area distributed in the center of the target, reducing the area distributed outside the center of the target and the background area, and improving the accuracy of target tracking.

[0146] Embodiment 4 - Determining the extraction model according to the number of feature points distributed in different regions

[0147] Taking the one-dimensional Gaussian model as an example, the range parameters a and b in Formula 4.1 are the truncation parameters of the Gaussian model, which are used to limit all feature points within the detection area of the target. Similarly, two additional range parameters can be added to limit the number of feature points in the first area and the second area in the target detection area. Among them, the first area can be an area centered on the center of the detection area of the target and with a length and width of 10% or 20% of the rectangular frame of the detection area, and the second area is the area in the target detection area except the first area. By means of the two additional range parameters, the number of feature points within the first area is restricted. Thus, based on the total number of feature points and the number of feature points within the first area, the number of feature points within the second area can be restricted. When the range parameters are determined, by adjusting the variance of the Gaussian model, the number of feature points distributed in the first area and the second area in the detection area of the target can be adjusted.

[0148] Exemplarily, please refer to Figure 8c and Figure 8d , Figure 8c and Figure 8d which are all schematic diagrams of an extraction model determined according to the number of feature points in different areas provided by the embodiments of the present application. When the category information of the target is a pedestrian, the second adjustment parameter corresponding to the pedestrian is determined, and then the variance in the Gaussian model can be determined through the second adjustment parameter, and further an extraction model including the determined variance can be obtained, such as Figure 8c the Gaussian model shown on the left. The number of feature points distributed in the first area and the second area in the detection area of the pedestrian can be as Figure 8c shown on the right, where Figure 8c the dashed box in is the first area, and the gray circles are the feature points. When the category information of the target is a vehicle, the target tracking device determines the first adjustment parameter corresponding to the vehicle to determine the variance in the Gaussian model, and further obtains an extraction model including the determined variance, such as Figure 8d the Gaussian model shown on the left; using Figure 8d the Gaussian model shown on the left to extract the feature points in the first area and the second area in the detection area of the vehicle, and the feature points shown as gray circles as Figure 8d shown on the right are obtained. Among them, the dashed rectangular box is the area of the first area, and the second area is the area around the first area.

[0149] It should be noted that adjusting the variance of the Gaussian model with the second adjustment parameter can either modify the original variance to the second adjustment parameter or adjust the variance of the Gaussian model according to the second adjustment parameter. By adopting the method of Embodiment 4 of the present application, more feature points extracted by the extraction model can be distributed in the central area of the target, and the number of feature points distributed in the surrounding area and the background area of the target can be reduced, thereby improving the utilization rate of the feature points and also improving the accuracy of target tracking.

[0150] The method of the embodiments of the present application is elaborated in detail above. The related devices of the embodiments of the present application are provided below.

[0151] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a target tracking device provided by an embodiment of the present application. The target tracking device 90 may include a detection unit 901, a first determination unit 902, a second determination unit 903, a third determination unit 904, and a fourth determination unit 905. The detailed descriptions of each unit are as follows:

[0152] The detection unit 901 is configured to detect a target in a first image to obtain a detection area of the target and attribute information of the target; the first determination unit 902 is configured to determine an extraction model according to the attribute information of the target; the second determination unit 903 is configured to determine a first feature point set according to the extraction model; the first feature point set is distributed in the detection area of the target, the detection area includes a first area and a second area other than the first area, and the number of feature points of the target in the first area is greater than the number of feature points of the target in the second area; the third determination unit 904 is configured to determine a second feature point set of the target in a second image according to the first feature point set, where the second image is the next frame image of the first image; the fourth determination unit 905 is configured to determine a position area of the target in the second image according to the second feature point set.

[0153] In a possible implementation manner, the attribute information of the target includes one or more of occlusion information of the target, the heading angle of the target, and category information of the target.

[0154] In a possible implementation manner, the first determination unit 902 is specifically configured to: determine the central position of the extraction feature points according to the attribute information of the target; and determine the extraction model according to the central position of the extraction feature points.

[0155] In a possible implementation manner, the occlusion information of the target includes the occlusion ratio of the target and occlusion area information of the target; the first determination unit 902 is specifically configured to: obtain an initial central position of the extraction feature points; determine an adjustment distance corresponding to the occlusion ratio of the target according to the corresponding relationship between the occlusion ratio and the adjustment distance; and determine the central position of the extraction feature points according to the initial central position of the extraction feature points, the adjustment distance corresponding to the occlusion ratio of the target, and the opposite direction of the direction corresponding to the occlusion area information.

[0156] In a possible implementation, the above-mentioned first determination unit 902 is specifically configured to: obtain the initial central position of the extracted feature points; determine the feature region of the target according to the heading angle of the target, where the feature region is a region that matches the preset feature information except for the background region; determine the central position of the extracted feature points according to the initial central position of the extracted feature points, the preset adjustment distance, and the direction of the center of the feature region.

[0157] In a possible implementation, the above-mentioned first determination unit 902 is specifically configured to: determine the density of the extracted feature points distributed in each region of the detection region of the target according to the attribute information of the target; determine the extraction model according to the density of the extracted feature points distributed in each region of the detection region of the target.

[0158] In a possible implementation, the above-mentioned first determination unit 902 is specifically configured to: obtain the initial density of the extracted feature points distributed in each region of the detection region of the target; determine the first adjustment parameter corresponding to the category information of the target according to the correspondence between the category information and the adjustment parameter; determine the density of the extracted feature points distributed in each region of the detection region of the target according to the initial density of the extracted feature points distributed in each region of the detection region of the target and the first adjustment parameter corresponding to the category information of the target.

[0159] In a possible implementation, the above-mentioned first determination unit 902 is specifically configured to: determine the number of the extracted feature points distributed in each region of the detection region of the target according to the attribute information of the target; determine the extraction model according to the number of the extracted feature points distributed in each region of the detection region of the target.

[0160] In a possible implementation, the above-mentioned first determination unit 902 is specifically configured to: obtain the initial number of the extracted feature points distributed in each region of the detection region of the target; determine the second adjustment parameter corresponding to the category information of the target according to the correspondence between the category information and the adjustment parameter; determine the number of the extracted feature points distributed in each region of the detection region of the target according to the initial number of the extracted feature points distributed in each region of the detection region of the target and the second adjustment parameter corresponding to the category information of the target.

[0161] In a possible implementation manner, the above-mentioned third determination unit 904 is specifically configured to: perform forward estimation on all feature points in the above-mentioned first feature point set according to the pyramid optical flow algorithm, and determine the corresponding feature points of each feature point in the above-mentioned first feature point set in the above-mentioned second image to obtain a third feature point set; perform backward estimation on all feature points in the above-mentioned third feature point set according to the pyramid optical flow algorithm, and determine the corresponding feature points of each feature point in the above-mentioned third feature point set in the above-mentioned first image to obtain a fourth feature point set; determine the above-mentioned second feature point set according to the above-mentioned first feature point set, the above-mentioned third feature point set, and the above-mentioned fourth feature point set.

[0162] In a possible implementation manner, the above-mentioned third determination unit 904 is specifically configured to: determine the distance between each feature point in the above-mentioned first feature point set and the corresponding feature point in the above-mentioned fourth feature point set; remove the feature points in the above-mentioned third feature point set corresponding to the distances greater than a preset threshold to obtain the above-mentioned second feature point set.

[0163] In a possible implementation manner, the above-mentioned fourth determination unit 905 is specifically configured to: determine the weights of the distance mappings corresponding to each feature point in the above-mentioned second feature point set according to the mapping relationship between distance and weight; determine the position area of the target in the above-mentioned second image according to the weights of the distance mappings corresponding to each above-mentioned feature point in the above-mentioned second feature point set and the detection area of the target in the above-mentioned first image.

[0164] In a possible implementation manner, the above-mentioned fourth determination unit 905 is specifically configured to: determine the center position of the target in the above-mentioned second image according to the positions and corresponding weights of each above-mentioned feature point in the above-mentioned second feature point set; determine the size and shape of the above-mentioned detection area; determine the position area of the target in the above-mentioned second image according to the center position of the target in the above-mentioned second image, the size and shape of the above-mentioned detection area.

[0165] In a possible implementation manner, the above-mentioned detection unit 901 is specifically configured to: input the above-mentioned first image into a preset target detection model to obtain a detection result, where the detection result includes the detection area of the target and the attribute information of the target.

[0166] It should be noted that for the functions of each functional unit in the target tracking device 900 described in the embodiments of the present application, reference may be made to the relevant descriptions of steps S301 - S305 in the above-mentioned Figure 3 method embodiments, which will not be elaborated here.

[0167] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 10As shown, the computer device 1000 includes: a processor 1001, a communication interface 1002, and a memory 1003. The above-mentioned processor 1001, communication interface 1002, and memory 1003 are interconnected with each other through an internal bus 1004.

[0168] The above-mentioned processor 1001 can be composed of one or more general-purpose processors. For example, a central processing unit (CPU), or a combination of a CPU and a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above-mentioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0169] The bus 1004 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The above-mentioned bus 1004 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 10 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0170] The memory 1003 can include a volatile memory, such as a random access memory (RAM); the memory 1003 can also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); the memory 1003 can also include a combination of the above types. The program code can be used to implement Figure 3 the method steps of the method embodiment of the target tracking method shown with the computing device as the execution subject.

[0171] It should be noted that for the functions of each functional unit in the computer device 1000 described in the embodiments of the present application, reference may be made to the relevant descriptions of steps S301 - S305 in the above-mentioned Figure 3 method embodiment in the above, and details are not described herein again.

[0172] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it can implement some or all of the steps described in any one of the above method embodiments, and implement the Figure 9 functions of any one of the functional modules described above.

[0173] The embodiments of the present application also provide a computer program product. When it runs on a computer or a processor, it causes the computer or the processor to execute one or more steps in any of the above methods. If each component module of the above-mentioned device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above-mentioned computer-readable storage medium.

[0174] In the above embodiments, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0175] It should be understood that the first, second, third, fourth, and various numerical numbers involved herein are only for the convenience of description and are not used to limit the scope of the present application.

[0176] It should be understood that the term "and / or" herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0177] It should also be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0178] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0179] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0180] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0181] The units described as separate components above may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0182] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0183] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0184] The steps in the method embodiments of the present application can be adjusted, combined, and deleted according to actual needs.

[0185] The modules in the device embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0186] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A target tracking method, characterized in that, comprising: detecting a target in a first image to obtain a detection region of the target and attribute information of the target; the attribute information of the target includes the heading angle of the target; acquiring an initial center position of the extracted feature points, determining a feature region of the target according to the heading angle of the target, the feature region being a region that matches preset feature information except for the background region, determining the center position of the extracted feature points according to the initial center position of the extracted feature points, a preset adjustment distance, and the direction of the center of the feature region, and determining an extraction model according to the center position of the extracted feature points; or, determining distribution information of the extracted feature points in each region of the detection region of the target according to the attribute information of the target, and determining an extraction model according to the distribution information; determining a first set of feature points according to the extraction model; the first set of feature points is distributed within the detection region of the target, the detection region includes a first region and a second region other than the first region, and the number of feature points of the target in the first region is greater than the number of feature points of the target in the second region; determining a second set of feature points of the target in a second image according to the first set of feature points, the second image being the next frame image of the first image; determining a position region of the target in the second image according to the second set of feature points.

2. The method according to claim 1, characterized in that, the attribute information of the target further includes one or more of the occlusion information of the target and the category information of the target.

3. The method according to claim 2, characterized in that, the occlusion information of the target includes the occlusion ratio of the target and the occlusion region information of the target; the determining the center position of the extracted feature points according to the attribute information of the target includes: acquiring an initial center position of the extracted feature points; determining an adjustment distance corresponding to the occlusion ratio of the target according to the corresponding relationship between the occlusion ratio and the adjustment distance; determining the center position of the extracted feature points according to the initial center position of the extracted feature points, the adjustment distance corresponding to the occlusion ratio of the target, and the opposite direction of the direction corresponding to the occlusion region information.

4. The method according to claim 1, characterized in that, the determining the distribution information of the extracted feature points in each region of the detection region of the target according to the attribute information of the target, and determining an extraction model according to the distribution information includes: determining the density of the extracted feature points in each region of the detection region of the target according to the attribute information of the target; determining an extraction model according to the density of the extracted feature points in each region of the detection region of the target.

5. The method according to claim 4, characterized in that, the determining the density of the extracted feature points in each region of the detection region of the target according to the attribute information of the target includes: acquiring the initial density of the extracted feature points in each region of the detection region of the target; Determine a first adjustment parameter corresponding to the category information of the target according to the corresponding relationship between the category information and the adjustment parameter; Determine the density of the extracted feature points distributed in each region of the detection region of the target according to the initial density of the extracted feature points distributed in each region of the detection region of the target and the first adjustment parameter corresponding to the category information of the target.

6. The method according to claim 1, wherein, the determining the distribution information of the extracted feature points distributed in each region of the detection region of the target according to the attribute information of the target, and determining the extraction model according to the distribution information includes: Determine the number of the extracted feature points distributed in each region of the detection region of the target according to the attribute information of the target; Determine the extraction model according to the number of the extracted feature points distributed in each region of the detection region of the target.

7. The method according to claim 6, wherein, the determining the number of the extracted feature points distributed in each region of the detection region of the target according to the attribute information of the target includes: Obtain the initial number of the extracted feature points distributed in each region of the detection region of the target; Determine a second adjustment parameter corresponding to the category information of the target according to the corresponding relationship between the category information and the adjustment parameter; Determine the number of the extracted feature points distributed in each region of the detection region of the target according to the initial number of the extracted feature points distributed in each region of the detection region of the target and the second adjustment parameter corresponding to the category information of the target.

8. The method according to any one of claims 1-7, wherein, the determining the second feature point set of the target in the second image according to the feature point set includes: Perform forward estimation on all feature points in the first feature point set according to the pyramid optical flow algorithm, and determine the feature points corresponding to each feature point in the first feature point set in the second image to obtain a third feature point set; Perform backward estimation on all feature points in the third feature point set according to the pyramid optical flow algorithm, and determine the feature points corresponding to each feature point in the third feature point set in the first image to obtain a fourth feature point set; Determine the second feature point set according to the first feature point set, the third feature point set and the fourth feature point set.

9. The method according to claim 8, wherein, the obtaining the second feature point set according to the first feature point set, the third feature point set and the fourth feature point set includes: Determine the distance between each feature point in the first feature point set and the corresponding feature point in the fourth feature point set; Remove the feature points in the third feature point set corresponding to the distances greater than the preset threshold to obtain the second feature point set.

10. The method according to claim 9, wherein, the determining the position region of the target in the second image according to the second feature point set includes: Determine the weight of the distance mapping corresponding to each feature point in the second feature point set according to the mapping relationship between the distance and the weight; Determine the position area of the target in the second image according to the weights of the distance mappings corresponding to each of the feature points in the second feature point set and the detection area of the target in the first image.

11. The method according to claim 10, wherein, the determining the position area of the target in the second image according to the weights corresponding to each of the feature points in the second feature point set and the detection area of the target in the first image includes: determining the central position of the target in the second image according to the positions and corresponding weights of each of the feature points in the second feature point set; determining the size and shape of the detection area; determining the position area of the target in the second image according to the central position of the target in the second image, the size and shape of the detection area.

12. The method according to any one of claims 1-7 or 9-11, wherein, the detecting the target in the first image to obtain the detection area of the target and the attribute information of the target includes: inputting the first image into a preset target detection model to obtain a detection result, where the detection result includes the detection area of the target and the attribute information of the target.

13. A target tracking device, wherein, it includes: a detection unit configured to detect a target in a first image to obtain the detection area of the target and the attribute information of the target; the attribute information of the target includes the heading angle of the target; a first determination unit configured to obtain the initial central position of the extracted feature points, determine the feature area of the target according to the heading angle of the target, where the feature area is an area that matches preset feature information except for the background area, determine the central position of the extracted feature points according to the initial central position of the extracted feature points, a preset adjustment distance, and the direction of the center of the feature area, and determine an extraction model according to the central position of the extracted feature points; or, determine the distribution information of the extracted feature points in each area of the detection area of the target according to the attribute information of the target, and determine the extraction model according to the distribution information; a second determination unit configured to determine a first feature point set according to the extraction model; the first feature point set is distributed within the detection area of the target, the detection area includes a first area and a second area other than the first area, and the number of feature points of the target in the first area is greater than the number of feature points of the target in the second area; a third determination unit configured to determine a second feature point set of the target in a second image according to the first feature point set, where the second image is the next frame image of the first image; a fourth determination unit configured to determine the position area of the target in the second image according to the second feature point set.

14. The device according to claim 13, wherein, the attribute information of the target further includes one or more of the occlusion information of the target and the category information of the target.

15. The device according to claim 14, wherein, The occlusion information of the target includes the occlusion ratio of the target and the occlusion area information of the target; The first determining unit is specifically configured to: Obtain the initial center position of the extracted feature points; Determine the adjustment distance corresponding to the occlusion ratio of the target according to the corresponding relationship between the occlusion ratio and the adjustment distance; Determine the center position of the extracted feature points according to the initial center position of the extracted feature points, the adjustment distance corresponding to the occlusion ratio of the target, and the opposite direction of the direction corresponding to the occlusion area information.

16. The device according to claim 13, wherein, The first determining unit is specifically configured to: Determine the density of the extracted feature points distributed in each area of the detection area of the target according to the attribute information of the target; Determine the extraction model according to the density of the extracted feature points distributed in each area of the detection area of the target.

17. The device according to claim 16, wherein, The first determining unit is specifically configured to: Obtain the initial density of the extracted feature points distributed in each area of the detection area of the target; Determine the first adjustment parameter corresponding to the category information of the target according to the corresponding relationship between the category information and the adjustment parameter; Determine the density of the extracted feature points distributed in each area of the detection area of the target according to the initial density of the extracted feature points distributed in each area of the detection area of the target and the first adjustment parameter corresponding to the category information of the target.

18. The device according to claim 13, wherein, The first determining unit is specifically configured to: Determine the number of the extracted feature points distributed in each area of the detection area of the target according to the attribute information of the target; Determine the extraction model according to the number of the extracted feature points distributed in each area of the detection area of the target.

19. The device according to claim 18, wherein, The first determining unit is specifically configured to: Obtain the initial number of the extracted feature points distributed in each area of the detection area of the target; Determine the second adjustment parameter corresponding to the category information of the target according to the corresponding relationship between the category information and the adjustment parameter; Determine the number of the extracted feature points distributed in each area of the detection area of the target according to the initial number of the extracted feature points distributed in each area of the detection area of the target and the second adjustment parameter corresponding to the category information of the target.

20. The device according to any one of claims 13-19, wherein, The third determining unit is specifically configured to: Perform forward estimation on all the feature points in the first feature point set according to the pyramid optical flow algorithm, and determine the corresponding feature points of each feature point in the first feature point set in the second image to obtain a third feature point set; Perform backward estimation on all the feature points in the third feature point set according to the pyramid optical flow algorithm, and determine the corresponding feature points of each feature point in the third feature point set in the first image to obtain a fourth feature point set; Determine the second set of feature points according to the first set of feature points, the third set of feature points, and the fourth set of feature points.

21. The apparatus according to claim 20, wherein, the third determining unit is specifically configured to: determine the distance between each feature point in the first set of feature points and the corresponding feature point in the fourth set of feature points; remove the feature points in the third set of feature points corresponding to the distances greater than a preset threshold to obtain the second set of feature points.

22. The apparatus according to claim 21, wherein, the fourth determining unit is specifically configured to: determine the weight of the distance mapping corresponding to each feature point in the second set of feature points according to the mapping relationship between the distance and the weight; determine the position area of the target in the second image according to the weight of the distance mapping corresponding to each feature point in the second set of feature points and the detection area of the target in the first image.

23. The apparatus according to claim 22, wherein, the fourth determining unit is specifically configured to: determine the center position of the target in the second image according to the position and the corresponding weight of each feature point in the second set of feature points; determine the size and shape of the detection area; determine the position area of the target in the second image according to the center position of the target in the second image, the size and shape of the detection area.

24. The apparatus according to any one of claims 13-19 or 21-23, wherein, the detection unit is specifically configured to: input the first image into a preset target detection model to obtain a detection result, where the detection result includes the detection area of the target and the attribute information of the target.

25. A computer device, wherein, it includes a processor and a memory, the processor and the memory are connected to each other, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1-12.

26. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Feature-point-assistance-based target tracking method of space-time context

    CN108022254A

  • Target tracking method and device, electronic equipment and storage medium

    CN111047626A

  • Image frame position detection method and device, electronic equipment and storage medium

    CN111079741A