An object tracking method, device, electronic device and storage medium

By using historical images to predict target positions and determine image acquisition types, and combining combination and weight information to update trajectory information, the problem of inefficient target tracking insurround-view autonomous driving is solved, and more efficient target tracking and lower calculation amount are achieved.

CN113989761BActive Publication Date: 2025-06-24CHINA AUTOMOTIVE INNOVATION CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111275859.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-06-24
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

In surround-view autonomous driving, image acquisition of multiple surveillance cameras leads to inefficient target tracking and analysis, and the calculation is large, especially when the four cameras are facing differently, increasing the difficulty of target matching.

Method used

The position prediction information is determined through the historical image, the image acquisition type is determined based on the position prediction information, and at least one combination and its weight information are determined based on the image acquisition type, the object track information is updated through the weight information, the possibility of objects being lost across cameras is reduced, and the tracking efficiency is improved.

Benefits of technology

The target tracking of surround-view autonomous driving is achieved, reducing the possibility of objects being lost across cameras, and to a certain extent reduces time consumption and improves the efficiency of target tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113989761B_ABST
    Figure CN113989761B_ABST
Patent Text Reader

Abstract

The present invention discloses an object tracking method, device, electronic device and storage medium, including: obtaining a current image and a historical image; when there is a preset object in the historical image, based on the historical image, obtaining historical state information of at least one first preset object in the historical image; determining position prediction information corresponding to each of the at least one first preset object according to the historical state information; determining an image acquisition type according to the position prediction information; when the image acquisition type is a cross-camera acquisition type, determining at least one combination, and determining distance information and corresponding first feature information corresponding to the at least one combination; determining weight information corresponding to each of the at least one combination according to the distance information and the first feature information; updating object trajectory information according to the weight information. According to the technical solution of the present invention, the target tracking of surround-view automatic driving is realized, the possibility of object loss across cameras is reduced, and the tracking efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and particularly to an object tracking method, device, electronic device, and storage medium. Background Art

[0002] At the present stage, autonomous driving is a relatively popular field, and there is a need for research on multi-object tracking. Currently, most of the pedestrian tracking methods across cameras are based on research in the field of surveillance videos.

[0003] The difference from surround-view autonomous driving is that there is a lot of overlap in the scenarios of multiple surveillance cameras, and the relative positions of the same target can be captured in the same scenario. However, the positions of the surround-view fisheye cameras commonly used in current surround-view autonomous driving are located at the front of the vehicle, the rear of the vehicle, and under the two rearview mirrors respectively, and the shooting directions of the four cameras are different. After the four images are de-distorted, some spatially corresponding images will disappear, and there is little similarity in the background images that can be obtained by the four cameras, which also increases a certain degree of difficulty for target matching in tracking.

[0004] As is well known, autonomous driving is a field with high requirements for timeliness. After detecting and classifying the targets in the images obtained by the four cameras, it is necessary to track these targets. Analyzing the images has already consumed a large amount of time. How to minimize the computational effort as much as possible in the tracking part has become a difficult point. Especially when there are more cameras, it will bring a multiple increase in time consumption. For example, since the current surround-view uses four cameras and the orientations of the four cameras are completely different, it is necessary to analyze the four cameras simultaneously each time for tracking analysis, which will bring certain efficiency problems. Summary of the Invention

[0005] The purpose of the present invention is to provide an object tracking method, device, electronic device, and storage medium, which determine position prediction information through historical images, determine the image acquisition type based on the position prediction information, and determine at least one combination and the weight information corresponding to each combination based on the image acquisition type, and update the object trajectory information through the weight information, realizing target tracking in surround-view autonomous driving, reducing the possibility of object loss across cameras, and reducing the time consumption to a certain extent, improving the target tracking efficiency.

[0006] To achieve the above purpose, the present invention provides the following solutions:

[0007] An object tracking method, the method comprising:

[0008] Obtain a current image and a historical image;

[0009] When there is a preset object in the historical image, based on the historical image, obtain the historical state information of at least one first preset object in the historical image;

[0010] According to the historical state information, determine the position prediction information corresponding to each of the at least one first preset object;

[0011] According to the position prediction information, determine the image acquisition type;

[0012] When the image acquisition type is a cross-camera acquisition type, based on the current image and the historical image, determine at least one combination, and determine the distance information and the corresponding first feature information corresponding to the at least one combination;

[0013] According to the distance information and the first feature information, determine the weight information corresponding to each of the at least one combination;

[0014] According to the weight information, update the object trajectory information.

[0015] Optionally, after determining the image acquisition type according to the position prediction information, further include:

[0016] When the image acquisition type is a non-cross-camera acquisition type, based on the current image and the historical image, determine at least one combination, and determine the intersection over union information and the corresponding second feature information corresponding to the at least one combination;

[0017] According to the intersection over union information and the second feature information, determine the weight information corresponding to each of the at least one combination.

[0018] Optionally, determining the image acquisition type according to the position prediction information includes:

[0019] When the position point corresponding to the position prediction information does not belong to the preset range, determine that the image acquisition type is the cross-camera acquisition type;

[0020] When the position point corresponding to the position prediction information belongs to the preset range, determine that the image acquisition type is a non-cross-camera acquisition type.

[0021] Optionally, the historical state information includes the historical time, the position information and the speed information corresponding to at least one first preset object, and determining the position prediction information corresponding to each of the at least one first preset object according to the historical state information includes:

[0022] Obtain the current time;

[0023] According to the historical time and the current time, obtain the time difference information;

[0024] Determine the position prediction information corresponding to each of the at least one combination according to the time difference information, the position information, and the speed information.

[0025] Optionally, when the image acquisition type is a cross-camera acquisition type, determine at least one combination based on the current image and the historical image, including:

[0026] Perform recognition processing on the image corresponding to the position prediction information in the current image, and use the preset object in the image as the second preset object;

[0027] Combine the first preset object in the historical image with each second preset object to obtain the at least one combination, where the image acquisition device corresponding to the current image and the image acquisition device corresponding to the historical image are adjacently arranged.

[0028] Optionally, after obtaining the current image and the historical image, further include:

[0029] When the preset object exists in the current image and does not exist in the historical image, update the object trajectory information, initialize the historical state information, and return to the step of obtaining the current image and the historical image.

[0030] Optionally, the updating the object trajectory information according to the weight information includes:

[0031] Determine a target combination according to the weight information;

[0032] Based on the target combination, perform object association matching and update the object trajectory information according to the matched objects.

[0033] On the other hand, the present invention also provides an object tracking device, and the device includes:

[0034] A first information acquisition module, configured to acquire a current image and a historical image;

[0035] A second information acquisition module, configured to, when a preset object exists in the historical image, acquire the historical state information of at least one first preset object in the historical image based on the historical image;

[0036] A first information determination module, configured to determine the position prediction information corresponding to each of the at least one first preset object according to the historical state information;

[0037] An image acquisition type determination module, configured to determine the image acquisition type according to the position prediction information;

[0038] A second information determination module, configured to, when the image acquisition type is a cross-camera acquisition type, determine at least one combination based on the current image and the historical image, and determine distance information and corresponding first feature information corresponding to the at least one combination;

[0039] A third information determination module, configured to determine weight information corresponding to each of the at least one combination according to the distance information and the first feature information;

[0040] An association matching module, configured to update object trajectory information according to the weight information.

[0041] On the other hand, the present invention further provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the above object tracking method.

[0042] On the other hand, the present invention further provides a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein, when the computer program instructions are executed by a processor, the above object tracking method is implemented.

[0043] An object tracking method, device, electronic device and storage medium provided by the present invention determine position prediction information through a historical image, determine an image acquisition type based on the position prediction information, and determine at least one combination and weight information corresponding to each combination based on the image acquisition type, and update object trajectory information through the weight information, realizing target tracking for surround view autonomous driving, reducing the possibility of an object being lost across cameras, and reducing time consumption to a certain extent, improving target tracking efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0045] Figure 1 is a flowchart of a method for an object tracking method provided by an embodiment of the present invention;

[0046] Figure 2 is a flowchart of a method after determining an image acquisition type according to position prediction information provided by an embodiment of the present invention;

[0047] Figure 3 is a flowchart of a method for determining an image acquisition type according to position prediction information provided by an embodiment of the present invention;

[0048] Figure 4 It is a flowchart of a method for determining position prediction information corresponding to at least one first preset object according to historical state information provided by an embodiment of the present invention;

[0049] Figure 5 It is a flowchart of a method for determining at least one combination based on a current image and a historical image when an image acquisition type is a cross-camera acquisition type provided by an embodiment of the present invention;

[0050] Figure 6 It is a flowchart of a method after acquiring a current image and a historical image provided by an embodiment of the present invention;

[0051] Figure 7 It is a flowchart of a method for updating object trajectory information according to weight information provided by an embodiment of the present invention;

[0052] Figure 8 It is a structural block diagram of an object tracking device provided by an embodiment of the present invention;

[0053] Figure 9 It is a schematic diagram of a preset range in an object tracking method provided by an embodiment of the present invention. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0056] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0057] The following describes an embodiment of an object tracking method of the present invention. Figure 1 It is a flowchart of an object tracking method provided by an embodiment of the present invention. It should be noted that this specification provides method operation steps as described in the embodiment or flowchart, but based on routine or non-creative labor, there may be more or fewer operation steps. The step order listed in the embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system product is executed, it can be executed in the order of the method shown in the embodiment or the accompanying drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). As Figure 1 As shown, this embodiment provides an object tracking method, which includes:

[0058] S101. Obtain the current image and the historical image.

[0059] Among them, the current image may refer to the images collected by multiple image acquisition devices at the current moment. The historical image may refer to the previous frame image relative to the current image among the images collected by multiple image acquisition devices. The historical image may be the latest frame image with an existing target trajectory. The multiple image acquisition devices may be arranged around the autonomous vehicle.

[0060] In practical applications, during the operation of the autonomous vehicle, it can control multiple image acquisition devices around the vehicle to collect images in real time, and store the collected images and the image acquisition time in the memory of the image acquisition device correspondingly. The controller of the object tracking device can obtain the image corresponding to the current time and the previous frame image of this image from the memory of the image acquisition device according to the current time, and use the image corresponding to the current time and the previous frame image of this image as the current image and the historical image respectively.

[0061] S102. When there is a preset object in the historical image, based on the historical image, obtain the historical state information of at least one first preset object in the historical image.

[0062] The preset object can refer to an object whose ratio of the width to the height of the detection box for object recognition in the image is less than a preset value; for example, the preset object can be a pedestrian. The first preset object can refer to the preset object in the historical image. It can be understood that there may or may not be a preset object in the historical image; further, there may be one preset object or multiple preset objects in the historical image. The historical state information can be the state information of at least one first preset object in the historical image. The historical state information can characterize the motion state of the first preset object at the moment corresponding to the historical image. The historical state information can include the shooting moment corresponding to the historical image, the position (x, y) of the first preset object in the coordinate system corresponding to the historical image, the width w and height h of the detection box of the first preset object in the historical image, and the change speeds corresponding to the above four variables (x, y, w, h).

[0063] In practical applications, when performing image recognition processing on the historical image, after identifying the preset object, the preset object can be used as the first preset object; if multiple preset objects are detected, multiple preset objects can be used as multiple first preset objects. Establish a coordinate system in the historical image, determine the position information of the midpoint of the detection box of the first preset object in the coordinate system of the historical image, and use the position information of at least one first preset object as the current coordinate information. By comparing the historical image with the previous frame of the historical image, the change speeds corresponding to the four variables in the historical image can be determined, so that the historical state information of the first preset object can be obtained.

[0064] S103. Determine the position prediction information corresponding to each of at least one first preset object according to the historical state information.

[0065] Among them, the position prediction information can refer to the information for predicting the position of the first preset object corresponding to the current frame. The position prediction information can be the position in the image coordinate system of the image acquisition device corresponding to the shooting range where the first preset object is located.

[0066] In practical applications, according to the historical state information, the position prediction information of each first preset object can be obtained through a Kalman filter.

[0067] S104. Determine the image acquisition type according to the position prediction information.

[0068] Among them, the image acquisition type can characterize the relative position relationship between the image acquisition device corresponding to the shooting range where the first preset object may appear in the current frame and the image acquisition device corresponding to it in the previous frame. The image acquisition type can include a cross-camera acquisition type and a non-cross-camera acquisition type.

[0069] In practical applications, if the position of the first preset object in the previous frame corresponds to the first image acquisition device, and it can be predicted according to the position prediction information that the first preset object may appear in the second image acquisition device in the current frame, the second image acquisition device is at an adjacent position to the first image acquisition device, and the image acquisition type can be a cross-camera acquisition type; similarly, if it can be predicted according to the position prediction information that the first preset object may still appear in the first image acquisition device in the current frame, the image acquisition type can be a non-cross-camera acquisition type.

[0070] S105. In the case where the image acquisition type is a cross-camera acquisition type, based on the current image and the historical image, determine at least one combination, and determine the distance information and the corresponding first feature information corresponding to at least one combination.

[0071] Among them, a combination can refer to a combination composed of the first preset object in the historical image and each second preset object in the current image. The distance information can refer to the Euclidean distance between the histogram features corresponding to the detection frames of the first preset object and the second preset object in each combination after being unified in size. The histogram feature can characterize the numerical distribution of the image within the detection frame. The first feature information corresponding to each combination can characterize the similarity between the first preset object and the second preset object in the combination.

[0072] In practical applications, the detection frames of the first preset object and the second preset object in each combination can be processed to have the same size; after the detection frames of the two are unified in size, the histogram features of the two detection frames are obtained, and then the Euclidean distance between the two histogram features is calculated. Use the Person Re-identification (ReID) model to extract features from the detection frames of the first preset object and the second predicted object in each combination; through the two extracted feature information, the feature similarity between the first preset object and the second preset object in each combination can be calculated, and this feature similarity can be used as the first feature information.

[0073] S106. According to the distance information and the first feature information, determine the weight information corresponding to each of the at least one combination.

[0074] Among them, the weight information of each combination can be used to characterize the matching degree between the first preset object and the second preset object in the combination.

[0075] In practical applications, according to the distance information corresponding to each combination and the first feature information corresponding to the combination, the weight information of the combination can be calculated.

[0076] Specifically, the weight information W of each combination can be calculated according to the following formula:

[0077] W = w1dist(F H1 , F H2 ) + w2cos(Fea1, Fea2) + w3(|w1 / h1 - w2 / h2|)

[0078] Wherein, w1, w2, and w3 are preset parameters respectively, preferably; w1, w2, and w3 can characterize the proportional relationship between dist(F H1 , F H2 ), cos(Fea1, Fea2), and (w1 / h1 - w2 / h2). w1, w2, and w3 can be set according to experiments. Through dist(F H1 , F H2 ), the Euclidean distance between the histogram features of two detection frames can be calculated; through cos(Fea1, Fea2), the feature similarity can be calculated as the first feature information. Specifically, F H1 and F H2 are respectively the histogram features of the latest detection frame of the existing trajectory and the detection frame of the current frame after being unified in size. Fea1 and Fea2 respectively represent the feature information extracted by the pedestrian ReID model. Fea1 is the feature information inside the detection frame of the latest frame of the existing target trajectory (i.e., the previous frame of the current frame), and Fea2 is the feature information of the detection frame of the current frame. h1 and h2 respectively represent the width and height of the latest detection frame of the existing trajectory and the width and height of the detection frame of the current frame.

[0079] Specifically, the calculation formula of the Euclidean distance is:

[0080] S107. Update the object trajectory information according to the weight information.

[0081] Wherein, the object trajectory information may refer to the trajectory information of a preset object captured by an image acquisition device installed on a vehicle.

[0082] In practical applications, based on the weight information corresponding to at least one combination, the objects can be matched according to the KM (Kuhn - Munkres) algorithm (Hungarian matching algorithm with weights), and the object trajectory information can be updated according to the matching result.

[0083] Determine the position prediction information through historical images, determine the image acquisition type based on the position prediction information, and based on the image acquisition type, determine at least one combination and the weight information corresponding to each combination. Update the object trajectory information through the weight information, realizing the target tracking of surround - view automatic driving, reducing the possibility of objects being lost across cameras, and to a certain extent reducing the time consumption and improving the target tracking efficiency.

[0084] Figure 2It is a flowchart of a method for determining an image acquisition type according to position prediction information provided by an embodiment of the present invention. In a possible implementation manner, as Figure 2 shown, after determining the image acquisition type according to the position prediction information, it may further include:

[0085] S201. When the image acquisition type is a non-cross-camera acquisition type, based on the current image and the historical image, determine at least one combination, and determine the intersection over union information corresponding to at least one combination and the corresponding second feature information.

[0086] Among them, the intersection over union information may refer to the intersection over union of the detection box of the first preset object and the detection box of the second preset object in the combination; specifically, the intersection over union may refer to the area ratio of the intersection and the union of two rectangular boxes. The second feature information may characterize the similarity between the first preset object and the second preset object in the combination.

[0087] In practical applications, the intersection over union of the two can be directly calculated according to the detection box of the first preset object and the detection box of the second preset object in the combination, or the intersection over union of the two after expansion can be calculated after expanding the above two detection boxes. In this embodiment, the method of calculating the intersection over union after expansion is used to calculate the intersection over union W IOU , and the specific calculation process is as follows:

[0088]

[0089] R e = w v v host + w yaw v yaw_rate

[0090] Among them, Rc1 and Rc2 respectively represent two detection boxes participating in the calculation, and R e represents the expansion ratio of the detection box; v host and v yaw_rate respectively represent the vehicle speed and the corresponding yaw angular velocity. w v and w yaw are respectively preset parameters, characterizing the corresponding weight relationship. w v and w yaw can be determined through multiple experiments; preferably, w v and w yaw can be 0.6 and 0.4 respectively, with obvious tracking effect and less false matching.

[0091] S202. According to the intersection over union information and the second feature information, determine the weight information corresponding to each of at least one combination.

[0092] In practical applications, the cosine distance can be used to measure the feature similarity, thereby obtaining the second feature information. The weight information W of each combination can be obtained through the intersection-over-union information and the second feature information. The specific calculation formula is as follows:

[0093] W = w1W IOU + w2 cos(Fea1,Fea2)

[0094] where w1 and w2 are preset parameters respectively, representing the proportional relationship between W IOU and cos(Fea1,Fea2). Preferably, when w1 and w2 are 0.6 and 0.4 respectively, the result of target association matching is the best. Fea1 and Fea2 respectively represent the feature information extracted by using the pedestrian ReID model. Fea1 is the feature information inside the latest frame detection box of the existing target trajectory, and Fea2 is the feature information of the current frame detection box.

[0095] It should be noted that in the case where the image acquisition type is a non-cross-camera acquisition type, the method for determining at least one combination can be: in the case where the image acquisition type is a non-cross-camera acquisition type, it can be determined that the image acquisition device corresponding to the current image is the same device as the image acquisition device corresponding to the first preset image; the current image with the same device identifier as the device identifier corresponding to the historical image where the first preset object is located in the combination can be obtained for analysis. The current image is identified and processed; in the case where the preset object exists in the current image, at least one preset object can be identified as the second preset object respectively. The first preset object is combined with each second preset object in the current image. It can be understood that the number of combinations is the same as the number of second preset objects.

[0096] After determining the image acquisition type, different weight determination methods can be corresponding to different image acquisition types; after determining that the image acquisition type is a non-cross-camera acquisition type, targeted processing can be performed after determining the image to be processed, which can reduce the amount of data for analysis and improve the analysis efficiency.

[0097] Figure 3 is a flowchart of a method for determining an image acquisition type according to position prediction information provided by an embodiment of the present invention. In a possible implementation manner, as Figure 3 shown, the above step S104 may include:

[0098] S301. When the position point corresponding to the position prediction information does not belong to the preset range, determine that the image acquisition type is a cross-camera acquisition type.

[0099] Among them, the preset range can be determined according to the shooting range of each image acquisition device. The preset range can be the same as the shooting range of the image acquisition device, or can be a range slightly smaller than the shooting range of the image acquisition device, and the present disclosure does not specifically limit this. In this embodiment, the preset range is a range smaller than the shooting range of the image acquisition device. The position point corresponding to the position prediction information can be a point representing the predicted position of the first preset object in the current image. Specifically, the position point corresponding to the position prediction information can be the midpoint of the bottom edge of the predicted detection frame.

[0100] In practical applications, through the position prediction information, the relative position of the predicted detection frame of each first preset object with respect to the historical image at the current moment can be obtained; furthermore, the specific position of the midpoint of the bottom edge of the detection frame and the preset range can be determined, and the image acquisition type can be determined according to whether the midpoint of the bottom edge is within the preset range. It can be understood that when the midpoint of the bottom edge is within the preset range, the image acquisition type is a non-cross-camera acquisition type; when the midpoint of the bottom edge is outside the preset range, the image acquisition type is a cross-camera acquisition type. For example, as Figure 9 shown, in the figure, W and H respectively represent the width and height of the image, the gray area represents the preset range, and A and B respectively represent the detection frames of two preset objects; among them, the midpoint of the bottom edge of A is within the preset range, and the image acquisition type is determined to be a non-cross-camera acquisition type; the midpoint of the bottom edge of B is outside the preset range, and the image acquisition type is determined to be a cross-camera acquisition type.

[0101] S302. When the position point corresponding to the position prediction information belongs to the preset range, determine that the image acquisition type is a non-cross-camera acquisition type.

[0102] Through the positional relationship between the position point corresponding to the position prediction information and the preset range, the relative positional relationship between the first preset object at the current moment and the image acquisition device corresponding to the historical moment can be determined, so that the image acquisition type can be determined quickly and accurately.

[0103] Figure 4 FIG. is a flowchart of a method for determining position prediction information corresponding to at least one first preset object according to historical state information provided by an embodiment of the present invention. In a possible implementation manner, the historical state information includes the historical moment, position information and speed information corresponding to at least one first preset object, as Figure 4 shown, the above step S103 may include:

[0104] S401. Obtain the current moment.

[0105] Among them, the current moment may refer to the shooting moment corresponding to the current image.

[0106] In practical applications, when an image acquisition device acquires an image, it will synchronously save the shooting time. The current time can be obtained according to the identification information of the current image.

[0107] S402. Obtain time difference information according to the historical time and the current time.

[0108] Among them, the historical time can refer to the shooting time corresponding to the historical image, that is, the shooting time corresponding to the previous frame of image. The time difference information can refer to the difference information between the historical time and the current time.

[0109] In practical applications, the time difference information can be obtained by taking the difference between the current time and the historical time.

[0110] S403. Determine the position prediction information corresponding to each of at least one combination according to the time difference information, the position information, and the speed information.

[0111] Among them, the position information of each first preset object can refer to the position coordinates of the first preset object in the historical image. Specifically, the position information can include the position (x, y) of the first preset object in the coordinate system corresponding to the historical image, the width w and the height h of the detection frame of the first preset object in the historical image. The speed information of each object can include the change speeds corresponding to the above four variables (x, y, w, h).

[0112] In practical applications, according to the time difference information, the position information, and the speed information, the position of the first preset object corresponding to the current time in the historical image coordinate system and the width and height of the detection frame can be calculated.

[0113] Figure 5 It is a flowchart of a method for determining at least one combination based on a current image and a historical image in the case where the image acquisition type is a cross-camera acquisition type provided by an embodiment of the present invention. In a possible implementation manner, as Figure 5 shown, in the case where the image acquisition type is a cross-camera acquisition type, determining at least one combination based on the current image and the historical image may include:

[0114] S501. Perform recognition processing on the image corresponding to the position prediction information in the current image, and use the preset object in the image as the second preset object.

[0115] In practical applications, the current image may include images captured by multiple image acquisition devices at the current moment. According to the relative position relationship between the position prediction information and the image where the first preset object is located, and the image acquisition device corresponding to the shooting range where the first preset object is located, the device identifier of the image acquisition device corresponding to the position prediction information can be determined. The image corresponding to the device identifier in the current image is the image corresponding to the position prediction information in the current image. It can be understood that the position prediction information can reflect the image acquisition device corresponding to the shooting area where the first preset object may appear at the current moment, and the image corresponding to the position prediction information in the current image can be determined according to the position prediction information. Perform recognition processing on this image, and use the preset object in the image as the second preset object. For example, when 2 preset objects are recognized in the image, the above 2 preset objects are respectively used as 2 second preset objects.

[0116] S502. Combine the first preset object in the historical image with each second preset object to obtain at least one combination, and the image acquisition device corresponding to the current image and the image acquisition device corresponding to the historical image are adjacently arranged.

[0117] It can be understood that two first preset objects may respectively correspond to different images and different image acquisition types according to the position prediction information. In the case where the image acquisition type of one of the first preset objects is the cross-camera acquisition type, combining this first preset object with each second preset object can obtain the same number of combinations as the second preset object. For example, assuming there are 2 second preset objects (object 21 and object 22 respectively), and the first preset object is object 11, then 2 combinations can be combined, including the first combination (object 11 and object 21) and the second combination (object 11 and object 22) respectively.

[0118] Figure 6 It is a flowchart of a method after obtaining the current image and the historical image provided by an embodiment of the present invention. In a possible implementation manner, as Figure 6 shown, after obtaining the current image and the historical image, it may further include:

[0119] S601. In the case where the current image has a preset object and the historical image does not have a preset object, update the object trajectory information, initialize the historical state information, and return to the step of obtaining the current image and the historical image.

[0120] It can be understood that in the case where the current image has a preset object and the historical image does not have a preset object, it can be explained that the current image is the first frame image in the trajectory corresponding to the preset object.

[0121] In practical applications, when a preset object exists in the current image but does not exist in the historical image, the trajectory information of the preset object can be newly added to the object trajectory information to update the object trajectory information, and the historical state information of the preset object is initialized. Specifically, the position of the detection box of the preset object in the coordinate system of the historical image is used as the position information in the historical state information of the preset object; the change speeds of the four variables therein are initialized to 0. After returning to the step of obtaining the current image and the historical image, the preset object will be used as the first preset object, and the historical state information of the preset object obtained in the subsequent steps is the above-mentioned initialized information.

[0122] Figure 7 It is a flowchart of a method for updating object trajectory information according to weight information provided by an embodiment of the present invention. In a possible implementation manner, as Figure 7 shown, the above step S107 may include:

[0123] S701. Determine a target combination according to the weight information.

[0124] Among them, the target combination can represent that the first preset object and the second preset object in the combination have a good matching degree.

[0125] In practical applications, according to the weight information of at least one combination, based on the KM algorithm, the combination corresponding to the perfect matching with the maximum weight value of the weighted bipartite graph is used as the target combination.

[0126] S702. Based on the target combination, perform object association matching and update the object trajectory information according to the matched objects.

[0127] In practical applications, the first preset object in the target combination corresponds to a trajectory, the second preset object in the target combination is added to the trajectory, and the object trajectory information of the trajectory is updated according to the relevant information of the second preset object.

[0128] Figure 8 It is a structural block diagram of an object tracking device provided by an embodiment of the present invention. On the other hand, as Figure 8 shown, this embodiment also provides an object tracking device, and the device includes:

[0129] The first information acquisition module 10 is configured to acquire a current image and a historical image;

[0130] The second information acquisition module 20 is configured to, when a preset object exists in the historical image, acquire the historical state information of at least one first preset object in the historical image based on the historical image;

[0131] The first information determination module 30 is configured to determine the position prediction information corresponding to each of at least one first preset object according to the historical status information;

[0132] The image acquisition type determination module 40 is configured to determine the image acquisition type according to the position prediction information;

[0133] The second information determination module 50 is configured to, when the image acquisition type is the cross-camera acquisition type, determine at least one combination based on the current image and the historical image, and determine the distance information and the corresponding first feature information corresponding to at least one combination;

[0134] The third information determination module 60 is configured to determine the weight information corresponding to each of at least one combination according to the distance information and the first feature information;

[0135] The association matching module 70 is configured to update the object trajectory information according to the weight information.

[0136] On the other hand, an embodiment of the present invention further provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the above object tracking method.

[0137] On the other hand, an embodiment of the present invention further provides a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein, when the computer program instructions are executed by a processor, the above object tracking method is implemented.

[0138] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be adopted in other sequences or simultaneously. Similarly, each module of the above object tracking device refers to a computer program or a program segment for executing one or more specific functions. In addition, the distinction of the above modules does not mean that the actual program codes must also be separated. In addition, the above embodiments can be arbitrarily combined to obtain other embodiments.

[0139] In the above embodiments, the descriptions of the embodiments each have their own emphasis. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. Those skilled in the art can also understand that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention can be implemented by electronic hardware, computer software, or a combination of both. To clearly show the interchangeability of hardware and software, the above various illustrative components, units, and steps have been generally described in terms of their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art can use various methods to implement the described functions for each specific application, but such implementation should not be construed as exceeding the scope protected by the embodiments of the present invention.

[0140] The above description has fully disclosed the specific implementation manners of the present invention. It should be noted that any modifications made by those skilled in the art to the specific implementation manners of the present invention do not depart from the scope of the claims of the present invention. Accordingly, the scope of the claims of the present invention is not limited solely to the foregoing specific implementation manners.

Claims

1. An object tracking method, characterized in that, The method includes: Obtaining a current image and a historical image; When there is a preset object in the historical image, based on the historical image, obtaining historical state information of at least one first preset object in the historical image; According to the historical state information, determining position prediction information corresponding to each of the at least one first preset object; According to the position prediction information, determining an image acquisition type; the image acquisition type represents the relative position relationship between an image acquisition device corresponding to a possible shooting range of a first preset object in the current frame and an image acquisition device corresponding to the first preset object in the previous frame; the image acquisition type includes a cross-camera acquisition type, and the cross-camera acquisition type is used to indicate that the image acquisition device corresponding to the position of the first preset object in the previous frame is different from the image acquisition device corresponding to the position where the first preset object appears in the current frame predicted according to the position prediction information; When the image acquisition type is the cross-camera acquisition type, based on the current image and the historical image, determining at least one combination, and determining distance information and corresponding first feature information corresponding to the at least one combination; According to the distance information and the first feature information, determining weight information corresponding to each of the at least one combination; According to the weight information, updating object trajectory information.

2. The method according to claim 1, wherein After determining the image acquisition type according to the position prediction information, it further includes: When the image acquisition type is a non-cross-camera acquisition type, based on the current image and the historical image, determining at least one combination, and determining intersection over union information and corresponding second feature information corresponding to the at least one combination; According to the intersection over union information and the second feature information, determining weight information corresponding to each of the at least one combination.

3. The method according to claim 1, wherein Determining the image acquisition type according to the position prediction information includes: When the position point corresponding to the position prediction information does not belong to a preset range, determining that the image acquisition type is the cross-camera acquisition type; When the position point corresponding to the position prediction information belongs to a preset range, determining that the image acquisition type is a non-cross-camera acquisition type.

4. The method according to claim 1, wherein The historical state information includes a historical time, position information and speed information corresponding to at least one first preset object, and determining the position prediction information corresponding to each of the at least one first preset object according to the historical state information includes: Obtaining the current time; According to the historical time and the current time, obtaining time difference information; According to the time difference information, the position information and the speed information, determining the position prediction information corresponding to each of the at least one combination.

5. The method according to claim 1, wherein When the image acquisition type is the cross-camera acquisition type, determining at least one combination based on the current image and the historical image includes: Performing recognition processing on the image corresponding to the position prediction information in the current image, and taking the preset object in the image as a second preset object; Combine the first preset object in the historical image with each second preset object to obtain the at least one combination, where the image acquisition device corresponding to the current image and the image acquisition device corresponding to the historical image are adjacently arranged.

6. The method according to claim 1, wherein After acquiring the current image and the historical image, it further includes: In the case where the preset object exists in the current image and does not exist in the historical image, update the object trajectory information, initialize the historical state information, and return to the step of acquiring the current image and the historical image.

7. The method according to claim 1, wherein Updating the object trajectory information according to the weight information includes: Determine the target combination according to the weight information; Based on the target combination, perform object association matching and update the object trajectory information according to the matched objects.

8. An object tracking device, characterized in that, The device includes: A first information acquisition module for acquiring a current image and a historical image; A second information acquisition module for, in the case where a preset object exists in the historical image, acquiring the respective historical state information of at least one first preset object in the historical image based on the historical image; A first information determination module for determining the respective position prediction information corresponding to the at least one first preset object according to the historical state information; An image acquisition type determination module for determining the image acquisition type according to the position prediction information; the image acquisition type characterizes the relative position relationship between the image acquisition device corresponding to the possible shooting range of the first preset object in the current frame and the image acquisition device corresponding to it in the previous frame; the image acquisition type includes a cross-camera acquisition type, and the cross-camera acquisition type is used to indicate that the image acquisition device corresponding to the position of the first preset object in the previous frame and the image acquisition device corresponding to the position where the first preset object appears in the current frame predicted according to the position prediction information are different image acquisition devices; A second information determination module for, in the case where the image acquisition type is a cross-camera acquisition type, determining at least one combination based on the current image and the historical image, and determining the distance information and the corresponding first feature information corresponding to the at least one combination; A third information determination module for determining the respective weight information corresponding to the at least one combination according to the distance information and the first feature information; An association matching module for updating the object trajectory information according to the weight information.

9. An electronic device, characterized in that, Includes: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to execute the executable instructions to implement the object tracking method according to any one of claims 1 to 7.

10. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the object tracking method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Route trajectory data exception detection method, system and equipment and storage medium

    CN109684916A

  • Human body behavior recognition method and system based on multi-target tracking

    CN110399808A