A target tracking method and device

By introducing re-identification features and motion prediction in target tracking, setting search areas and matching features, the problem of over-reliance of target tracking results in the prior art is solved, and the accuracy and reliability of target tracking are improved.

CN115761655BActive Publication Date: 2025-06-27ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211441963.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2025-06-27
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

Existing target tracking algorithms rely too much on the detector's detection results, resulting in inaccurate target tracking results, especially when the target shape or size changes.

Method used

By extracting the feature maps of the previous image frame and the current image frame, re-identification features are obtained, and the prediction offset of the target in the current frame is determined based on the motion prediction results, and then setting the search area to perform feature matching to determine the tracking result.

Benefits of technology

Improve the accuracy and reliability of target tracking, reduce tracking errors caused by inaccurate detection results, and enhance the accuracy of matching targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761655B_ABST
    Figure CN115761655B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a target tracking method and apparatus to improve the accuracy of target tracking. The method includes: extracting features from the feature map of the previous image frame and the feature map of the current image frame to obtain the first-level recognition features of the previous image frame and the second-level recognition features of the current image frame, and fusing the first feature map and the second feature map and performing motion prediction to obtain the predicted target offset of each first target object; for each first target object: determining a set area of the first target object in the current image frame according to the first target information of the first target object and the predicted target offset; determining the feature similarity according to the first-level recognition features and the second-level recognition features included in the set area, and determining the position distance according to the first target information and the second target information within the set area; determining the tracking result of each first target object according to the feature similarity and the position distance of each first target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of computer vision target tracking, and particularly to a target tracking method and device. Background Art

[0002] With the development of computer vision technology based on deep learning, computer vision technology has been rapidly applied in the transportation field. By detecting and tracking targets such as people and vehicles in surveillance videos and performing real-time analysis, this technology can complete the timely reporting of traffic problems, improve the efficiency of traffic management work, and relieve highway operation and maintenance and police force resources. Existing target tracking algorithms are detection-based tracking, that is, usually existing detectors detect targets, and then the tracking algorithm performs data association of targets in the front and back frames. However, the target tracking result of this method is too dependent on the detection result of the detector. When the detection result of the target is inaccurate due to the shape and size of the target, it will affect the accuracy of target tracking. Summary of the Invention

[0003] Embodiments of this application provide a target tracking method and device to improve the accuracy of target tracking.

[0004] In a first aspect, embodiments of this application provide a target tracking method, including:

[0005] Performing feature extraction on the first feature map of the previous image frame and the second feature map of the current image frame respectively to obtain the first re-identification feature of each first target object in the previous image frame and the second re-identification feature of each second target object in the current image frame, and fusing the first feature map and the second feature map, and performing motion prediction on the fused feature map to obtain the predicted target offset of each first target object in the current image frame;

[0006] For each first target object: determining the set area of the first target object in the current image frame according to the first target information of the first target object and the predicted target offset of the first target object; performing feature matching between the first re-identification feature of the first target object and the second re-identification feature of the second target object included in the set area to obtain at least one feature similarity, and matching the first target information of the first target object with the second target information of the second target object included in the set area to obtain at least one position distance;

[0007] Determining the tracking result of each first target object according to at least one feature similarity and at least one position distance of each first target object.

[0008] Based on the above solution, when the present application performs matching of target objects, re-identification features and motion prediction are introduced, and the matching result is jointly determined by the re-identification features and the motion prediction result, avoiding the problem of tracking errors caused by only using target bounding boxes or feature maps to track two target objects with close distances or similar appearances in existing target tracking methods, making the matching result more accurate and reliable, and improving the tracking effect of the target.

[0009] In a possible implementation manner, the first feature map and the second feature map are determined by respectively performing feature extraction on the previous image frame and the current image frame through a first network in the target tracking network. The first target information and the second target information are determined by respectively performing feature extraction on the first feature map and the second feature map through a detection head network in the target tracking network. The target tracking network further includes a re-identification network and a motion network. The step of respectively performing feature extraction on the first feature map of the previous image frame and the second feature map of the current image frame to obtain the first re-identification feature of each first target object in the previous image frame and the second re-identification feature of each second target object in the current image frame includes:

[0010] Inputting the previous image frame and the current image frame into the re-identification network respectively for feature extraction, so that the re-identification network outputs the first re-identification feature of each first target object in the previous image frame and the second re-identification feature of each second target object in the current image frame;

[0011] The step of fusing the first feature map and the second feature map and performing motion prediction on the fused feature map to obtain the predicted target offset of each first target object in the current image frame includes:

[0012] Adding the first feature map and the second feature map bit by bit to obtain a fused feature map, and inputting the fused feature map into the motion network for motion prediction, so that the motion network outputs the predicted target offset of each first target object in the current image frame.

[0013] In a possible implementation, the first target information includes the position of the first target center point, the width of the first target object, and the height of the first target object. Determining the set area of the first target object in the current image frame according to the first target information of the first target object and the predicted target offset of the first target object includes: determining the predicted target center point position of the first target object in the current image frame according to the position of the first target center point and the predicted target offset of the first target object; taking the area obtained by magnifying the width of the first target object and the height of the first target object by a set multiple with the predicted target center point position as the center as the set area.

[0014] Based on the above solution, through the target position of the previous frame, the width and height are magnified by a certain proportion as the search area of the target in the next frame. Through the limitation of the area, invalid matching of targets with too large a distance between the previous and next frames can be reduced, and the running speed can be improved.

[0015] In a possible implementation, the method further includes: before matching the first re-identification feature of the first target object with the second re-identification feature of the second target object included in the set area, determining the third re-identification feature of the target object with the same identification number as the first target object in the image frame before the previous image frame; matching the first re-identification feature of the first target object with the second re-identification feature of the second target object included in the set area to obtain at least one feature similarity, including: determining the top S re-identification features with higher confidence levels among the third re-identification feature and the first re-identification feature; for each second target object in the set area, determining the feature similarity according to the S re-identification features and the second re-identification feature.

[0016] In a possible implementation, the feature similarity satisfies the conditions shown in the following formula:

[0017]

[0018] Among them, the dis reid represents the feature similarity, represents the feature vector of the kth re-identification feature among the S re-identification features, and n R represents the feature vector of the second re-identification feature.

[0019] Based on the above solution, compared with the traditional cosine similarity calculation formula, this formula determines the final feature similarity through the mean value of the similarities between multiple re-identification features with higher confidence levels corresponding to the first target object and the second re-identification feature, which can reduce the contingency and risk of a single feature similarity result and improve the robustness of the algorithm.

[0020] In a possible implementation, the second target information includes the position of the second target center point. The first target information of the first target object is matched with the second target information of the second target object included in the set area to obtain at least one position distance, including:

[0021] For each second target object in the set area, the position distance is determined by the predicted target center point position of the first target object, the second target center point position, and the width of the first target object.

[0022] In a possible implementation, the position distance satisfies the conditions described by the following formula:

[0023]

[0024] where the dis g represents the position distance, the x represents the predicted target center point position, the x′ represents the second target center point position, and σ represents the width of the first target object.

[0025] Based on the above solution, compared with the general Euclidean distance formula, the distance between two target points in this application does not increase or decrease infinitely with the change of position, but is strictly controlled within [0,1]. This effectively avoids the large difference in the calculated distance similarity when the two targets are far away from each other. At the same time, it fuses better with the feature similarity distance than the simple Euclidean distance. In addition, the formula has distance weighting. Within a certain range, the distance similarity changes violently with the increase of the distance between the two targets. When it exceeds the set range, the distance similarity becomes less obvious with the increase of the distance and gradually tends to be stable. This non-linear distance similarity calculation method is more in line with the actual matching law between targets.

[0026] In a possible implementation, determining the tracking result of each first target object according to at least one feature similarity and at least one position distance of each first target object includes: for each first target object, weighting each feature similarity and its corresponding position distance to obtain the matching weight of each second target object; taking the second target object with the smallest matching weight as the tracking result of the first target object.

[0027] Based on the above solution, the target tracking results of the front and rear frames are jointly determined by the position distance and the feature similarity, reducing the dependence on a single weight and improving the accuracy of the tracking result.

[0028] In a possible implementation, the first target information further includes a first identification number, the second target information further includes a second identification number and a second target status. After using the second target object with the minimum matching weight as the tracking result of the first target object, the method further includes: setting the second identification number as the first identification number, and setting the second target status as an updated status.

[0029] In a possible implementation, the method further includes: for each second target object, when the second identification number of the second target object is blank, setting the second identification number of the second target object as an identification number different from the identification numbers of other target objects, and setting the second target status of the second target object as a created status.

[0030] In a possible implementation, the first target information further includes a first target status. The method further includes: if the position distances between the first target object and each second target object included in the set area are all greater than a set threshold, setting the first target status as a lost status; when the lost status lasts for a set duration, setting the target status of the first target object as a deleted status.

[0031] In a second aspect, an embodiment of the present application provides a target tracking device, including:

[0032] A first determination module, configured to perform feature extraction on the first feature map of the previous image frame and the second feature map of the current image frame respectively, to obtain the first re-identification features of each first target object in the previous image frame and the second re-identification features of each second target object in the current image frame, and to fuse the first feature map and the second feature map, and perform motion prediction on the fused feature map, to obtain the predicted target offset of each first target object in the current image frame;

[0033] A second determination module, configured to, for each first target object: determine the set area of the first target object in the current image frame according to the first target information of the first target object and the predicted target offset of the first target object; perform feature matching between the first re-identification features of the first target object and the second re-identification features of the second target objects included in the set area, to obtain at least one feature similarity, and perform matching between the first target information of the first target object and the second target information of the second target objects included in the set area, to obtain at least one position distance;

[0034] Determine the tracking result of each first target object according to at least one feature similarity and at least one position distance of each first target object.

[0035] In some embodiments, the first feature map and the second feature map are determined by respectively performing feature extraction on the previous image frame and the current image frame through a first network in the target tracking network. The first target information and the second target information are determined by respectively performing feature extraction on the first feature map and the second feature map through a detection head network in the target tracking network. The target tracking network further includes a re-identification network and a motion network. When the first determination module respectively performs feature extraction on the first feature map of the previous image frame and the second feature map of the current image frame to obtain the first re-identification feature of each first target object in the previous image frame and the second re-identification feature of each second target object in the current image frame, it specifically is used for:

[0036] Input the previous image frame and the current image frame into the re-identification network respectively for feature extraction, so that the re-identification network outputs the first re-identification feature of each first target object in the previous image frame and the second re-identification feature of each second target object in the current image frame;

[0037] When the first determination module fuses the first feature map and the second feature map and performs motion prediction on the fused feature map to obtain the predicted target offset of each first target object in the current image frame, it specifically is used for:

[0038] Add the first feature map and the second feature map bit by bit to obtain the fused feature map, and input the fused feature map into the motion network for motion prediction, so that the motion network outputs the predicted target offset of each first target object in the current image frame.

[0039] In some embodiments, the first target information includes the position of the center point of the first target, the width of the first target object, and the height of the first target object. When the second determination module determines the set area of the first target object in the current image frame according to the first target information of the first target object and the predicted target offset of the first target object, it specifically is used for: determining the predicted target center point position of the first target object in the current image frame according to the position of the center point of the first target and the predicted target offset of the first target object; taking the area obtained by magnifying the width of the first target object and the height of the first target object by a set multiple with the predicted target center point position as the center as the set area.

[0040] In some embodiments, the second determination module is further configured to: before performing feature matching between the first re-identification feature of the first target object and the second re-identification feature of the second target object included in the set area, determine the third re-identification feature of the target object with the same identification number as the first target object in the image frame before the previous image frame; when the second determination module performs feature matching between the first re-identification feature of the first target object and the second re-identification feature of the second target object included in the set area to obtain at least one feature similarity, it is specifically configured to: determine the top S re-identification features with higher confidence among the third re-identification feature and the first re-identification feature; for each second target object in the set area, determine the feature similarity according to the S re-identification features and the second re-identification feature.

[0041] In some embodiments, the feature similarity satisfies the conditions shown in the following formula:

[0042]

[0043] wherein, the dis reid represents the feature similarity, represents the feature vector of the k-th re-identification feature among the S re-identification features, and n R represents the feature vector of the second re-identification feature.

[0044] In some embodiments, the second target information includes the position of the center point of the second target. When the second determination module matches the first target information of the first target object with the second target information of the second target object included in the set area to obtain at least one position distance, it is specifically configured to:

[0045] For each second target object within the set area, determine the position distance based on the predicted center point position of the first target object, the position of the center point of the second target, and the width of the first target object.

[0046] In some embodiments, the position distance satisfies the conditions shown in the following formula:

[0047]

[0048] wherein, the dis g represents the position distance, x represents the predicted center point position, x' represents the position of the center point of the second target, and σ represents the width of the first target object.

[0049] In some embodiments, when the second processing module determines the tracking result of each first target object according to at least one feature similarity and at least one position distance of each first target object, it is specifically configured to:

[0050] For each first target object, weight each feature similarity and its corresponding position distance to obtain the matching weight of each second target object.

[0051] Use the second target object with the minimum matching weight as the tracking result of the first target object.

[0052] In some embodiments, the first target information further includes a first identification number, the second target information further includes a second identification number and a second target status. After using the second target object with the minimum matching weight as the tracking result of the first target object, the second processing module is further configured to:

[0053] Set the second identification number to the first identification number, and set the second target status to the updated status.

[0054] In some embodiments, the second processing module is further configured to: for each second target object, when the second identification number of the second target object is blank, set the second identification number of the second target object to an identification number different from the identification numbers of other target objects, and set the second target status of the second target object to the created status.

[0055] In some embodiments, the first target information further includes a first target status. The second processing module is further configured to: if the position distance between the first target object and each second target object included in the set area is greater than the set threshold, set the first target status to the lost status; when the lost status lasts for the set duration, set the target status of the first target object to the deleted status.

[0056] In a third aspect, an embodiment of the present application provides an execution device, including:

[0057] A memory for storing program instructions;

[0058] A processor for obtaining the program instructions stored in the memory and executing the methods described in the first aspect and different implementation manners of the first aspect according to the obtained program instructions.

[0059] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores computer instructions. When the computer instructions run on a computer, the computer is caused to execute the methods described in the first aspect and different implementation manners of the first aspect.

[0060] The technical effects brought by any implementation manner in the second aspect to the fourth aspect can be referred to the technical effects brought by the first aspect and different implementation manners of the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0062] Figure 1 Schematic diagram of a target tracking method provided by an embodiment of the present application;

[0063] Figure 2A Schematic diagram of an application scenario provided by an embodiment of the present application;

[0064] Figure 2B Schematic diagram of a server structure provided by an embodiment of the present application;

[0065] Figure 3 Schematic diagram of the flow of a target tracking method provided by an embodiment of the present application;

[0066] Figure 4 Schematic diagram of a target tracking provided by an embodiment of the present application;

[0067] Figure 5 Schematic diagram of the training process of a target tracking network provided by an embodiment of the present application;

[0068] Figure 6 Schematic diagram of a target tracking device provided by an embodiment of the present application;

[0069] Figure 7 Schematic diagram of an execution device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application usually described and illustrated in the drawings here can be arranged and designed in various different configurations.

[0071] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0072] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0073] With the development of computer vision technology based on deep learning, computer vision technology has been rapidly applied in the transportation field. This technology can complete the timely reporting of traffic problems, improve the efficiency of traffic management work, and relieve highway operation and maintenance and police resources by detecting and tracking targets such as people and vehicles in surveillance videos and performing real-time analysis. In the prior art, after target detection, matching is performed depending on the positions of the detected target boxes. This method relies on the detection results and is prone to tracking errors. In addition, in the prior art, the target tracker cannot be corrected in a timely manner based on the detection results, and the target trajectory is prone to drift.

[0074] Based on the above problems, the embodiments of the present application provide a target tracking method and device. The target information of the front and rear frames is obtained through a Detection Head network, and feature re-extraction is performed through a Re-Motion Head network to obtain re-identification features and motion prediction results. By determining a set area in the rear frame image according to the motion prediction result and the target information, and determining the position distance by predicting the target position and the target position in the set area, and performing feature matching between the re-identification feature of the target in the front frame image and the re-identification feature of the target object in the set area to obtain the feature similarity. Further, the matching weight can be determined by the position distance and the feature similarity, and the matching weight is sent to the Hungarian algorithm to complete the association and tracking of the same target between different frames. In some scenarios, the target state of the target can be updated according to the association result and the tracking result, such as Figure 1 shown.

[0075] The target tracking method provided by the embodiments of the present application can be implemented by an execution device. In some embodiments, the execution device may be an electronic device, and the electronic device may be implemented by one or more servers. Figure 2A Taking a server 100 as an example. Refer to Figure 2AAs shown, it is a schematic diagram of a possible application scenario provided by an embodiment of the present application, including an electronic device 100 and a collection device 200. The server 100 can be implemented by a physical server or by a virtual server. The server can be implemented by a single server, can be implemented by a server cluster composed of multiple servers, and can be implemented by a single server or a server cluster to implement the target tracking method provided by the present application. The collection device 200 is a device with an image collection function, including electric police devices, electronic monitoring devices, surveillance cameras, video recorders, terminal devices with video collection functions (such as laptops, computers, mobile phones, TVs), etc. The collection device 200 can send the collected video to be detected to the server 100 through a network. Optionally, the server 100 can be connected to the terminal device 300, receive the target tracking task sent by the terminal device 300, and perform target tracking based on the received video to be detected sent by the collection device 200. In some scenarios, the server 100 can send the target tracking result to the terminal device 300. The terminal device 300 can be a TV, a mobile phone, a tablet computer, a personal computer, and so on.

[0076] As an example, refer to Figure 2B As shown, the server 100 may include a processor 110, a communication interface 120, and a memory 130. Of course, other components may also be included in the server 100, Figure 2B which are not shown in the figure.

[0077] The communication interface 120 is used to communicate with the collection device 100 and the terminal device 300, receive the video to be detected sent by the collection device 100, or receive the target tracking task sent by the terminal device 300, or send the target tracking result to the terminal device 300.

[0078] In the embodiment of the present application, the processor 110 can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiment of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiment of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0079] The processor 110 is the control center of the server 100, connecting various parts of the entire server 100 through various interfaces and routes. By running or executing software programs / modules stored in the memory 130 and invoking the data stored in the memory 130, it performs various functions of the server 100 and processes data. Optionally, the processor 110 may include one or more processing units. The processor 110 can be, for example, a control component such as a processor, a microprocessor, a controller, etc., and can be, for example, a general-purpose central processing unit (CPU), a general-purpose processor, a digital signal processing (DSP), an application specific integrated circuits (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.

[0080] The memory 130 can be used to store software programs and modules. The processor 110 executes various functional applications and data processing by running the software programs and modules stored in the memory 130. The memory 130 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to business processing, etc. The memory 130, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 130 can include at least one type of storage medium, for example, it can include flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (RAM), a static random access memory (SRAM), a programmable read-only memory (PROM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic memory, a disk, an optical disc, and so on. The memory 130 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 130 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function for storing program instructions and / or data.

[0081] In some other embodiments, the execution device may be a terminal device. In some scenarios, when the terminal device can receive the video to be detected sent by the acquisition device, it identifies foreign objects in the camera lens according to the video to be detected. The terminal device may include a display device, which may be a liquid crystal display, an organic light-emitting diode (OLED) display, a projection display device, etc., and the present application does not make specific limitations thereto.

[0082] It should be noted that the structures shown above Figure 2A and Figure 2B are only examples, and the embodiments of the present application do not limit this.

[0083] The embodiments of the present application provide a target tracking method. Figure 3 Exemplarily, the process of the target tracking method is shown. This process can be executed by a target tracking device, which may be located in the server 100 as shown in Figure 2B . For example, it may be the processor 110 or the server 100. The target tracking device may also be located in the terminal device. The specific process is as follows:

[0084] 301. Feature extraction is respectively performed on the first feature map of the previous image frame and the second feature map of the current image frame to obtain the first re-identification features of each first target object in the previous image frame and the second re-identification features of each second target object in the current image frame, and the first feature map and the second feature map are fused, and motion prediction is performed on the fused feature map to obtain the predicted target offset of each first target object in the current image frame.

[0085] In some embodiments, the previous image frame and the current image frame are consecutive video frames collected by the acquisition device. The acquisition device may be an electronic police device, an electronic monitoring device, a surveillance camera, a video recorder, a terminal device with a video acquisition function (such as a notebook, a computer, a mobile phone, a TV), etc.

[0086] Exemplarily, after the acquisition device collects the video frames, the server obtains the previous image frame and the current image frame from the acquisition device. Among them, the previous image frame is the previous image frame adjacent to the current image frame. In some scenarios, the previous image frame and the current image frame are two video image frames within the set duration collected by the acquisition device.

[0087] In some embodiments, the server receives a video file sent by a collection device. The video file includes the previous image frame and the current image frame. The video file can be an encoded file of the video. Thus, the server can decode the received video file to obtain the previous image frame and the current image frame. Encoding the video can effectively reduce the file size of the video file, facilitating transmission. Thus, the transmission speed of the video can be increased to improve the efficiency of subsequent confirmation of video events. The manner of obtaining the encoded bitstream data can be any applicable manner, including but not limited to: Real Time Streaming Protocol (RTSP), Open Network Video Interface Forum (ONVIF) standard, or private protocol manner, etc.

[0088] In some embodiments, the object tracking network is composed of a first network, a detection head network, a re-identification network, and a motion network. The first feature map and the second feature map are determined by respectively performing feature extraction on the previous image frame and the current image frame through the first network in the object tracking network. Specifically, the previous image frame and the current image frame can be respectively input into the first network in the object tracking network to obtain the first feature map of the previous image frame and the second image frame of the current image frame. Among them, the first network can be the backbone network Backbone (DLA-34). Optionally, the first network can also be hrnet, resnet, or swin network.

[0089] In some embodiments, feature extraction can be respectively performed on the first feature map of the previous image frame and the second feature map of the current image frame to obtain the first re-identification feature of each first target object in the previous image frame and the second re-identification feature of each second target object in the current image frame. Specifically, the re-identification network can respectively perform feature extraction on the first feature map of the previous image frame and the second image frame of the current image frame, so that the re-identification network outputs the first re-identification feature of each first target object in the previous image frame and the second re-identification feature of each second target object in the current image frame. Among them, the re-identification network can be the Re-ID Embeddings network, and the Re-ID features corresponding to the first feature map and the second feature map are respectively extracted through Re-ID Embeddings.

[0090] In some embodiments, the first feature map and the second feature map can be fused, and the fused feature map is input into a motion network for motion prediction to obtain the predicted target offset of each first target object in the previous image frame in the current image frame. Among them, the motion network can be a Motion-MLP network, and the target offset is used to represent the offset between the predicted target center point position of the first target object in the current image frame and the target center point position of the first target object in the previous image frame. As an example, when there are three first target objects in the previous image frame, the fused feature map can be input into the Motion-MLP network so that the Motion-MLP network outputs the predicted target offsets of the three target objects in the previous image frame in the current image frame respectively. For example, if the target center point position of a certain target object in the previous image frame is (334, 556), the predicted target offset of this target object in the current image frame can be expressed as (80, 69).

[0091] 302. For each first target object: According to the first target information of the first target object and the predicted target offset of the first target object, determine the set area of the first target object in the current image frame.

[0092] In some embodiments, the first target information includes the first target center point position, the width of the first target object, and the height of the first target object. Among them, the width and height of the first target object are determined according to the width and height of the circumscribed rectangle of the first target object. That is, the width of the first target object is equal to the width of the circumscribed rectangle of the first target object, and the height of the first target object is equal to the height of the circumscribed rectangle of the first target object.

[0093] In some embodiments, the first feature map and the second feature map can be respectively input into the detection head network to output the first target information of the first target object and the second target information of the second target object through the detection head network. Among them, the detection head network can be divided into a first detection network, a second detection network, and a third detection network. As an example, after the first feature head is input into the detection head network, the heat map of the target center distribution corresponding to the first feature map can be output through the first detection network, the first target center point position of the first target object corresponding to the first feature map can be output through the second detection network, and the width of the first target object and the height of the first target object in the first feature map can be output through the third detection network. Similarly, the second feature map can be input into the detection head network to output the second target information of the second target object through the detection head network.

[0094] In some embodiments, the first feature map and the second feature map can be added bitwise to obtain a fused feature map, and the fused feature map is input into a motion network for motion prediction, so that the motion network outputs the predicted target offset of each first target object in the current image frame.

[0095] In some embodiments, target tracking can be performed by the target tracking module Re-Motion Tracker. Specifically, the set area of the first target object in the current image frame can be determined according to the first target information of the first target object and the predicted target offset of the first target object. The specific method is as follows: after determining the first target information of the first target object, the predicted target center point position of the first target object in the current image frame can be determined according to the first target center point position and the predicted target offset of the first target object. After determining the predicted target center point position of the first target object, the area obtained by magnifying the width and height of the first target object by a set multiple with the predicted target center point position as the center can be used as the set area. As an example, if the first target center point position of the first target object is (334, 556) and the predicted target offset of the first target object is (80, 69), then the predicted target center point position of the first target object in the current image frame can be expressed as (414, 625). Further, when the width of the first target object is 50 and the height is 40, the area obtained by magnifying the width and height of the first target object by 2 times is the area enclosed by (374, 575) and (454, 675).

[0096] 303. Feature matching is performed between the first re-identification feature of the first target object and the second re-identification feature of the second target object included in the set area to obtain at least one feature similarity, and matching is performed between the first target information of the first target object and the second target information of the second target object included in the set area to obtain at least one position distance.

[0097] In some embodiments, after determining the set region of the first target object in the current image frame, the re-identification feature of the first target object may be feature-matched with the second re-identification feature of the second target object in the set region to obtain at least one feature similarity. As an example, when the set region includes three second target objects, the re-identification feature of the first target object is feature-matched with the second re-identification features of the three second target objects respectively to obtain three feature similarities. Similarly, the first target information of the first target object may be matched with the second target information of the second target objects included in the set region to obtain at least one position distance. As a kind of distance, when the set region includes three second target objects, the first target information of the first target object may be matched with the second target information of the three second target objects respectively to obtain three position distances. It can be understood that the number of second target objects in the set region is equal to the number of feature similarities and is equal to the number of position distances, that is, one second target object corresponds to one feature similarity and one position distance. When the set region includes three second target objects, the first target object corresponds to three feature similarities and three position weights.

[0098] In some embodiments, the first target information further includes a first identification number. Before feature-matching the first re-identification feature of the first target object with the second re-identification feature of the second target object included in the set region, determine the third re-identification feature of the target object corresponding to the first identification number in the image frame before the previous image frame. In some scenarios, when performing target tracking, the re-identification features and confidence levels of each target object with the same identification number in different image frames may be saved. In other scenarios, the S re-identification features with relatively high confidence levels corresponding to the target objects with the same identification number in different image frames may be saved. When feature-matching the first re-identification feature of the first target object with the second re-identification feature of the second target object included in the set region to obtain at least one feature similarity, it may be achieved in the following manner: determine the S re-identification features with the top confidence levels among the third re-identification feature and the first re-identification feature; for each second target object in the set region, determine the feature similarity according to the S re-identification features and the second re-identification feature. In some scenarios, the feature similarity satisfies the conditions shown in the following formula:

[0099]

[0100] where dis reid represents the feature similarity, represents the feature vector of the kth re-identification feature among the S re-identification features, and n R represents the feature vector of the second re-identification feature.

[0101] Based on the above solution, compared with the traditional cosine similarity calculation formula, the final feature similarity is determined by the mean of the similarities between multiple re-identification features with relatively high confidence levels corresponding to the first target object and the second re-identification feature, which can reduce the contingency and risk of a single feature similarity result and improve the robustness of the algorithm.

[0102] In some embodiments, the second target information includes the position of the center point of the second target. The first target information of the first target object is matched with the second target information of the second target object included in the set area to obtain at least one position distance, which can be determined by the following method: for each second target object within the set area, the position distance is determined by the predicted center point position of the first target object, the center point position of the second target, and the width of the first target object. In some scenarios, the position distance satisfies the conditions described by the following formula:

[0103]

[0104] where dis g represents the position distance, x represents the predicted center point position, x′ represents the center point position of the second target, and σ represents the width of the first target object.

[0105] Based on the above solution, compared with the general Euclidean distance formula, the distance between two target points in this application does not increase or decrease infinitely with the change of position, but is strictly controlled within [0,1], which effectively avoids the large difference in distance similarity calculated when the two targets are far apart. At the same time, it better fuses with the feature similarity distance than the simple Euclidean distance. In addition, the formula has distance weighting. Within a certain range, the distance similarity changes violently with the increase of the distance between the two targets. When it exceeds the set range, the distance similarity becomes less obvious with the increase of the distance and gradually tends to be stable. This non-linear distance similarity calculation method better conforms to the law of actual target matching.

[0106] 304. Determine the tracking result of each first target object according to at least one feature similarity and at least one position distance of each first target object.

[0107] In some embodiments, for each first target object, each feature similarity and its corresponding position distance can be weighted to obtain the matching weight of each second target object. Further, the second target object with the smallest matching weight can be used as the tracking result of the first target object. As an example, for each first target object, when there are three second target objects included in the set area, the feature similarity and position distance of each second target object can be weighted to obtain the matching weights corresponding to the three second target objects respectively. Among them, the matching weight satisfies the conditions described by the following formula:

[0108]

[0109] Among them, dis represents the matching weight of the second target object. Represents the weight of the position distance, dis g Represents the position distance, dis reid Represents the feature similarity.

[0110] In some embodiments, the greater the matching weight, the smaller the matching degree between the first target object and the second target object, and the smaller the matching weight, the greater the matching degree between the first target object and the second target object. The matching weight of each second target object is passed through the Hungarian algorithm to determine the tracking result of the first target object.

[0111] In some embodiments, the first target information further includes a first identification number, and the second target information further includes a second identification number and a second target status. After determining the tracking result of the first target object, the second identification number can be set to the first identification number, and the second target status can be set to the updated status.

[0112] In some embodiments, for each second target object, when the second identification number of the second target object is blank, the second identification number of the second target object can be set to an identification number different from the identification numbers of other target objects, and the second target status of the second target object can be set to the creation status. Among them, other target objects include the first target object in the previous image frame and the second target object in the current image frame.

[0113] In some embodiments, after determining that the second target object is the tracking result of the first target object, the target information corresponding to the second identification number can be updated to the second target information.

[0114] In some embodiments, the first target information further includes a first target status. If the position distance between the first target object and each second target object included in the set area is greater than the set threshold, the first target status is set to the lost status. In some scenarios, when the lost status lasts for a set duration, the target status of the first target object can be set to the deleted status. After that, the target object in the deleted status can no longer be managed and concerned about. As an example, when the lost status lasts for 2 seconds, the target status of the first target object is set to the deleted status. The set duration can also be set to other durations, and the present application does not make specific limitations on this.

[0115] As an example, the previous image frame can be represented as I t , and the current image frame can be represented as I t+1 , and I t and I t+1Separate inputs are fed into the backbone network for feature extraction to obtain the corresponding feature maps of I t and the corresponding feature maps of I t+1 corresponding feature maps. Further, the feature maps can be fused and the fused feature maps are input into the Re-Motion Head network for feature extraction to obtain re-identification features, and motion features are extracted to obtain the re-identification features of the target object and the target offset. The feature maps are input into the Detection Head network to obtain the target information of the target object. Further, the output of the Re-Motion Head network and the output of the Detection Head network can be used as the target tracking module Re-MotionTracker for target tracking to obtain the target tracking result, as Figure 4 shown.

[0116] Next, the training method of the target tracking network involved in this application will be described. The method flow is as Figure 5 shown, specifically as follows:

[0117] 501. Obtain the training set.

[0118] In some embodiments, road surveillance videos at different shooting locations, times, heights, and angles can be collected, and then each frame of the image is annotated. The annotation content includes the target type, the target identification number ID, and the bounding rectangle. Among them, the target types include 6 categories: vehicles, pedestrians, non-motor vehicles, construction signs, road cones, and roadblocks. The same target has the same target identification number. In some scenarios, the data set can be cleaned. For targets that are too small or areas where the target is severely occluded, they are marked as ignored areas, and the bounding rectangles in the annotation that are too large or too small are adjusted to ensure accurate annotation. In some scenarios, data augmentation can be performed by changing the brightness, hue, adding noise, rotation, mirroring, etc.

[0119] In some embodiments, the training set includes multiple samples, each sample includes at least two image frames, and each image frame includes the target type, the target identification number ID, and the bounding rectangle.

[0120] 502. Construct the target tracking network, and input the first sample in the training set into the target tracking network to output the predicted target type in the first sample, the predicted width and height of the target object, the predicted target center point position, the predicted re-identification result, and the predicted target offset through the target tracking network.

[0121] In some embodiments, the target tracking network can be based on CenterNet, with a re-identification branch and a motion prediction branch added to construct a learnable target tracking network that unifies object detection, re-identification, and tracking.

[0122] 503. Adjust the network parameters of the target tracking network according to the loss function.

[0123] In some embodiments, the loss function of the network includes the center heatmap classification loss function L headmap , the width and height regression loss function L wh , the target center point position loss function L offset , the re-identification loss function L id and the motion model loss function L motion . The total loss function of the target tracking network is the sum of each loss function.

[0124] In some embodiments, the center heatmap loss function is determined by the predicted target type output by the target tracking network and the target type of the target in the first sample. The width and height regression loss function is determined by the predicted width and height of the target object and the width and height of the target in the first sample. The target center point position loss function is determined by the predicted target center point position and the target center point position of the target in the first sample. In some scenarios, the center heatmap loss function, the width and height regression loss, and the target center point position loss function are the loss functions of the target detection branch, and the loss functions of centerNet can be reused.

[0125] In some embodiments, before training the re-identification network, the identification number of the target object can be encoded. The re-identification task can be regarded as a multi-classification task, so the softmax loss function can be used. In some scenarios, the re-identification loss function satisfies the conditions shown in the following formula:

[0126]

[0127] where L id represents the re-identification loss value, y k represents the encoding of the identification number of the k-th target object, and f(z k ) represents the probability of predicting the identification number of the k-th object. In some embodiments, the motion model loss function satisfies the conditions shown in the following formula:

[0128]

[0129] where L motion represents the motion loss value, represents the predicted target offset, is the offset of the target center point of the target object from the image It frame to the image It+i frame, which can be determined by the target boxes of the target object in the image It frame and the image It+i frame respectively.

[0130] In some embodiments, the network parameters can be updated through a network loss function such that each loss function is lower than the set threshold corresponding to each function respectively, and the total loss function is less than the set value, and the network parameters are saved.

[0131] Based on the same technical concept, an embodiment of the present application provides a target tracking device 600, as Figure 6 shown. The device 600 can execute any step in the above target tracking method. To avoid repetition, it will not be elaborated here. The device 600 includes a first determination module 601 and a second determination module 602.

[0132] The first determination module 601 is configured to perform feature extraction on the first feature map of the previous image frame and the second feature map of the current image frame respectively, to obtain the first re-identification features of each first target object in the previous image frame and the second re-identification features of each second target object in the current image frame, and to fuse the first feature map and the second feature map, and perform motion prediction on the fused feature map, to obtain the predicted target offset of each first target object in the current image frame;

[0133] The second determination module 602 is configured to, for each first target object: determine the set region of the first target object in the current image frame according to the first target information of the first target object and the predicted target offset of the first target object; perform feature matching between the first re-identification features of the first target object and the second re-identification features of the second target objects included in the set region, to obtain at least one feature similarity, and perform matching between the first target information of the first target object and the second target information of the second target objects included in the set region, to obtain at least one position distance;

[0134] Determine the tracking result of each first target object according to at least one feature similarity and at least one position distance of each first target object.

[0135] In some embodiments, the first feature map and the second feature map are determined by respectively performing feature extraction on the previous image frame and the current image frame through a first network in the target tracking network, the first target information and the second target information are determined by respectively performing feature extraction on the first feature map and the second feature map through a detection head network in the target tracking network, the target tracking network further includes a re-identification network and a motion network, and the first determination module 601, when performing feature extraction on the first feature map of the previous image frame and the second feature map of the current image frame respectively to obtain the first re-identification features of each first target object in the previous image frame and the second re-identification features of each second target object in the current image frame, specifically is configured to:

[0136] Input the previous image frame and the current image frame into the re-identification network respectively for feature extraction, so that the re-identification network outputs the first re-identification features of each first target object in the previous image frame and the second re-identification features of each second target object in the current image frame;

[0137] When the first determination module 601 fuses the first feature map and the second feature map and performs motion prediction on the fused feature map to obtain the predicted target offset of each first target object in the current image frame, it specifically is used for:

[0138] Add the first feature map and the second feature map bit by bit to obtain the fused feature map, and input the fused feature map into the motion network for motion prediction, so that the motion network outputs the predicted target offset of each first target object in the current image frame.

[0139] In some embodiments, the first target information includes the position of the center point of the first target, the width of the first target object, and the height of the first target object. When the second determination module 602 determines the set area of the first target object in the current image frame according to the first target information of the first target object and the predicted target offset of the first target object, it specifically is used for: determining the predicted target center point position of the first target object in the current image frame according to the position of the center point of the first target and the predicted target offset of the first target object; taking the area obtained by magnifying the width and height of the first target object by a set multiple with the predicted target center point position as the center as the set area.

[0140] In some embodiments, the second determination module 602 is further used for: determining the third re-identification features of the target object with the same identification number as the first target object in the image frame before the previous image frame before matching the first re-identification features of the first target object with the second re-identification features of the second target object included in the set area; when the second determination module 602 matches the first re-identification features of the first target object with the second re-identification features of the second target object included in the set area to obtain at least one feature similarity, it specifically is used for: determining the top S re-identification features with higher confidence in the third re-identification features and the first re-identification features; for each second target object in the set area, determining the feature similarity according to the S re-identification features and the second re-identification features.

[0141] In some embodiments, the feature similarity satisfies the conditions shown in the following formula:

[0142]

[0143] Among them, the dis reid represents the feature similarity, represents the feature vector of the k-th re-identification feature among the S re-identification features, and n R represents the feature vector of the second re-identification feature.

[0144] In some embodiments, the second target information includes the position of the second target center point. The second determination module 6002, when matching the first target information of the first target object with the second target information of the second target object included in the set area to obtain at least one position distance, is specifically configured to:

[0145] For each second target object within the set area, determine the position distance through the predicted target center point position of the first target object, the position of the second target center point, and the width of the first target object.

[0146] In some embodiments, the position distance satisfies the conditions described by the following formula:

[0147]

[0148] Among them, the dis g represents the position distance, x represents the predicted target center point position, x' represents the position of the second target center point, and σ represents the width of the first target object.

[0149] In some embodiments, the second processing module 602, when determining the tracking result of each first target object according to at least one feature similarity and at least one position distance of each first target object, is specifically configured to:

[0150] For each first target object, weight each feature similarity and its corresponding position distance to obtain the matching weight of each second target object;

[0151] Take the second target object with the smallest matching weight as the tracking result of the first target object.

[0152] In some embodiments, the first target information further includes a first identification number, the second target information further includes a second identification number and a second target status. After the second processing module 602 takes the second target object with the smallest matching weight as the tracking result of the first target object, it is further configured to:

[0153] Set the second identification number to the first identification number, and set the second target status to the updated status.

[0154] In some embodiments, the second processing module 602 is further configured to: for each second target object, when the second identification number of the second target object is blank, set the second identification number of the second target object to an identification number different from the identification numbers of other target objects, and set the second target state of the second target object to the creation state.

[0155] In some embodiments, the first target information further includes a first target state, and the second processing module 602 is further configured to: if the position distance between the first re-identification feature of the first target object and each second target object included in the set area is greater than a set threshold, set the first target state to the lost state; when the lost state lasts for a set duration, set the target state of the first target object to the deleted state.

[0156] Based on the same inventive concept, an embodiment of the present application provides an execution device 700, and this device 700 can implement any step of the target tracking method discussed above. Please refer to Figure 7 . This device includes a memory 701 and a processor 702.

[0157] The memory 701 is used to store program instructions;

[0158] The processor 702 is used to call the program instructions stored in the memory and execute the above-mentioned target tracking method according to the obtained program.

[0159] In the embodiments of the present application, the processor 702 may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware processor, or executed by a combination of hardware and software modules in the processor.

[0160] The memory 701, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 701 may include at least one type of storage medium. For example, it may include flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, and so on. The memory 701 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 701 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0161] Based on the same technical concept, the embodiments of the present application provide a computer-readable storage medium, including: computer program code, when the computer program code runs on a computer, causing the computer to execute the target tracking method as described above. Since the principle of solving problems by the above computer-readable storage medium is similar to that of the target tracking method, the implementation of the above computer-readable storage medium can refer to the implementation of the method, and the repeated parts will not be described again.

[0162] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0163] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0164] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0165] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0166] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A target tracking method, characterized in that, Including: Performing feature extraction on the first feature map of the previous image frame and the second feature map of the current image frame respectively to obtain the first re-identification features of each first target object in the previous image frame and the second re-identification features of each second target object in the current image frame, and fusing the first feature map and the second feature map, and performing motion prediction on the fused feature map to obtain the predicted target offset of each first target object in the current image frame; For each first target object: determining the set area of the first target object in the current image frame according to the first target information of the first target object and the predicted target offset of the first target object; performing feature matching on the first re-identification features of the first target object and the second re-identification features of the second target objects included in the set area to obtain at least one feature similarity, and matching the first target information of the first target object with the second target information of the second target objects included in the set area to obtain at least one position distance; Determining the tracking result of each first target object according to at least one feature similarity and at least one position distance of each first target object.

2. The method according to claim 1, characterized in that, The first feature map and the second feature map are determined by performing feature extraction on the previous image frame and the current image frame respectively through a first network in the target tracking network. The first target information and the second target information are determined by performing feature extraction on the first feature map and the second feature map respectively through a detection head network in the target tracking network. The target tracking network further includes a re-identification network and a motion network. The performing feature extraction on the first feature map of the previous image frame and the second feature map of the current image frame respectively to obtain the first re-identification features of each first target object in the previous image frame and the second re-identification features of each second target object in the current image frame includes: Inputting the previous image frame and the current image frame into the re-identification network for feature extraction respectively, so that the re-identification network outputs the first re-identification features of each first target object in the previous image frame and the second re-identification features of each second target object in the current image frame; The fusing the first feature map and the second feature map, and performing motion prediction on the fused feature map to obtain the predicted target offset of each first target object in the current image frame includes: Adding the first feature map and the second feature map bit by bit to obtain a fused feature map, and inputting the fused feature map into the motion network for motion prediction, so that the motion network outputs the predicted target offset of each first target object in the current image frame.

3. The method according to claim 1, characterized in that, The first target information includes the position of the first target center point, the width of the first target object, and the height of the first target object. The determining the set area of the first target object in the current image frame according to the first target information of the first target object and the predicted target offset of the first target object includes: Determine the predicted target center point position of the first target object in the current image frame according to the first target center point position and the predicted target offset of the first target object; Use the area obtained by magnifying the width and height of the first target object by a set multiple centered on the predicted target center point position as the set area.

4. The method according to claim 3, characterized in that, The first target information further includes a first identification number, and the method further includes: Before performing feature matching between the first re-identification feature of the first target object and the second re-identification feature of the second target object included in the set area, determine the third re-identification feature of the target object corresponding to the first identification number in the image frame before the previous image frame; Perform feature matching between the first re-identification feature of the first target object and the second re-identification feature of the second target object included in the set area to obtain at least one feature similarity, including: Determine the top S re-identification features with higher confidence levels among the third re-identification feature and the first re-identification feature; For each second target object in the set area, determine the feature similarity according to the S re-identification features and the second re-identification feature.

5. The method according to claim 4, wherein The feature similarity satisfies the conditions shown in the following formula: Among them, the dis reid represents the feature similarity, represents the feature vector of the k-th re-identification feature among the S re-identification features, and n R represents the feature vector of the second re-identification feature.

6. The method according to claim 3, wherein The second target information includes a second target center point position. Matching the first target information of the first target object with the second target information of the second target object included in the set area obtains at least one position distance, including: For each second target object within the set area, determine the position distance through the predicted target center point position of the first target object, the second target center point position, and the width of the first target object.

7. The method according to claim 6, wherein The position distance satisfies the conditions described in the following formula: Among them, the dis g represents the position distance, x represents the position of the center point of the predicted target, x' represents the position of the center point of the second target, and σ represents the width of the first target object.

8. The method according to claim 1, wherein Determining the tracking result of each first target object according to at least one feature similarity and at least one position distance of each first target object includes: For each first target object, weight each feature similarity and its corresponding position distance to obtain the matching weight of each second target object; Use the second target object with the smallest matching weight as the tracking result of the first target object.

9. The method according to claim 8, wherein The first target information further includes a first target state, and the second target information further includes a second identification number and a second target state. After using the second target object with the smallest matching weight as the tracking result of the first target object, the method further includes: Set the second identification number to the first identification number, and set the second target state to the updated state.

10. The method according to claim 9, wherein The method further includes: For each second target object, when the second identification number of the second target object is blank, set the second identification number of the second target object to an identification number different from the identification numbers of other target objects, and set the second target state of the second target object to the created state.

11. The method according to any one of claims 1-10, characterized in that, The first target information further includes a first target state, and the method further includes: If the position distance between the first target object and each second target object included in the set area is greater than the set threshold, then the first target state is set to the lost state; When the lost state lasts for a set duration, the target state of the first target object is set to the deleted state.

12. A target tracking device, characterized in that, Including: A first determination module, configured to respectively perform feature extraction on the first feature map of the previous image frame and the second feature map of the current image frame, to obtain the first re-identification features of each first target object in the previous image frame and the second re-identification features of each second target object in the current image frame, and to fuse the first feature map and the second feature map, and perform motion prediction on the fused feature map, to obtain the predicted target offset of each first target object in the current image frame; A second determination module, for each first target object: determine the set area of the first target object in the current image frame according to the first target information of the first target object and the predicted target offset of the first target object; perform feature matching between the first re-identification features of the first target object and the second re-identification features of the second target objects included in the set area, to obtain at least one feature similarity, and match the first target information of the first target object with the second target information of the second target objects included in the set area, to obtain at least one position distance; Determine the tracking result of each first target object according to at least one feature similarity and at least one position distance of each first target object.

13. An execution device, characterized in that, Including: A memory, configured to store program instructions; A processor, configured to obtain the program instructions stored in the memory, and execute the method according to any one of claims 1-11 according to the obtained program instructions.

Citation Information

Patent Citations

  • Single-target person tracking method integrating pedestrian re-identification and face detection

    CN112668483A

  • Kind of dr radiography lung contour extraction method based on fully convolutional network

    US20180130202A1