Object tracking method, device and equipment and computer readable storage medium
By using an object detection model and a target segmentation model with zero sample recognition ability, combined with mask correlation technology, the tracking error judgment problem caused by the change of the morphology and movement orientation of the target object in the prior art is solved, and the accuracy of object tracking and the adaptability of small models are improved.
Patent Information
- Application Number
- CN202510127842.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-27
AI Technical Summary
Existing target tracking technologies are prone to misjudgment or tracking failure when the target object changes greatly or the direction of the motion changes, and small models are difficult to adapt to the infinity of scenarios and the diversity of task requirements, resulting in low tracking accuracy.
A target detection model with zero sample recognition ability and language recognition ability is adopted, combined with the target segmentation model, and through the mask association of the target frame image, accurate tracking of the objects to be tracked and newly added objects is achieved.
It improves the accuracy of object tracking, solves the problem that the accuracy of traditional tracking models decreases over time, and enables small models to iterate, deploy and apply quickly.
Smart Images

Figure CN119991729A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an object tracking method, device, equipment and computer-readable storage medium. Background Art
[0002] At present, existing target tracking technical solutions are usually implemented using various small models, and often adopt a method of first detecting and then matching. Specifically, a target detection algorithm is first used to detect the target in the video, and then a motion state matching algorithm is used to determine whether it is the same target object, that is, an association between the trajectory and the detection frame is established based on the position, size and other information of the detection frame, and then the detection results of each frame are associated based on the Hungarian algorithm or target feature information, and the trajectory information is updated according to the association results; however, when the target object changes greatly in shape, such as when the posture or movement direction changes, the existing detection method is prone to misjudgment or tracking failure, and the existing small model has the problem of difficulty in adapting to the infinity of scenes and the diversity of task requirements, which ultimately leads to the problem of low tracking accuracy. Summary of the invention
[0003] In order to solve the above technical problems, the embodiments of the present application hope to provide an object tracking method, device, equipment and computer-readable storage medium, which can solve the problem of low tracking accuracy in related technologies.
[0004] The technical solution of this application is implemented as follows:
[0005] An object tracking method, the method comprising:
[0006] Determine a target frame image after a current frame image from a target video sequence; wherein the current frame image includes a plurality of objects to be tracked;
[0007] Based on the target detection model and the target segmentation model, a first mask corresponding to the target frame image is determined; wherein the target detection model is a model obtained by training an initial detection model using a target sample data set and a model fine-tuning technology, and has zero-sample recognition capability and language recognition capability; the mask is a pixel matrix corresponding to the target frame image;
[0008] Based on the target tracking model and the target frame image, determining a second mask corresponding to the target frame image;
[0009] Determine a target mask corresponding to the target frame image based on the first mask and the second mask;
[0010] Based on the target tracking model, the target mask and the target frame image, tracking results of each of the objects to be tracked and the target object in the target frame image are obtained; wherein the target object is a newly added object in the target frame image except for the multiple objects to be tracked.
[0011] In the above solution, determining the target frame image after the current frame image from the target video sequence includes:
[0012] Determining an image frame interval based on a motion parameter of the object to be tracked;
[0013] Based on the current frame image and the image frame number interval, the target frame image is acquired from the target video sequence.
[0014] In the above solution, determining the first mask corresponding to the target frame image based on the target detection model and the target segmentation model includes:
[0015] Detect the target frame image based on the target detection model to obtain a first detection frame of each of the objects to be tracked; wherein the first detection frame is obtained by screening multiple candidate detection frames of each of the objects to be tracked;
[0016] A mask to be processed corresponding to each first detection frame is determined based on the target segmentation model, the target frame image and the plurality of first detection frames, and the first mask is determined based on the plurality of masks to be processed.
[0017] In the above solution, determining the target mask corresponding to the target frame image based on the first mask and the second mask includes:
[0018] Performing a first operation on the first mask and the second mask to obtain an update mask, so as to update the second mask based on the update mask;
[0019] Obtaining a third mask of the target object based on the first mask and the second mask;
[0020] A second operation is performed on the update mask and the third mask to obtain the target mask.
[0021] In the above solution, the object tracking method further includes:
[0022] receiving a click instruction for the current frame image, and determining a target tracking object from the plurality of objects to be tracked based on the click instruction;
[0023] Acquire a second detection frame corresponding to the target tracking object, and determine a fourth mask corresponding to the target tracking object based on the target segmentation model, the second detection frame and the current frame image;
[0024] A tracking result of the target tracking object is determined based on the target tracking model, the fourth mask, the current frame image, and a plurality of consecutive frame images after the current frame image in the target video sequence.
[0025] In the above scheme, the step of determining from the target video sequence a target frame image after the current frame image and before the target frame image further includes:
[0026] Based on each initial sample image in the initial sample data set and the content information of each of the initial sample images, a target sample data set is constructed; wherein the target sample data set includes target images of different scene complexities;
[0027] Based on the target sample data set and model fine-tuning technology, the initial detection model is trained to obtain the target detection model.
[0028] In the above solution, the target sample data set is constructed based on each initial sample image in the initial sample data set and the content information of each initial sample image, including:
[0029] Based on each of the content information, determine a first target label and a second target label corresponding to each of the initial sample images; wherein the first target label represents description information of scene complexity corresponding to the initial sample image; and the second target label represents quantity information of objects in the sample initial image;
[0030] The target sample data set is constructed based on a plurality of the initial sample images, a plurality of the first target labels, and a plurality of the second target labels.
[0031] In the above solution, determining the first target label and the second target label corresponding to each of the initial sample images based on each of the content information includes:
[0032] Constructing a first prompt word, and determining each first label based on the first prompt word, each of the initial sample images and the first network model; wherein the first prompt word is used to obtain environmental description information of the image;
[0033] Constructing a second prompt word, and determining each second label and each second target label based on the second prompt word, each of the initial sample images and the second network model; wherein the second prompt word is used to obtain the number of each type of sample objects in the image;
[0034] Each first target label is determined based on each first label and each second label.
[0035] An object tracking device, the device comprising:
[0036] An acquisition unit, used to determine a target frame image after a current frame image from a target video sequence; wherein the current frame image includes a plurality of objects to be tracked;
[0037] A detection unit, configured to determine a first mask corresponding to the target frame image based on a target detection model and a target segmentation model; wherein the target detection model is a model obtained by training an initial detection model using a target sample data set and a model fine-tuning technology, and has zero-sample recognition capability and language recognition capability; and the mask is a pixel matrix corresponding to the target frame image;
[0038] A segmentation unit, configured to determine a second mask corresponding to the target frame image based on a target tracking model and the target frame image;
[0039] A processing unit, configured to determine a target mask corresponding to the target frame image based on the first mask and the second mask;
[0040] A tracking unit is used to obtain tracking results of each of the objects to be tracked and the target object in the target frame image based on the target tracking model, the target mask and the target frame image; wherein the target object is a newly added object in the target frame image except for the multiple objects to be tracked.
[0041] An object tracking device, the device comprising: a processor, a memory, and a communication bus;
[0042] The communication bus is used to realize the communication connection between the processor and the memory;
[0043] The processor is used to execute the object tracking program stored in the memory to implement the steps of the above object tracking method.
[0044] A computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the object tracking method as described above.
[0045] The object tracking method, apparatus, device and computer-readable storage medium provided in the embodiments of the present application first determine a target frame image after a current frame image from a target video sequence, and the current frame image includes a plurality of objects to be tracked, and then determine a first mask corresponding to the target frame image based on a target detection model and a target segmentation model, and the target detection model is a model obtained by training an initial detection model using a target sample data set and a model fine-tuning technology, and has zero-sample recognition capability and language recognition capability, and the mask is a pixel matrix corresponding to the target frame image, and then determine a second mask corresponding to the target frame image based on the target tracking model and the target frame image, and then determine a target mask corresponding to the target frame image based on the first mask and the second mask, and then obtain each object in the target frame image based on the target tracking model, the target mask and the target frame image. The tracking results of the newly added objects in the to-be-tracked object and the target frame image other than the multiple to-be-tracked objects. In this way, by adopting a target detection model with zero-sample recognition and language recognition capabilities, it can not only solve the problem of the traditional tracking model's accuracy decreasing over time, but also have the zero-sample / few-sample generalization capabilities that are difficult to achieve with traditional tracking algorithms. Then, combined with the segmentation results of a large model such as a target segmentation model, the problem of low tracking efficiency caused by the existing small model's difficulty in zero-sample generalization and poor domain adaptation is solved, so that the small model can be quickly iterated, deployed and applied, and the annotated frame mask (i.e., target mask) is obtained through the large model, so that the old target mask can be updated to a more accurate large model segmentation result, thereby solving the problem of the target tracking model (i.e., small model) decreasing in accuracy over time, thereby improving the accuracy of object tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A flowchart of an object tracking method provided in an embodiment of the present application;
[0047] Figure 2 A flowchart of another object tracking method provided in an embodiment of the present application;
[0048] Figure 3 A schematic diagram of obtaining a target tracking object in an object tracking method provided in an embodiment of the present application;
[0049] Figure 4 A schematic diagram of a visualization result display in an object tracking method provided in an embodiment of the present application;
[0050] Figure 5 A schematic diagram of the structure of an object tracking device provided in an embodiment of the present application;
[0051] Figure 6 A schematic diagram of the structure of an object tracking device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0053] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0054] It should be noted that the existing target tracking technical solutions are mostly implemented using various small models, and the method of first detection and then matching is often adopted. Specifically, the target detection algorithm is first used to detect the target in the video, and then the motion state matching algorithm is used to determine whether it is the same target object, such as the Simple Online and Realtime Tracking (SORT) method and the Deep Simple Online and Realtime Tracking (DeepSORT) method. The core idea is to track the target by associating the detection frame with the known trajectory, that is, to establish the association between the trajectory and the detection frame based on the position, size and other information of the detection frame, and then associate the detection results of each frame based on the Hungarian algorithm or target feature information, and update the trajectory information according to the association results. On the one hand, when the target object changes greatly in shape, such as when the posture or movement direction changes, the existing method is prone to misjudgment or tracking failure. In addition, this method requires the detection of the target in the scene in each frame, which has low computational efficiency; on the other hand, the existing small model has the problem of being difficult to adapt to the infinity of scenes and the diversity of task requirements. In the process of implementing conventional target tracking tasks, from detection models to tracking models, data fitting small models are used. During the tracking process, it is necessary to continuously provide the mask or detection box of the tracked target as an annotation, and then continuously track the target based on the existing annotations. Therefore, the accuracy of the annotation is particularly important. In order to achieve the accuracy requirements of the mask or detection box and the robustness of the model in new scenes, the detection and tracking small models need to be retrained using a large amount of scene detection box annotations or mask annotations that have been accumulated over a long period of time when applied to new scenes. Therefore, in real scenarios or project implementation, it is difficult for small models to quickly iterate, deploy, and apply and meet the capability requirements required by the project.
[0055] Based on this, the embodiment of the present application provides an object tracking method, which can be applied to an object tracking device. Figure 1 As shown, the method comprises the following steps:
[0056] Step 101: determine a target frame image after a current frame image from a target video sequence.
[0057] The current frame image includes a plurality of objects to be tracked.
[0058] In an embodiment of the present application, a target video sequence may refer to a video consisting of multiple frames of images in a certain arrangement order acquired under a target scene, and the target scene may refer to a traffic scene, such as intersections, parks, highways, ports and other road vehicle traffic scenes; the target video sequence includes multiple frames of ordered images; the current frame image may refer to the first frame image in the target video sequence, or it may be any frame image specified in the target video sequence; the object to be tracked may refer to an object (such as a vehicle) contained in the current frame image (such as a vehicle image); of course, the target frame image (recorded as the Fth frame image) also includes the object to be tracked included in the current frame image. In addition, the target frame image may also include newly added objects (i.e., objects not included in the current frame image); the target frame image is one or more frames after the current frame image, and specifically one or more target frame images can be determined from the target video sequence according to the motion parameters of the object to be tracked in the current frame image.
[0059] Step 102: Determine a first mask corresponding to the target frame image based on the target detection model and the target segmentation model.
[0060] Among them, the target detection model is obtained by training the initial detection model using the target sample data set and model fine-tuning technology, and has zero-sample recognition capability and language recognition capability; the mask is the pixel matrix corresponding to the target frame image.
[0061] In an embodiment of the present application, the first mask may refer to the first pixel matrix corresponding to the target frame image; the target detection model (i.e., the large model) may be a model for detecting the object to be tracked, and the target detection model is a model obtained by training the initial detection model using a target sample data set and model fine-tuning technology; the target segmentation model (i.e., the large model) may be a model for segmenting each object to be tracked; the first mask may refer to a mask obtained based on the large model and the target frame image; the first mask is essentially a pixel matrix, and the first mask may be obtained by masks corresponding to multiple objects to be tracked included in the target frame image. In a feasible implementation, the target detection model may be an open set object detection large model (Grounding DINO); the target segmentation model may refer to a segmentation everything large model (SegmentAnything Model, SAM).
[0062] Step 103: Determine a second mask corresponding to the target frame image based on the target tracking model and the target frame image.
[0063] In an embodiment of the present application, the second mask may refer to a second pixel matrix corresponding to the target frame image; the target tracking model may refer to a small tracking model, and compared with a large tracking model, the small tracking model has fewer parameters and lower computational requirements; the second mask is a prediction mask obtained by predicting the target frame image with the small tracking model, and the second mask is essentially also a pixel matrix; the target frame image can be input into the target tracking model for inference prediction to obtain a prediction mask corresponding to the target frame image.
[0064] Step 104: Determine a target mask corresponding to the target frame image based on the first mask and the second mask.
[0065] In an embodiment of the present application, the target mask may refer to the mask after the target frame image is updated; after obtaining the first mask and the second mask, the first mask and the second mask may be operated to obtain the target mask of the target frame image object, and the essence of the target mask is also a pixel matrix, and the target mask serves as a frame annotation of the target frame image.
[0066] Step 105: Based on the target tracking model, the target mask and the target frame image, obtain the tracking results of each object to be tracked and the target object in the target frame image.
[0067] The target object is a newly added object in the target frame image except for the multiple objects to be tracked.
[0068] In an embodiment of the present application, the target mask and the target frame image can be input into the target tracking model. When a new object (i.e., the target object) is not detected in the target frame image, the tracking result of the mask that does not contain the new object in the corresponding F+1 frame can be inferred; when a new object (i.e., the target object) is detected in the target frame image, the tracking result of the mask that contains the new object in the corresponding F+1 frame can be inferred.
[0069] The object tracking method provided in the embodiment of the present application, by adopting a target detection model with zero-sample recognition and language recognition capabilities, can not only solve the problem of the accuracy of traditional tracking models decreasing over time, but also has zero-sample / few-sample generalization capabilities that are difficult to achieve with traditional tracking algorithms. Then, combined with the segmentation results of a large model such as a target segmentation model, the problem of low tracking efficiency caused by the difficulty of zero-sample generalization and poor domain adaptation of existing small models is solved, so that the small model can be quickly iterated, deployed and applied, and by obtaining the annotated frame mask (i.e., the target mask) through the large model, the old target mask can be updated to a more accurate large model segmentation result, thereby solving the problem of the accuracy of the small target tracking model decreasing over time, thereby improving the accuracy of object tracking.
[0070] Based on the above embodiments, the present application provides another object tracking method. Figure 3 As shown, the method may include the following steps:
[0071] It should be noted that although the pre-trained large model has good zero-sample generalization ability, fine-tuning the large model with a small amount of scene data can make the large model have better reasoning accuracy in this specific scene. Therefore, it is necessary to build a fine-tuning dataset when there is a high reasoning accuracy requirement. This application proposes a data complexity classification method that can efficiently complete the construction of a scene fine-tuning dataset without manual screening, as shown in the following step 201:
[0072] Step 201: The object tracking device constructs a target sample data set based on each initial sample image and content information of each initial sample image.
[0073] Among them, the target sample data set includes target images of different scene complexities.
[0074] In an embodiment of the present application, the initial sample image may refer to an image corresponding to the application scenario of the tracking task, and in a traffic scenario, the initial sample image may be obtained through methods such as the Real Time Streaming Protocol (RTSP); after obtaining multiple initial sample images, the content information of each initial sample image may be analyzed to obtain a scene complexity label corresponding to each initial sample image, and then each initial sample image may be labeled to obtain a target sample data set.
[0075] It should be noted that step 201 can be implemented in the following ways:
[0076] Step 201A: The object tracking device determines a first target label and a second target label corresponding to each initial sample image based on content information of each initial sample image.
[0077] The first target label represents the description information of the scene complexity corresponding to the initial sample image; the second target label represents the quantity information of the objects in the sample initial image.
[0078] In an embodiment of the present application, the content of the initial sample image includes the environmental description information of the image and the number of objects in the image; for each initial sample image, different network models can be used to analyze the environmental description information of the initial sample image and the number of objects in the image, so as to obtain the first target label and the second target label corresponding to the initial sample image.
[0079] It should be noted that step 201A can be implemented in the following ways:
[0080] Step 201a1: The object tracking device constructs a first prompt word, and determines each first label based on the first prompt word, each initial sample image and the first network model.
[0081] The first prompt word is used to obtain the environmental description information of the image; the second prompt word is used to obtain the number of each type of sample objects in the image.
[0082] In the embodiment of the present application, setting a standardized prompt word (Prompt) can quickly allow the network model to return an answer that meets the intention. The first network model can refer to the large model of graphic scene understanding; the first prompt word in the return format of [weather, light intensity, traffic participant size, occlusion, time] can be set, and then the large model of graphic scene understanding can be used to read the initial sample image in sequence, and the first prompt word is also input into the large model of graphic scene understanding to obtain the first label in the specified format (i.e. [weather, light intensity, traffic participant size, occlusion, time]), and the data set with the first label is recorded as: Among them, frame i is the RGB pixel matrix of the i-th image, is frame i Corresponding to the first label, n is the total number of initial sample images.
[0083] Step 201a2: The object tracking device constructs a second prompt word, and determines each second label and each second target label based on the second prompt word, each initial sample image and the second network model.
[0084] The second prompt word is used to obtain the number of sample objects of each category in the image.
[0085] In an embodiment of the present application, the second network model is a model that has zero-sample recognition and language recognition capabilities and can perform zero-sample target detection, such as a GroundingDINO model; since in scene level classification, the number of objects in an image is a range value rather than an exact value, the detection accuracy requirement for the object is not high, therefore, GroundingDINO can be used to perform zero-sample target detection to determine the number of objects in the image.
[0086] Specifically, first, the corresponding second prompt word is set according to the actual task and traffic scene. For example, if the target data set needs to be used for the detection of vehicles and pedestrians in the intersection scene, the second prompt word can be set to [vehicle.person.]; then the second prompt word and the data set with the first label can be input into the second model for detection, and multiple detection frames of each sample object in the initial sample image, the confidence corresponding to each detection frame (score, and score∈(0,1)) and the category to which each pair of sample objects belongs can be obtained. Then, for each sample object, multiple detection frames are screened based on the confidence, and when score≥confidence threshold (such as 0.3), the category detection result of the sample object is determined to be a credible result. For frame i , the corresponding target object credible result is recorded as Y i ,and in, is the label of the credible result in the initial sample image, m is the number of credible results; finally, frame i The number of sample objects in the image is classified to obtain the number of sample object level labels Y in the final single frame image. i * , and Y i * ∈{Y i s ,Y i m ,Y i h}. When m∈[3,6), the magnitude label (i.e., the second target label) is Y i s ; When m∈[6,12), the quantity level label is Y i m ; When m∈[12,+∞), the quantity level label is Y i h .
[0087] Step 201a4: The object tracking device determines each first target tag based on each first tag and each second tag.
[0088] In this embodiment, for each initial sample image, the first label and the second label can be matched to obtain a first target label, and the first target label is recorded as For each initial sample image frame i , while extracting tags Weather, light intensity, traffic participant size, occlusion, time, and label Y i *The number of target labels is constructed, and the first target label list [weather, light intensity, traffic participant size, occlusion, time, number of objects] and the scene level label (L SIMPLE ,L MEDIUM ,L HARD ) to get the frame i The corresponding scene level label L, where the scene level label L corresponds to L SIMPLE ,L MEDIUM ,L HARD The scene level mapping relationship is as follows:
[0089] ["Sunny", "No glare", "Same size", "No obstruction", "Daytime", "Y i s ”], L SIMPLE
[0090] ["Sunny" / "Cloudy" / "Rainy","No strong light","Inconsistent size","Obstruction","Daytime" / "Nighttime","Y i m ”], L MEDIUM
[0091] ["Sunny" / "Cloudy" / "Rainy" / "Fog" / "Snowy","No strong light","Inconsistent size","Obstructed","Daytime" / "Nighttime","Y i h ”], L HARD
[0092] It should be noted that after the initial sample images are obtained, the obtained initial sample images can be screened according to the complexity of the scene environment (i.e., the scene), and the scene complexity can be judged to divide the scene into three levels: SIMPLE , L MEDIUM and L HARD , let the scene level label be L, and L∈{L SIMPLE ,L MEDIUM ,L HARD}, where the corresponding relationship of scene level labels is described as follows: L SIMPLE : The target is clear, the target is not blocked, there is no strong light, it is daytime, sunny, the target size is relatively simple, there are 3-6 target objects within the effective field of view of a single frame, and it is not rainy or foggy; L MEDIUM : The target is clear, the target is blocked, there is no strong light, all day, sunny / cloudy / rainy days, the target size is not single (far or near), 6-12 target objects within the effective field of view of a single frame, and it is not rainy or foggy; L MEDIUM : The target is not clear, the target is blocked, there is strong light, all day, all weather, the target size is not single (far or near), there are more than 12 target objects in the effective field of view of a single frame and all weather.
[0093] It should be noted that in order to ensure the rationality of data distribution, the distribution ratio of the number of scenes at each level is controlled within L SIMPLE ,L MEDIUM ,L HARD =2:5:3. Finally, the constructed initial data set is obtained, which is recorded as D={(frame1,L1),...,(frame i ,L i )),...,(frame n ,L n )}.
[0094] Step 201B: The object tracking device constructs a target sample data set based on multiple initial sample images, multiple first target labels, and multiple second target labels.
[0095] In an embodiment of the present application, the sample objects included in each initial sample image can also be manually labeled to obtain the detection frame and category of each sample object, and the detection frame and the category of the object are determined as the second target label; each initial sample image in the target sample data set has a first target label and a second target label, and each initial sample image can be marked with the first target label and the second target label corresponding to it, thereby constructing a target sample data set based on multiple labeled initial sample images.
[0096] It should be noted that the above-mentioned scene dataset (i.e., target sample dataset) construction method utilizes the scene recognition and open set detection capabilities of a multimodal large model to achieve scene classification work that requires manual judgment and screening when constructing traditional datasets, thereby reducing labor costs and realizing the automated construction of high-quality datasets.
[0097] Step 202: The object tracking device trains the initial detection model based on the target sample data set and the model fine-tuning technology to obtain a target detection model.
[0098] In the embodiment of the present application, the model fine-tuning technology may specifically refer to low-rank adaptation technology (Low Rank Adaptation, LoRA); in order to enhance the prediction accuracy of the initial detection model in a specific scenario, the initial detection model may be fine-tuned to improve the detection capability of the model in a specific scenario, so that the detection model can better learn the scene domain knowledge, and thus better adapt to the needs and characteristics of the traffic scene at the data distribution level; specifically, the target sample data set may be input into the initial detection model, and the LoRA technology may be used to fine-tune the initial detection model during the training process, and the core of the LoRA technology fine-tuning is to construct a low-rank weight matrix based on the pre-trained model weight matrix, and the specific steps are as follows:
[0099] Step 1, Weight matrix initialization: Denote the weight matrix pre-trained by the initial detection model as W, and denote the weight matrix of the target detection model after fine-tuning update as W + W0, where W0 is the weight matrix after fine-tuning update, and W0 << W. Using the low-rank decomposition method, decompose the weight matrix W0 into the product of matrix A and matrix B, that is: W0 = B × A, where both matrix A and matrix B are low-rank matrices, matrix A is initialized as a zero matrix, and matrix B is initialized as a Gaussian matrix.
[0100] Step 2, Weight update: First, set the LoRA optimization objective:
[0101]
[0102] Constrain and update the pre-trained weights by low-rank decomposition W + W0 = W + B × A, where B ∈ R d×r , A ∈ R r×k , and r << min(d, k). During training, W is not updated, and only matrices A and B are used to update the trainable parameters. Finally, the weight of the target detection model after fine-tuning is generated as W + W0.
[0103] Step 203, The object tracking device determines the image frame number interval based on the motion parameters of the object to be tracked.
[0104] Step 204, The object tracking device obtains the target frame image from the target video sequence based on the current frame image and the image frame number interval.
[0105] In the embodiments of the present application, for target tracking in the global scene, to achieve global tracking, it is necessary to continuously detect whether there are new objects in the scene. Therefore, it is necessary to first determine the target frame image according to the image frame number interval. Specifically, the image frame number interval (denoted as F) is mainly related to the motion parameters of the object to be tracked (such as moving speed). When the moving speed of the object to be tracked is relatively fast, F should take a smaller value to capture the motion deformation and new objects of the object to be tracked. When the moving speed of the object to be tracked is relatively slow, F can take a relatively large value.
[0106] In the embodiments of the present application, after determining the image frame number interval, starting from the current frame image, one frame of image is obtained every image frame number interval, which is the target frame image. Denote the RGB pixel matrix corresponding to the target frame image as frame rF , frame rF has a size of (h, w, c).
[0107] Step 205, The object tracking device detects the target frame image based on the target detection model to obtain the first detection box of each object to be tracked.
[0108] The first detection frame is obtained by screening multiple candidate detection frames of each object to be tracked.
[0109] In an embodiment of the present application, the target frame image can be input into the target detection model for detection to obtain multiple detection frames for each object to be tracked and the confidence of each detection frame, and then the multiple detection frames of each object to be tracked are screened based on the confidence to obtain the first detection frame of each object to be tracked. Specifically, the trained Grounding DINO model can be used to detect each object to be tracked in the target frame image and obtain its corresponding detection frame, that is, the above-mentioned trained Grounding DINO model is used to detect n objects to be tracked in the target frame image and generate multiple two-dimensional detection frames corresponding to each object to be tracked and the confidence of each detection frame, and set the confidence threshold α. In order to ensure the accuracy of subsequent segmentation, the detection frame with score>α is selected as the first detection frame of each object to be tracked. The first detection frame set obtained after screening is recorded as: B={[x1 min ,y1 min ,x1 max ,y1 max ],...,[xn min ,yn min ,xn max ,yn max ]}, where n is the number of first detection boxes, (xn min ,yn min ) is the pixel position coordinate of the upper left corner of the detection box, (xn max ,yn max ) is the pixel position coordinate of the lower right corner of the detection box.
[0110] Step 206: The object tracking device determines a mask to be processed corresponding to each first detection frame based on the target segmentation model, the target frame image and the multiple first detection frames, and determines a first mask corresponding to the target frame image based on the multiple masks to be processed.
[0111] In the embodiment of the present application, multiple first detection frames can be used as prompt words of the target segmentation model, and multiple first detection frames and the target frame image are input into the target segmentation model for segmentation, and multiple masks of each object to be tracked and the confidence of each mask are obtained, and then the mask with the highest confidence is selected as the mask to be processed for each object to be tracked, and finally the mask set to be processed of the object to be tracked in the target frame image is obtained as Masks rF ={mask1,...,mask j ,...mask P}, where Masks rFIt is a matrix of dimension (p,h,w), where h represents the image height, w represents the image width, and p is the number of masks, which is consistent with the number of objects to be tracked in the target frame image; mask j It represents the processing mask of a single object to be tracked. It is a Boolean matrix with dimension (h, w). When the corresponding element is the mask corresponding to a certain object to be tracked, the element value is True, otherwise it is False. Therefore, the processing mask of each object to be tracked can be numerically encoded to ensure the uniqueness of each mask. The unique ID is the corresponding non-repeating integer value, and the non-tracking pixel is set to 0. The encoded mask set is recorded as The first mask corresponding to the final target frame image is It should be noted that after obtaining the mask, the subsequent target tracking model can continuously predict the mask of the same object to be tracked in subsequent frames based on the given mask.
[0112] Step 207: The object tracking device determines a second mask corresponding to the target frame image based on the target tracking model and the target frame image.
[0113] In the embodiment of the present application, the second mask is a mask predicted by the target tracking model; the output result of the target tracking model (ie, the second mask) is Pred rF , then Pred rF =T(frame rF ), where T represents the encoding and decoding structure of the target tracking model, Pred rF is the inference result of the target tracking model, and its dimension size is (h,w). It should be noted that when no new objects are detected, for the same object to be tracked, Pred rF The ID in The ID in Pred rF The values in the pixel matrix are represented as integers.
[0114] Step 208: The object tracking device performs a first operation on the first mask and the second mask to obtain an updated mask, so as to update the second mask based on the updated mask.
[0115] In the embodiment of the present application, the first operation may refer to a matrix intersection operation; a matrix intersection operation may be performed on the first mask and the second mask to obtain an update mask, that is, Since the segmentation ability of the target segmentation model is more accurate than the predicted mask (i.e., the second mask), the second mask of the target tracking model can be updated to the calculated update mask refMask rF .
[0116] Step 209: The object tracking device obtains a third mask of the target object based on the first mask and the second mask.
[0117] In the embodiment of the present application, the second mask may be firstly negated, and then a matrix intersection operation may be performed on the obtained operation result and the first mask to obtain a mask of the newly added object (ie, the third mask), that is,
[0118] Step 210: The object tracking device performs a second operation on the update mask and the third mask to obtain a target mask.
[0119] In the embodiment of the present application, the second operation may refer to a matrix union operation; the target mask is the frame annotation of the target frame image; after obtaining the update mask and the third mask, the update mask and the third mask may be subjected to a matrix union operation (ie, merged) to obtain the target mask, i.e., newMask rF =refMask rF ∪addMask rF .
[0120] Step 211: The object tracking device obtains tracking results of each to-be-tracked object and the target object in the target frame image based on the target tracking model, the target mask and the target frame image.
[0121] The target object is a newly added object in the target frame image except for the multiple objects to be tracked.
[0122] In the embodiment of the present application, the annotation frame mask newMask rF The target frame image and the target frame image are used as the input to the target tracking model T(·), and the tracking result of the corresponding F+1 frame containing the target object (or not containing the target object) mask can be inferred, that is, the tracking result of each object to be tracked and the target object.
[0123] It should be noted that for target tracking in the global scene, first use RTSP streaming tools and the like to collect data for the application scenario; then use the capabilities of the multimodal graphic scene understanding large model and the target detection large model to achieve image scene complexity matching, and use this as a basis for label matching to complete the construction of the scene data set; then, use the constructed scene data set (i.e., target sample data set) to fine-tune the initial detection large model to improve the annotation accuracy of the tracking model; finally, generate a mask of the tracked target by calling the large model (i.e., target segmentation model), and use the mask as the frame annotation of the tracking small model (i.e., target tracking model), that is, based on the characteristics of the target object's movement speed in the scene, call the target detection large model and the segmentation large model every F frames to obtain the annotated frame mask; update the mask annotation of the old target object to the segmentation result of the large model, and cache some inference results of the small model; compare the results of the large and small models to obtain the mask annotation of the new target object, and achieve tracking of targets in the entire scene.
[0124] It should be noted that the above embodiment may further include the following steps:
[0125] Step 212: The object tracking device receives a click instruction for the current frame image, and determines a target tracking object from a plurality of objects to be tracked based on the click instruction.
[0126] Step 213: The object tracking device obtains a second detection frame corresponding to the target tracking object, and determines a fourth mask corresponding to the target tracking object based on the target segmentation model, the second detection frame and the current frame image.
[0127] In the embodiment of the present application, for tracking of a specified target object, a target tracking object may be determined from multiple objects to be tracked in an interactive manner, specifically, Figure 3 As shown, the mouse can be used to frame the target object to be tracked (i.e., the target tracking object) in the current frame image with a rectangular frame, and the target tracking object is framed in the rectangular frame (i.e., the second detection frame). The pixel coordinates of the second detection frame of the target tracking object are recorded as R = {[x1 left ,y1 left ,x1 right ,y1 right ]}, where (x1 left ,y1 left ) is the coordinate of the upper left corner pixel of the second detection box, (x1 right ,y1 right ) is the pixel coordinate of the lower right corner of the second detection frame. Afterwards, the second detection frame and the current frame image are input into the target segmentation model to obtain multiple masks corresponding to the target tracking object and the confidence corresponding to each mask, and then the mask with the highest confidence is determined as the fourth mask corresponding to the target tracking object.
[0128] Step 214: The object tracking device determines a tracking result of the target tracking object based on the target tracking model, the fourth mask, the current frame image, and a plurality of consecutive frame images after the current frame image in the target video sequence.
[0129] In an embodiment of the present application, the fourth mask, the current frame image, and the continuous multi-frame images can be input into the target tracking model, and the target tracking model is used to infer the mask of the target tracking object in the subsequent frames frame by frame. The mask inference result is overlaid on the target tracking object of the corresponding frame image through post-processing to obtain the visualization result output by the target tracking model, such as Figure 4 The overlay mask shown is the visualization result of tracking the specified vehicle (ie, the target tracking object).
[0130] It should be noted that for tracking of specified targets, the pixel coordinates of the specified tracking target or the selected detection box coordinates are obtained by interacting with the front-end interface through the mouse. For global scene target tracking, the fine-tuned target detection large model is used to obtain the target detection box coordinates of the entire scene, and the above coordinates are used as the input of the segmentation model to generate the tracking annotations corresponding to the target.
[0131] It should be noted that this application 1. implements a target tracking method by combining large and small models: this method can not only solve the problem of the accuracy of traditional tracking models decreasing over time, but also has the zero-sample / few-sample generalization capabilities that are difficult to achieve with traditional tracking algorithms. Combined with the detection and segmentation results of the large model, the problem of low tracking efficiency caused by the difficulty of zero-sample generalization and poor domain adaptation of the existing small model is solved, so that the small model can be quickly iterated, deployed, and applied and meet the capability requirements required by the project. The annotated frame mask is obtained through the large model, and the old target mask is updated to a more accurate large model segmentation result, which solves the problem of the accuracy of the traditional tracking small model decreasing over time. 2. A scene data set construction method is proposed, that is, the scene recognition and detection of large models with the open set detection capabilities of the multimodal large model are utilized to realize the scene classification work that requires manual judgment and screening to complete when constructing the traditional data set, thereby reducing labor costs and realizing the automatic construction of high-quality data sets. 3. A target tracking method based on a large traffic perception model is proposed, that is, a unified architecture combining large and small models to realize tracking tasks is provided, and a complete solution is constructed from fine-tuning the data set construction to the large and small models to coordinate the completion of designated targets and global scene target tracking. This solution can be quickly reused between different projects, and to a certain extent solves the problem that the existing small semi-supervised tracking model is difficult to adapt to the infinity of scenes and the diversity of task requirements. 4. A target tracking interactive method is provided, that is, a variety of methods for obtaining annotated frames of the target to be tracked are proposed, which can not only interactively obtain the tracking results of the selected target, but also complete the acquisition of object annotated frames and the generation of tracking results in the global scene. Compared with the traditional tracking model annotation frame acquisition method that relies on manual annotation, it can not only automatically obtain annotated frames, but also realize flexible adaptation of application forms.
[0132] It should be noted that, for the description of the same steps and the same contents in this embodiment as those in other embodiments, reference can be made to the description in other embodiments and will not be repeated here.
[0133] The object tracking method provided in the embodiment of the present application, by adopting a target detection model with zero-sample recognition and language recognition capabilities, can not only solve the problem of the accuracy of traditional tracking models decreasing over time, but also has zero-sample / few-sample generalization capabilities that are difficult to achieve with traditional tracking algorithms. Then, combined with the segmentation results of a large model such as a target segmentation model, the problem of low tracking efficiency caused by the difficulty of zero-sample generalization and poor domain adaptation of existing small models is solved, so that the small model can be quickly iterated, deployed and applied, and by obtaining the annotated frame mask (i.e., the target mask) through the large model, the old target mask can be updated to a more accurate large model segmentation result, thereby solving the problem of the accuracy of the small target tracking model decreasing over time, thereby improving the accuracy of object tracking.
[0134] Based on the above embodiments, the present invention provides an object tracking device, which can be applied to Figure 1 and Figure 2 In the object tracking method provided in the corresponding embodiment, refer to Figure 5 As shown, the object tracking device 3 may include: an acquisition unit 31, a detection unit 32, a segmentation unit 33, a processing unit 34 and a tracking unit 35, wherein:
[0135] An acquisition unit 31 is used to determine a target frame image after a current frame image from a target video sequence; wherein the current frame image includes a plurality of objects to be tracked;
[0136] The detection unit 32 is used to determine a first mask corresponding to the target frame image based on the target detection model and the target segmentation model; wherein the target detection model is a model obtained by training the initial detection model using the target sample data set and the model fine-tuning technology, and has zero-sample recognition capability and language recognition capability; the mask is a pixel matrix corresponding to the target frame image;
[0137] A segmentation unit 33, configured to determine a second mask corresponding to the target frame image based on the target tracking model and the target frame image;
[0138] The processing unit 34 is used to determine a target mask corresponding to the target frame image based on the first mask and the second mask;
[0139] The tracking unit 35 is used to obtain tracking results of each to-be-tracked object and the target object in the target frame image based on the target tracking model, the target mask and the target frame image; wherein the target object is a newly added object in the target frame image except for the multiple to-be-tracked objects.
[0140] In other embodiments of the present application, the acquisition unit 31 is specifically configured to perform the following steps:
[0141] Determining the image frame interval based on the motion parameters of the object to be tracked;
[0142] Based on the current frame image and the image frame number interval, a target frame image is obtained from a target video sequence.
[0143] In other embodiments of the present application, the segmentation unit 33 is specifically configured to perform the following steps:
[0144] Detect the target frame image based on the target detection model to obtain a first detection frame of each object to be tracked; wherein the first detection frame is obtained by screening multiple candidate detection frames of each object to be tracked;
[0145] A mask to be processed corresponding to each first detection frame is determined based on the target segmentation model, the target frame image and the plurality of first detection frames, and a first mask is determined based on the plurality of masks to be processed.
[0146] In other embodiments of the present application, the processing unit 34 is specifically configured to perform the following steps:
[0147] Performing a first operation on the first mask and the second mask to obtain an update mask, so as to update the second mask based on the update mask;
[0148] Obtaining a third mask of the target object based on the first mask and the second mask;
[0149] A second operation is performed on the update mask and the third mask to obtain a target mask.
[0150] In other embodiments of the present application, the tracking unit 35 is specifically configured to perform the following steps:
[0151] receiving a click instruction for the current frame image, and determining a target tracking object from a plurality of objects to be tracked based on the click instruction;
[0152] Obtain a second detection frame corresponding to the target tracking object, and determine a fourth mask corresponding to the target tracking object based on the target segmentation model, the second detection frame and the current frame image;
[0153] A tracking result of the target tracking object is determined based on the target tracking model, the fourth mask, the current frame image, and a plurality of continuous frame images after the current frame image in the target video sequence.
[0154] In other embodiments of the present application, the detection unit 32 is specifically configured to perform the following steps:
[0155] Based on each initial sample image in the initial sample data set and the content information of each initial sample image, a target sample data set is constructed; wherein the target sample data set includes target images of different scene complexities;
[0156] Based on the target sample data set and model fine-tuning technology, the initial detection model is trained to obtain the target detection model.
[0157] In other embodiments of the present application, the detection unit 32 is specifically configured to perform the following steps:
[0158] Based on each content information, determine a first target label and a second target label corresponding to each initial sample image; wherein the first target label represents description information of scene complexity corresponding to the initial sample image; and the second target label represents quantity information of objects in the sample initial image;
[0159] Based on a plurality of initial sample images, a plurality of first target labels, and a plurality of second target labels, a target sample data set is constructed.
[0160] In other embodiments of the present application, the detection unit 32 is specifically configured to perform the following steps:
[0161] Constructing a first prompt word, and determining each first label based on the first prompt word, each initial sample image and the first network model; wherein the first prompt word is used to obtain environmental description information of the image;
[0162] Constructing a second prompt word, and determining each second label and each second target label based on the second prompt word, each initial sample image and the second network model; wherein the second prompt word is used to obtain the number of each type of sample objects in the image;
[0163] Each first target label is determined based on each first label and each second label.
[0164] It should be noted that the specific description of the steps performed by each unit can be found in Figure 1 and Figure 2 The object tracking method provided in the corresponding embodiment will not be described again here.
[0165] The object tracking device provided in the embodiment of the present application, by adopting a target detection model with zero-sample recognition and language recognition capabilities, can not only solve the problem of the traditional tracking model's accuracy decreasing over time, but also has the zero-sample / few-sample generalization capability that is difficult to achieve with traditional tracking algorithms. Then, combined with the segmentation results of a large model such as a target segmentation model, the problem of low tracking efficiency caused by the difficulty of zero-sample generalization and poor domain adaptation of the existing small model is solved, so that the small model can be quickly iterated, deployed and applied, and by obtaining the annotated frame mask (i.e., the target mask) through the large model, the old target object mask can be updated to a more accurate large model segmentation result, thereby solving the problem of the accuracy of the small target tracking model decreasing over time, thereby improving the accuracy of object tracking.
[0166] Based on the above embodiments, the embodiments of the present application provide an object tracking device, which can be applied to Figure 1 and Figure 2 In the object tracking method provided in the corresponding embodiment, refer to Figure 6 As shown, the object tracking device 4 may include: a processor 41, a memory 42 and a communication bus 43, wherein:
[0167] The communication bus 43 is used to realize the communication connection between the processor 41 and the memory 42;
[0168] The processor 41 is used to execute the object tracking program in the memory 42 to implement the following steps:
[0169] Determine a target frame image after the current frame image from the target video sequence; wherein the current frame image includes a plurality of objects to be tracked;
[0170] Based on the target detection model and the target segmentation model, determine the first mask corresponding to the target frame image; wherein the target detection model is a model obtained by training the initial detection model using the target sample data set and the model fine-tuning technology, and has zero-sample recognition capability and language recognition capability; the mask is a pixel matrix corresponding to the target frame image;
[0171] Based on the target tracking model and the target frame image, determine a second mask corresponding to the target frame image;
[0172] Based on the first mask and the second mask, determining a target mask corresponding to the target frame image;
[0173] Based on the target tracking model, the target mask and the target frame image, the tracking results of each object to be tracked and the target object in the target frame image are obtained; wherein the target object is a newly added object in the target frame image except the multiple objects to be tracked.
[0174] In other embodiments of the present application, the processor 41 is used to execute the object tracking program in the memory 42 to determine the target frame image after the current frame image from the target video sequence, so as to implement the following steps:
[0175] Determining the image frame interval based on the motion parameters of the object to be tracked;
[0176] Based on the current frame image and the image frame number interval, a target frame image is obtained from a target video sequence.
[0177] In other embodiments of the present application, the processor 41 is used to execute the object tracking program in the memory 42 based on the target detection model and the target segmentation model to determine the first mask corresponding to the target frame image to implement the following steps:
[0178] Detect the target frame image based on the target detection model to obtain a first detection frame of each object to be tracked; wherein the first detection frame is obtained by screening multiple candidate detection frames of each object to be tracked;
[0179] A mask to be processed corresponding to each first detection frame is determined based on the target segmentation model, the target frame image and the plurality of first detection frames, and a first mask is determined based on the plurality of masks to be processed.
[0180] In other embodiments of the present application, the processor 41 is used to execute the object tracking program in the memory 42 to determine the target mask corresponding to the target frame image based on the first mask and the second mask, so as to implement the following steps:
[0181] Performing a first operation on the first mask and the second mask to obtain an update mask, so as to update the second mask based on the update mask;
[0182] Obtaining a third mask of the target object based on the first mask and the second mask;
[0183] A second operation is performed on the update mask and the third mask to obtain a target mask.
[0184] In other embodiments of the present application, the processor 41 is used to execute the object tracking method in the object tracking program in the memory 42 to implement the following steps:
[0185] receiving a click instruction for the current frame image, and determining a target tracking object from a plurality of objects to be tracked based on the click instruction;
[0186] Obtain a second detection frame corresponding to the target tracking object, and determine a fourth mask corresponding to the target tracking object based on the target segmentation model, the second detection frame and the current frame image;
[0187] Based on the target tracking model, the fourth mask, the current frame image and the continuous multiple frame images after the current frame image in the target video sequence, the tracking result of the target tracking object is determined; wherein the candidate frame image is the continuous multiple frame images after the current frame image in the target video sequence.
[0188] In other embodiments of the present application, the processor 41 is used to execute the object tracking method in the object tracking program in the memory 42 to implement the following steps:
[0189] Based on each initial sample image in the initial sample data set and the content information of each initial sample image, a target sample data set is constructed; wherein the target sample data set includes target images of different scene complexities;
[0190] Based on the target sample data set and model fine-tuning technology, the initial detection model is trained to obtain the target detection model.
[0191] In other embodiments of the present application, the processor 41 is used to execute the object tracking program in the memory 42 to construct a target sample data set based on each initial sample image and content information of each initial sample image in the initial sample data set, so as to implement the following steps:
[0192] Based on each content information, determine a first target label and a second target label corresponding to each initial sample image; wherein the first target label represents description information of scene complexity corresponding to the initial sample image; and the second target label represents quantity information of objects in the sample initial image;
[0193] Based on a plurality of initial sample images, a plurality of first target labels, and a plurality of second target labels, a target sample data set is constructed.
[0194] In other embodiments of the present application, the processor 41 is used to execute the object tracking program in the memory 42 to determine the first target label and the second target label corresponding to each initial sample image based on the content information, so as to implement the following steps:
[0195] Constructing a first prompt word, and determining each first label based on the first prompt word, each initial sample image and the first network model; wherein the first prompt word is used to obtain environmental description information of the image;
[0196] Constructing a second prompt word, and determining each second label and each second target label based on the second prompt word, each initial sample image and the second network model; wherein the second prompt word is used to obtain the number of each type of sample objects in the image;
[0197] Each first target label is determined based on each first label and each second label.
[0198] It should be noted that the specific description of the steps performed by the processor can be referred to Figure 1 and Figure 2 The object tracking method provided in the corresponding embodiment will not be described again here.
[0199] The object tracking device provided in the embodiment of the present application, by adopting a target detection model with zero-sample recognition and language recognition capabilities, can not only solve the problem of the traditional tracking model's accuracy decreasing over time, but also has the zero-sample / few-sample generalization capability that is difficult to achieve with traditional tracking algorithms. Then, combined with the segmentation results of a large model such as a target segmentation model, the problem of low tracking efficiency caused by the difficulty of zero-sample generalization and poor domain adaptation of the existing small model is solved, so that the small model can be quickly iterated, deployed and applied, and by obtaining the annotated frame mask (i.e., the target mask) through the large model, the old target object mask can be updated to a more accurate large model segmentation result, thereby solving the problem of the accuracy of the small target tracking model decreasing over time, thereby improving the accuracy of object tracking.
[0200] Based on the foregoing embodiments, the embodiments of the present application provide a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement Figure 1 and Figure 2 The corresponding embodiments provide steps in the object tracking method.
[0201] It should be noted that the above-mentioned computer-readable storage medium can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM) and other memories; it can also be various electronic devices including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0202] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0203] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0204] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course, by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0205] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products of the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0206] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0207] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0208] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. An object tracking method, characterized in that: The method comprises: Determine a target frame image after a current frame image from a target video sequence; wherein the current frame image includes a plurality of objects to be tracked; Based on the target detection model and the target segmentation model, a first mask corresponding to the target frame image is determined; wherein the target detection model is a model obtained by training an initial detection model using a target sample data set and a model fine-tuning technology, and has zero-sample recognition capability and language recognition capability; the mask is a pixel matrix corresponding to the target frame image; Based on the target tracking model and the target frame image, determining a second mask corresponding to the target frame image; Determine a target mask corresponding to the target frame image based on the first mask and the second mask; Based on the target tracking model, the target mask and the target frame image, tracking results of each of the objects to be tracked and the target object in the target frame image are obtained; wherein the target object is a newly added object in the target frame image except for the multiple objects to be tracked.
2. The method according to claim 1, characterized in that The step of determining a target frame image after the current frame image from the target video sequence comprises: Determining an image frame interval based on a motion parameter of the object to be tracked; Based on the current frame image and the image frame number interval, the target frame image is acquired from the target video sequence.
3. The method according to claim 1, characterized in that The determining, based on the target detection model and the target segmentation model, a first mask corresponding to the target frame image comprises: Detect the target frame image based on the target detection model to obtain a first detection frame of each of the objects to be tracked; wherein the first detection frame is obtained by screening multiple candidate detection frames of each of the objects to be tracked; A mask to be processed corresponding to each first detection frame is determined based on the target segmentation model, the target frame image and the plurality of first detection frames, and the first mask is determined based on the plurality of masks to be processed.
4. The method according to claim 1, characterized in that The determining, based on the first mask and the second mask, a target mask corresponding to the target frame image includes: Performing a first operation on the first mask and the second mask to obtain an update mask, so as to update the second mask based on the update mask; Obtaining a third mask of the target object based on the first mask and the second mask; A second operation is performed on the update mask and the third mask to obtain the target mask.
5. The method according to claim 1, characterized in that The method further comprises: receiving a click instruction for the current frame image, and determining a target tracking object from the plurality of objects to be tracked based on the click instruction; Acquire a second detection frame corresponding to the target tracking object, and determine a fourth mask corresponding to the target tracking object based on the target segmentation model, the second detection frame and the current frame image; A tracking result of the target tracking object is determined based on the target tracking model, the fourth mask, the current frame image, and a plurality of consecutive frame images after the current frame image in the target video sequence.
6. The method according to claim 1, characterized in that The step of determining from the target video sequence a target frame image after the current frame image and before the target frame image further includes: Based on each initial sample image in the initial sample data set and the content information of each of the initial sample images, a target sample data set is constructed; wherein the target sample data set includes target images of different scene complexities; Based on the target sample data set and model fine-tuning technology, the initial detection model is trained to obtain the target detection model.
7. The method according to claim 6, characterized in that The step of constructing a target sample data set based on each initial sample image in the initial sample data set and content information of each initial sample image comprises: Based on each of the content information, determine a first target label and a second target label corresponding to each of the initial sample images; wherein the first target label represents description information of scene complexity corresponding to the initial sample image; and the second target label represents quantity information of objects in the sample initial image; The target sample data set is constructed based on a plurality of the initial sample images, a plurality of the first target labels, and a plurality of the second target labels.
8. The method according to claim 7, characterized in that The determining, based on each of the content information, a first target label and a second target label corresponding to each of the initial sample images comprises: Constructing a first prompt word, and determining each first label based on the first prompt word, each of the initial sample images and the first network model; wherein the first prompt word is used to obtain environmental description information of the image; Constructing a second prompt word, and determining each second label and each second target label based on the second prompt word, each of the initial sample images and the second network model; wherein the second prompt word is used to obtain the number of each type of sample objects in the image; Each first target label is determined based on each first label and each second label.
9. An object tracking device, characterized in that: The device comprises: An acquisition unit, used to determine a target frame image after a current frame image from a target video sequence; wherein the current frame image includes a plurality of objects to be tracked; A detection unit, configured to determine a first mask corresponding to the target frame image based on a target detection model and a target segmentation model; wherein the target detection model is a model obtained by training an initial detection model using a target sample data set and a model fine-tuning technology, and has zero-sample recognition capability and language recognition capability; and the mask is a pixel matrix corresponding to the target frame image; A segmentation unit, configured to determine a second mask corresponding to the target frame image based on a target tracking model and the target frame image; A processing unit, configured to determine a target mask corresponding to the target frame image based on the first mask and the second mask; A tracking unit is used to obtain tracking results of each of the objects to be tracked and the target object in the target frame image based on the target tracking model, the target mask and the target frame image; wherein the target object is a newly added object in the target frame image except for the multiple objects to be tracked.
10. An object tracking device, characterized in that: The device comprises: a processor, a memory and a communication bus; The communication bus is used to realize the communication connection between the processor and the memory; The processor is used to execute the object tracking program stored in the memory to implement the steps of the object tracking method according to any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the object tracking method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Target tracking method and device, electronic equipment and storage medium
CN110610510A
Target tracking method and device, electronic equipment and readable storage medium
CN112101207A
Video mask self-encoding method and system
CN116363560A
Image processing method and device, equipment and computer storage medium
CN118799554A
Target tracking method, apparatus and system, and computer-readable storage medium
WO2021114702A1
Cited By
Classroom behavior detection method based on multi-modal large model
CN120183048A