Method, device and equipment for detecting object in video stream and storage medium
By generating a template feature set in the video stream and using a similarity matching update method, combined with a Siamese model to extract search features, the stability and accuracy issues of object tracking algorithms in video streams under complex scenarios are solved, and the accuracy of object location detection is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD
- Filing Date
- 2022-11-10
- Publication Date
- 2026-04-24
AI Technical Summary
In existing technologies, motion target tracking algorithms for objects in video streams suffer from insufficient stability and accuracy in complex interference scenarios, especially when the object itself undergoes significant changes, resulting in a decrease in precision and accuracy.
By determining the initial position information of objects in the first image of the video stream, a template feature set is generated, and the template features are updated based on a similarity matching update method. Combined with the Siamese model to extract search features, the position information of the objects is accurately determined.
It improves the accuracy and precision of object location detection in video streams, and enhances stability and tracking capabilities in complex scenes.
Smart Images

Figure CN116824422B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to, but is not limited to, the field of computer vision technology, and in particular to a method, apparatus, device, and storage medium for detecting objects in a video stream. Background Technology
[0002] Moving target tracking has been a popular research area in computer vision since its inception, and has been extensively studied by many scholars in recent years. With the continuous improvement of computer performance and the rapid development of artificial intelligence technology, higher requirements are placed on practical moving target tracking algorithms, such as maintaining the stability and accuracy of moving target tracking algorithms in complex and noisy scenarios. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure provide at least one method, apparatus, device, and storage medium for detecting objects in a video stream.
[0004] The technical solution of this disclosure embodiment is implemented as follows:
[0005] On one hand, embodiments of this disclosure provide a method for detecting objects in a video stream, comprising: determining initial position information of an object to be detected in a first image of the video stream; determining a template image corresponding to the first image based on the initial position information, and generating a template feature set based on template features of the template image; determining a search image corresponding to a second image of the video stream based on the initial position information; wherein the frame number of the first image is located before the frame number of the second image; determining the similarity between the search features of the search image and each feature in the template feature set; updating the template features using an update method matching the similarity to obtain updated template features; and determining second position information of the object to be detected in the second image based on the updated template features and the search features.
[0006] On the other hand, embodiments of this disclosure provide an object detection device in a video stream, comprising: a first determining module, configured to determine initial position information of an object to be detected in a first image of the video stream; a second determining module, configured to determine a template image corresponding to the first image based on the initial position information, and generate a template feature set based on template features of the template image; a third determining module, configured to determine a search image corresponding to a second image of the video stream based on the initial position information; wherein the frame number of the first image is located before the frame number of the second image; a fourth determining module, configured to determine the similarity between the search features of the search image and each feature in the template feature set; a first updating module, configured to update the template features using an update method matching the similarity, to obtain updated template features; and a fifth determining module, configured to determine second position information of the object to be detected in the second image based on the updated template features and the search features.
[0007] In another aspect, embodiments of this disclosure provide a computer device including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement some or all of the steps in the above-described method.
[0008] In another aspect, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above-described method.
[0009] In another aspect, embodiments of this disclosure provide a computer program including computer-readable code, which, when executed in a computer device, causes a processor in the computer device to perform some or all of the steps in the above-described method.
[0010] In another aspect, embodiments of this disclosure provide a computer program product, the computer program product including a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is read and executed by a computer, it implements some or all of the steps in the above method.
[0011] In related technologies, a template image corresponding to the first frame of a video stream is obtained, and the template features of the template image are used as the basis for determining the object position information in all subsequent video images. However, the template features of the template image are fixed, which leads to a decrease in the accuracy and precision of moving target tracking when the object itself changes significantly in all subsequent video images of the video stream.
[0012] In this embodiment, firstly, the initial position information of the object to be detected in the first image of the video stream is determined; based on the initial position information, a template image corresponding to the first image is determined, and a template feature set is generated based on the template features of the template image; based on the initial position information, a search image corresponding to the second image of the video stream is determined. This allows for accurate acquisition of the template features of the template image, the search features of the search image, and the template feature set used for confidence evaluation of the search features. Secondly, the similarity between the search features of the search image and each feature in the template feature set can be determined; the template features are updated using an update method matching the similarity, quickly and accurately obtaining the updated template features; finally, based on the updated template features and search features, the second position information of the object to be detected in the second image can be determined more accurately. Thus, by updating the template features of the template image, the accuracy and precision of object position detection in subsequent video images of the video stream are improved.
[0013] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure. Attached Figure Description
[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the specification, serve to illustrate the technical solutions of this disclosure.
[0015] Figure 1 A schematic diagram illustrating the implementation flow of the first method for detecting objects in a video stream provided in this embodiment of the disclosure;
[0016] Figure 2 A schematic diagram illustrating the implementation flow of the second method for detecting objects in a video stream provided in this embodiment of the disclosure;
[0017] Figure 3 A schematic diagram illustrating the implementation flow of the third method for detecting objects in a video stream provided in this embodiment of the disclosure;
[0018] Figure 4 A schematic diagram illustrating the implementation flow of the fourth method for detecting objects in a video stream provided in this embodiment of the disclosure;
[0019] Figure 5 A schematic diagram of the composition structure of a twin model provided in an embodiment of this disclosure;
[0020] Figure 6 A schematic diagram illustrating the implementation process of the fifth method for detecting objects in a video stream provided in this embodiment of the disclosure;
[0021] Figure 7 This is a schematic diagram of a sample pair in a sample set provided in an embodiment of the present disclosure;
[0022] Figure 8 This is a schematic diagram of the composition structure of a device for detecting objects in a video stream, provided in an embodiment of this disclosure.
[0023] Figure 9 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this disclosure. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this disclosure. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0025] In the following description, references to "some embodiments" describe a subset of all possible embodiments; however, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict. The terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein.
[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used herein is for descriptive purposes only and is not intended to limit this disclosure.
[0027] This disclosure provides a method for detecting objects in a video stream, which can be executed by a processor of a computer device. The computer device can refer to a server, laptop computer, tablet computer, desktop computer, smart TV, set-top box, mobile device (e.g., mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device), or any other device with object detection capabilities. Figure 1 This is a schematic diagram illustrating the implementation flow of a method for detecting objects in a video stream provided in an embodiment of this disclosure, as shown below. Figure 1 As shown, the method includes the following steps S101 to S106:
[0028] Step S101: Determine the initial position information of the object to be detected in the first image of the video stream.
[0029] Here, a video stream can refer to video data containing an object to be detected. A video stream can be composed of multiple consecutive frames containing images of the object to be detected, such as the first image in the video stream, the second image adjacent to the first image, etc. The object to be detected can refer to an object whose position can be detected, such as vehicles, animals, and pedestrians. The object to be detected in the video stream can be moving or stationary, such as a video stream containing the reciprocating motion of the object to be detected.
[0030] The initial position information of an object can refer to its location in the first image. This initial position information may include the center position information representing the object's displacement state, and the boundary position information representing the object's contour. For example, if the first frame of a video stream is defined as the first image, the initial position information of the object to be detected in the first image can be represented by a first bounding box. This first bounding box can be a rectangle with its center coordinates at (100, 50) and its length and width at (20, 20), etc. The number of objects to be detected in the first image is not limited. For example, if the first image includes a first vehicle and a second vehicle, the initial position information can include the position information of both vehicles. When the first image includes at least two objects to be detected, the first and second objects can belong to the same type or different types. For example, the first object can be a motor vehicle, and the second object can be a pedestrian.
[0031] During the implementation of step S101, the process may include: responding to a user's annotation operation on an object to be detected in the first image, reading the annotation information carried by the annotation operation; wherein the annotation information includes the position information and size information of the annotation box; and determining the position information and size information of the annotation box as the initial position information. The annotation operation may refer to the user's operation of annotating the first image using a preset annotation tool (such as Labelimg), for example, using the annotation tool to annotate vehicles on the first image to obtain an annotation box containing the vehicles. Step S101 may also include: receiving the initial position information uploaded by the user, thereby determining the initial position information of the object to be detected in the first image of the video stream.
[0032] Step S102: Determine the template image corresponding to the first image based on the initial position information, and generate a template feature set based on the template features of the template image.
[0033] Here, the template image can refer to a template used in the video stream to detect the position information of the object to be detected in images following the first image. For example, the position information of the object to be detected in the second image is determined at least based on the template image, and the position information of the object to be detected in the third image is also determined at least based on the template image, and so on. Template features can refer to features used to characterize the template image, such as performing feature extraction processing on the template image to obtain the template features of the template image.
[0034] A template feature set can be an object used to store one or more features. The features in the template feature set can be used to determine how the template features are updated in subsequent processes. The features in the template feature set and the number of features can be updated. For example, first, an empty template feature set is initialized, and template features are added to the template feature set, so the template feature set contains 1 feature (i.e., the template feature). Then, the template feature is updated for the first time using a preset update method, and the updated template feature is added to the template feature set, so the current template feature set contains 2 features. Second, the template feature is updated for the second time using the preset update method, and the updated template feature is added to the template feature set, so the current template feature set contains 3 features, and so on. Finally, the maximum number of features in the template feature set can be preset to 5. When the template feature is updated for the fifth time, a preset replacement method can be used to replace the updated template feature in the template feature set, and so on.
[0035] The size of the template image can be fixed, for example, the size of the template image is 127*127. Step S102 may include: obtaining the annotation region matching the initial position information from the first image; performing size transformation processing on the image represented by the annotation region to obtain a template image whose size meets the preset conditions. For example: the initial position information includes the center coordinates and the boundary length and width, the center coordinates are (100,100), and the boundary length and width are (90,90); (100,100) can be determined as the center position of the annotation region, and (90,90) can be determined as the length and width of the annotation region; the first image is cropped to obtain the image represented by the annotation region, and the image represented by the annotation region is enlarged by image fitting to obtain a template image with a size of 127*127.
[0036] Step S103: Determine the search image corresponding to the second image of the video stream based on the initial position information.
[0037] Here, the frame number can refer to the decoding order number of each image when decoding a video stream to obtain multiple frames. The frame number of the first image precedes the frame number of the second image. The frame numbers of the first and second images can be adjacent or not adjacent, etc. For example, the first image is the first frame of the video stream, and the second image is the second frame of the video stream, etc. The search image (also called the detection image) can be an image used to detect the position information of the object to be detected in the second image. The position information of the object to be detected in the second image can be determined from the search image. The size of the search image can be larger than the size of the template image. For example, the size of the search image is 255*255, and the size of the template image is 127*127. Step S103 may include: determining the annotation area for initial position information matching, enlarging the size of the annotation area to obtain the enlarged annotation area; performing size transformation processing on the image represented by the enlarged annotation area to obtain the search image whose size meets the preset conditions, etc.
[0038] Step S104: Determine the similarity between the search features of the search image and each feature in the template feature set.
[0039] Here, similarity can include the degree of similarity or distance between features. The similarity between features can be Euclidean distance, cosine similarity, cosine distance, cosine distance, Minkowski distance, Manhattan distance, Chebyshev distance, Mahalanobis distance, Pearson correlation coefficient, etc. Step S104 can include: extracting features from the search image to obtain search features of a preset dimension; since the features in the template feature set are variable, for the search features corresponding to the second image, only the cosine similarity between the search features corresponding to the second image and the template features corresponding to the first image needs to be determined.
[0040] Step S105: Update the template features using an update method that matches the similarity to obtain the updated template features.
[0041] Here, the correspondence between similarity and update method can be preset. For example, if the similarity is greater than the similarity threshold, the first update method is used to update the template features; if the similarity is less than or equal to the similarity threshold, the second update method is used to update the template features. The first update method can be to superimpose the template features with the preset reference features to obtain the updated template features; the second update method can be to subtract the template features from the preset reference features to obtain the updated template features, etc.
[0042] Step S106: Based on the updated template features and the search features, determine the second location information of the object to be detected in the second image.
[0043] Here, the updated template features and search features can be convolved to obtain convolutional features; the element with the largest value in the convolutional features can be determined, and the coordinates of the element with the largest value in the second image can be used as the center coordinates of the object to be detected in the second image; and the boundary length and width in the initial position information can be transformed according to a preset size transformation formula to obtain the boundary coordinates of the object to be detected in the second image. After implementing step S106, it may also include: based on the initial position information and the second position information, determining the motion trajectory of the object to be detected between the first image and the second image to achieve tracking of the object to be detected.
[0044] In related technologies, a template image corresponding to the first frame of a video stream is obtained, and the template features of the template image are used as the basis for determining the object position information in all subsequent video images. However, the template features of the template image are fixed, which leads to a decrease in the accuracy and precision of moving target tracking when the object itself changes significantly in all subsequent video images of the video stream.
[0045] In this embodiment, firstly, the initial position information of the object to be detected in the first image of the video stream is determined; based on the initial position information, a template image corresponding to the first image is determined, and a template feature set is generated based on the template features of the template image; based on the initial position information, a search image corresponding to the second image of the video stream is determined. This allows for accurate acquisition of the template features of the template image, the search features of the search image, and the template feature set used for confidence evaluation of the search features. Secondly, the similarity between the search features of the search image and each feature in the template feature set can be determined; the template features are updated using an update method matching the similarity, quickly and accurately obtaining the updated template features; finally, based on the updated template features and search features, the second position information of the object to be detected in the second image can be determined more accurately. Thus, by updating the template features of the template image, the accuracy and precision of object position detection in subsequent video images of the video stream are improved.
[0046] This disclosure provides a method for detecting objects in a video stream, such as... Figure 2 As shown, the method includes the following steps S201 to S206:
[0047] Steps S201 to S204 correspond to steps S101 to S104, respectively, and can be implemented with reference to the specific implementation of steps S101 to S104; step S206 corresponds to step S106, and can be implemented with reference to the specific implementation of step S106.
[0048] Step S205: If all the similarities are greater than or equal to the similarity threshold, update the template features based on the search features to obtain the updated template features.
[0049] Here, if there are 3 features in the current template feature set, the preset similarity threshold is 0.6, and the cosine similarity between the template feature and each feature in the template feature set is determined to be 0.8, 0.9, and 0.7 respectively, and all cosine similarities are greater than the similarity threshold, then the template feature and the search feature can be superimposed, subtracted, or convolved to obtain the updated template feature.
[0050] In this embodiment of the disclosure, by determining that the object to be detected in the second image is detected when all similarities are greater than or equal to the similarity threshold, the confidence level of accurately obtaining the second location information is high. The search features determined in the second image are more accurate, so the template features can be updated based on the search features to accurately obtain the updated template features.
[0051] In some embodiments, step S105 may include steps S211 to S212 as follows:
[0052] Step S211: If at least one of the similarities is less than the similarity threshold, obtain the feature with the highest similarity to the search feature from the template feature set.
[0053] Here, if there are 3 features in the current template feature set, the preset similarity threshold is 0.6, and the cosine similarity between the template feature and each feature in the template feature set is determined to be 0.8, 0.5, and 0.7 respectively, and there is one cosine similarity less than the similarity threshold; then the feature with the largest cosine similarity to the search feature can be obtained from the template feature set (that is, the first feature with a cosine similarity of 0.8). Since the search feature has the largest cosine similarity to the first feature in the template feature set, the first feature is obtained from the template feature set.
[0054] Step S212: Update the template features based on the features with the highest similarity to the search features to obtain the updated template features.
[0055] Here, if the cosine similarity between the search feature and the first feature in the template feature set is the largest, the template feature and the first feature in the template feature set can be superimposed, subtracted or convolved to obtain the updated template feature.
[0056] In this embodiment of the disclosure, if the confidence level of detecting the object to be detected in the second image is low when at least one similarity is greater than the similarity threshold, and the deviation of the search features determined by the second image is large, then the template features can be updated based on the features with the highest similarity to the search features in the template feature set, so as to accurately obtain the updated template features.
[0057] In some embodiments, step S205 above may include the following step S2051:
[0058] Step S2051: Perform a weighted summation on the search features and the template features to obtain the updated template features.
[0059] Here, a first weight (also known as the learning rate for template feature updates) and a second weight of the template features can be predetermined. Based on the first and second weights, the search features and template features are weighted and summed to obtain the updated template features. Alternatively, the template feature update method can also include weighted summing of the features with the highest similarity to the search features and the template features to obtain the updated template features.
[0060] In some embodiments, the template feature can be updated using the following formula:
[0061] φ t '(z)=λφ t (z)+(1-λ)φ t (x) (1);
[0062] In formula (1), φ t '(z) can represent the updated template features, z can represent the template image, λ can represent the weights corresponding to the template features, and φ t (z) can represent the template features before the update, (1-λ) can represent the weights corresponding to the search features, and φ t (x) can represent the search feature corresponding to the t-th image, or the feature in the template feature set that has the highest similarity to the search feature. x can represent the search image, etc.
[0063] In this embodiment of the disclosure, the updated template features can be accurately obtained by performing a weighted summation of the search features and template features.
[0064] This disclosure provides a method for detecting objects in a video stream, such as... Figure 3 As shown, the method includes the following steps S301 to S308:
[0065] Steps S301 to S306 correspond to the aforementioned steps S101 to S106 respectively. When implementing these steps, you can refer to the specific implementation methods of the aforementioned steps S101 to S106.
[0066] Step S307: If all the similarities are greater than or equal to the similarity threshold, add the search features to the template feature set to obtain the updated template feature set.
[0067] Here, if the current template feature set contains 3 features, the preset similarity threshold is 0.6, and the cosine similarities between the template feature and each feature in the template feature set are determined to be 0.8, 0.9, and 0.7 respectively (all cosine similarities are greater than the similarity threshold), then while updating the template feature set based on the template feature and the search feature, the search feature can also be added to the template feature set, resulting in an updated template feature set containing 4 features. If at least one preset similarity is less than or equal to the similarity threshold, then it is not necessary to add the search feature to the template feature set, and the template feature set is not updated.
[0068] Step S308: Based on the updated template feature set and the second location information, determine the third location information of the object to be detected in the third image of the video stream.
[0069] Here, the frame number of the second image precedes the frame number of the third image. For example, the second image is the next frame in the video stream, and the third image is a subsequent frame in the video stream adjacent to the next frame. Step S308 includes: determining the search image corresponding to the third image based on the second position information; determining the similarity between the search features corresponding to the search image in the third image and each feature in the updated template feature set; updating the updated template features again based on the update method matching the similarity; and then determining the third position information of the object to be detected in the third image based on the updated template features and the search features of the third image; simultaneously, determining whether to update the current template feature set based on the relationship between the similarity and the similarity threshold.
[0070] In this embodiment of the disclosure, by searching for the similarity between the feature and each feature in the template feature set, the template feature set can be updated accurately, and then the updated template features can be updated accurately based on the updated template feature set, which helps to obtain the third location information of the object to be detected in the third image more accurately.
[0071] In some embodiments, step S307 may include the following steps S3071 to S3072:
[0072] Step S3071: If the number of features in the template feature set is greater than the number threshold, the feature with the lowest similarity to the search feature is obtained from the template feature set.
[0073] Here, before implementing step S3071, the following steps are included: determining whether the number of features in the current template feature set is greater than a threshold. If the number of features is greater than the threshold, the feature with the lowest similarity to the search feature is obtained from the current template feature set. For example, if the current number of features is 5, the threshold is 4, and the cosine similarity between the search feature and each feature in the template feature set is 0.8, 0.6, 0.75, 0.7, and 0.9 respectively, then the second feature is determined from the template feature set. If the number of features is less than or equal to the threshold, the template feature is added to the current template feature set. For example, if the current number of features is 4 and the threshold is 4, the template feature is added to the current template feature set to increase the number of features in the template feature set.
[0074] Step S3072: Remove the feature with the lowest similarity to the search feature from the template feature set, and add the search feature to the template feature set to obtain the updated template feature set.
[0075] Here, if the cosine similarity between the second feature in the template feature set and the template feature is the smallest, then the second feature is removed from the template feature set, and the search feature is added to the template feature set so that the number of features in the template feature set does not exceed the number threshold.
[0076] In this embodiment of the disclosure, by removing the feature with the lowest similarity to the search feature from the template feature set, the number of features in the template feature set is limited, which can effectively reduce the amount of computation without reducing the detection accuracy of the object.
[0077] This disclosure provides a method for detecting objects in a video stream, such as... Figure 4 As shown, the method includes the following steps S401 to S407:
[0078] Step S401 corresponds to the aforementioned step S101, and can be implemented with reference to the specific implementation of the aforementioned step S101; Steps S403 to S406 correspond to the aforementioned steps S102 to S105 respectively, and can be implemented with reference to the specific implementation of the aforementioned steps S102 to S105.
[0079] Step S402: Using the twin network in the preset twin model, feature extraction is performed on the template image and the search image to obtain the template features and the search features.
[0080] Here, Siamese networks (SN) are also called conjoined networks. The conjoint aspect in a Siamese model is achieved through shared weights. A Siamese model takes two inputs (e.g., the first input and the second input) and feeds them into two separate branch networks (e.g., the first network and the second network). The similarity between the two inputs is evaluated through the final loss calculation. Because of weight sharing, Siamese models limit the difference between the first and second networks to a certain extent. Therefore, they are typically used to handle problems where the two inputs are not very different, such as comparing the similarity between two images, two sentences, or two words. For similarities with large input differences, such as the similarity between an image and its corresponding text description, or the similarity between an article title and an article paragraph, pseudo-Siamese networks can be used. Siamese models can include Fully-Convolutional Siamese Networks (SiamFC), Region-Proposal Siamese Networks (SiamRPN), etc.
[0081] Taking the trained SiamRPN model as an example, the Siam model can include a Siamese network and a region generation network. The Siamese network can include two branches: a first Siamese network for extracting template features from the template image, and a second Siamese network for extracting search features from the search image. Both the first and second Siamese networks can be Convolutional Neural Networks (CNNs). For example, inputting a template image of size 127*127*3 into the first Siamese network yields template features of dimension 6*6*128; inputting a search image of size 255*255*3 into the second Siamese network yields search features of dimension 22*22*128.
[0082] Step S407: Using the region generation network in the twin model, perform convolution processing on the updated template features and the search features to obtain the center coordinates and boundary coordinates of the object to be detected in the second image.
[0083] Here, the region generation network in the Siamese model can include two branches: a classification branch and a regression branch. The classification branch is used to distinguish between objects and background in the search image, while the regression branch is used to determine the location information of objects. For example, the updated template features and search features can be input into the classification branch for convolution to obtain classification features; simultaneously, the updated template features and search features can be input into the regression branch for convolution to obtain regression features. The elements in the classification features can represent the probability that each anchor belongs to the background or object category, while the elements in the regression features can represent the offset of each anchor between adjacent images. Both the classification and regression features can have a dimension of 17*17, etc., and an anchor can refer to a pixel on the image used to represent the center position of an object.
[0084] Given the classification and regression features, the element with the largest value in the classification features can be determined as the target anchor point, and the offset corresponding to the target anchor point can be determined from the regression features. The offset and the center coordinates in the initial position information are added or subtracted to obtain the center coordinates in the second position information. At the same time, a scaling ratio matching the offset can be determined, and the boundary length and width in the initial position information are scaled according to the scaling ratio to obtain the boundary coordinates in the second position information.
[0085] Before implementing step S402, the process may further include: receiving a user-uploaded object tracking dataset (such as the GOT-10K dataset) and object detection dataset (such as the COCO object detection dataset), and at least determining the object tracking dataset and object detection dataset as sample sets for training the untrained Siamese model; performing data augmentation (such as translation, resizing, grayscale transformation, etc.) on the data in the sample set to obtain an expanded sample set; responding to user annotations on the expanded sample set, training the untrained Siamese model using the annotated and expanded sample set to obtain a trained Siamese model; wherein, during the training process, a spatial attention mechanism may also be introduced, such as adjusting the weight ratio between the region where the object is located and the region where the background is located in the image of the sample set. The sample set may also include a tracking training set (ILSVRC) and a tracking training set (Youtube-BB), etc.
[0086] In this embodiment, the Siamese model is trained using a spatial attention mechanism based on an augmented sample set. The augmented sample set is obtained by performing data augmentation on data from selected target tracking and target detection datasets. Therefore, using the trained Siamese model, template features and search features can be obtained quickly and accurately, thereby rapidly determining the second location information of the object to be detected.
[0087] The following describes the application of the object detection method in a video stream provided in this disclosure in a real-world scenario, using a scenario where the location information of objects in a video stream is detected using a trained SiamRPN model as an example. Figure 5 As shown, the Siamese model may include a Siamese network 501 and a region generation network 502, etc. The Siamese network 501 may include a convolutional neural network 5011, a convolutional neural network 5022, and a feature space attention unit 5013, etc. The region generation network 502 may include convolutional layers 5021, 5022, ..., 5026, etc. The convolutional neural network 5011 can be used to extract template features 512 from the template image 510, where the size of the template image 510 can be 127*127*3, and the dimension of the template features 512 can be 6*6*256; the convolutional neural network 5012 can be used to extract search features 513 from the search image 511, where the size of the search image 511 can be 255*255*3, and the dimension of the search features 513 can be 22*22*256, etc.; the feature space attention unit 5013 can be used to determine the weights corresponding to objects and background in the template image, etc.
[0088] Convolutional layer 5021 can perform convolution processing on template feature 512 to obtain first convolutional feature 514, the dimension of first convolutional feature 514 can be 4*4*(2k*256), where k can represent the number of preset anchor points, such as k=5; convolutional layer 5022 can perform convolution processing on search feature 513 to obtain second convolutional feature 515, the dimension of second convolutional feature 515 can be 20*20*256; convolutional layer 5023 can perform convolution processing on template feature 512 to obtain third convolutional feature 516, the dimension of third convolutional feature 516 can be 4*4*(4k*256); convolutional layer 5024 can perform convolution processing on search feature 513 to obtain fourth convolutional feature 517, the dimension of fourth convolutional feature 517 can be 20*20*256, and so on. Convolutional layer 5025 can perform convolution processing on the first convolutional feature 514 and the second convolutional feature 515 to obtain classification feature 518, the dimension of classification feature 518 can be 17*17*2k; convolutional layer 5026 can perform convolution processing on the third convolutional feature 516 and the fourth convolutional feature 517 to obtain regression feature 519, the dimension of regression feature 519 can be 17*17*4k, etc.
[0089] This disclosure provides a method for detecting objects in a video stream, such as... Figure 6 As shown, the method may include the following steps S601 to S613:
[0090] Step S601: Perform data augmentation on the sample set to obtain an expanded sample set.
[0091] Here, the GOT-10K and COCO object detection datasets can be introduced into the sample set, and the data in the GOT-10K and COCO object detection datasets can be enhanced using a series of data augmentation techniques (such as translation, resizing, grayscale transformation, etc.) to obtain an expanded sample set. This expands the types of positive sample pairs in the sample set, and at the same time, the same data augmentation techniques are used to increase the types of negative sample pairs of the same and different categories. For example... Figure 7 As shown, the positive sample pairs in the expanded sample set can be images 701 and 702, the negative sample pairs of the same class can be images 703 and 704, and the negative sample pairs of different classes can be images 705 and 706, etc. The size of each image in the sample pair can be different. For example, the size of image 701 is smaller than the size of image 702. Image 701 can be used as a template image in the training process, and image 702 can be used as the search image corresponding to the template image in the training process, etc.
[0092] For example, in a video stream containing multiple objects A, B, and C, the object to be detected is A. Using a Siamese model, the objects to be detected are classified on the video stream images. The predicted type value for object A is labeled as 1, the predicted type value for object B (a negative sample pair of the same class) is labeled as 0, and the predicted type value for object C (a negative sample pair of a different class) is also labeled as 0. In related techniques, if the sample set is not expanded to include sample pairs of different classes, and object C has a high similarity to object A, then there is a high probability that the predicted type value for object C will be labeled as 1, leading to target drift. Here, by performing data augmentation on the images in the sample set to add negative sample pairs of different classes, the probability of labeling the predicted type value for object C as 1 is lower when object C is present, helping to reduce the occurrence of target drift.
[0093] Step S602: Train the initial model based on the expanded sample set to obtain the trained twin model.
[0094] Here, a spatial feature attention mechanism can be used to adjust the parameters of the preset initial model based on the expanded sample set to obtain a trained twin model. The initial model can refer to an untrained twin model, and the twin model can be trained offline.
[0095] Step S603: Obtain the template image corresponding to the first image in the video stream, and use the twin model to extract features from the template image to obtain template features.
[0096] Here, the system can receive initial location information uploaded by the user. This initial location information represents the position of the object to be detected in the first image of the video stream. Based on the initial location information, the first image is cropped to obtain a template image. Then, the template features corresponding to the template image are extracted using a Siamese model, thus initializing the Siamese model.
[0097] Step S604: Obtain the second image from the video stream.
[0098] Here, the video stream can be decoded to obtain the first image, the second image, and the third image, etc.
[0099] Step S605: Determine the search image corresponding to the second image.
[0100] Here, the second image in the video stream can be cropped based on the initial position information to obtain the search image of the second image.
[0101] Step S606: Determine the search features of the search image.
[0102] Here, the twin model can be used to extract features from the search image to obtain the search features.
[0103] Step S607: Determine the cosine similarity between the search feature and each feature in the preset template feature set.
[0104] Here, to determine the cosine similarity corresponding to the search features of the second image, the template feature set may initially only include the template features of the first image; to determine the cosine similarity corresponding to the search features of the third image, the template feature set may include the template features of the first image and the search features of the second image.
[0105] Step S608: Determine whether all cosine similarities are greater than the similarity threshold.
[0106] Here, it is determined whether each cosine similarity corresponding to the search feature is greater than the similarity threshold. If they are all greater than the threshold, proceed to step S609 to update the template feature set. If there is at least one cosine similarity less than or equal to the similarity threshold, proceed to step S611 and do not need to update the template feature set.
[0107] Step S609: Update the template feature set to obtain the updated template feature set.
[0108] Here, if the number of features in the template feature set is equal to the number threshold, the feature with the smallest cosine similarity to the search feature in the template feature set is removed, and the search feature is added to the template feature set; if the number of features in the template feature set is less than the number threshold, the search feature is added to the template feature set.
[0109] Step S610: Update the template features using an update method that matches cosine similarity to obtain the updated template features.
[0110] Here, when all cosine similarities are greater than the similarity threshold, the template features are updated based on the search features to obtain the updated template features; when at least one cosine similarity is less than or equal to the similarity threshold, the template features are updated based on the feature with the highest cosine similarity to the search features in the template feature set to obtain the updated template features.
[0111] Step S611: Stop updating the template feature set.
[0112] Here, there is no need to increase the number of features in the template feature set or replace the features in the template feature set.
[0113] Step S612: Determine whether the frame number of the current image is equal to the preset frame number.
[0114] Here, the video stream can be decoded to obtain images with a preset number of frames, such as the first image and the second image. All images are traversed according to their frame numbers to obtain the position information of the object to be detected in each image. During the traversal, if the frame number of the currently processed image is less than the preset number of frames, step S613 is entered to continue traversing the images of the video stream; if the frame number of the currently processed image is equal to the preset number of frames, the position detection of the object to be detected can be terminated.
[0115] Step S613: Determine the frame number of the current image in the video stream.
[0116] Here, the frame number of the currently processed image can be recorded, which helps to end the object position detection in a timely manner.
[0117] In this embodiment, expanding the sample set can reduce the imbalanced distribution of training data. For example, by introducing the GOT-10K dataset and the COCO object detection dataset, and using a series of data augmentation techniques to expand the types of positive sample pairs, negative sample pairs of the same and different categories in the sample set, it helps to improve the SiamRPN model's ability to discriminate objects and its regression accuracy, thereby improving object detection accuracy. Simultaneously, an online adaptive update mechanism is proposed to update the template features extracted from the first image online. For example, by calculating the cosine similarity between the search features of the second image and the features in the template feature set, the cosine similarity is used as a confidence evaluation index for the object detection result, and the template features of the first image are updated online adaptively based on the obtained cosine similarities.
[0118] Compared to the SiamRPN model in related technologies, the method of expanding the sample set in this embodiment enables the trained Siamese model to have better discrimination ability and regression accuracy for the object to be detected, and can better distinguish between the object and similar objects in the background, thereby improving the model's tracking accuracy and precision. Compared to the SiamRPN model in related technologies, which uses the features of the template image in the first image as fixed template features, this embodiment proposes an online adaptive update mechanism for template features. This mechanism adaptively updates the template features extracted from the template image online, which helps to capture changes in the object in subsequent images, improving the model's tracking accuracy and precision.
[0119] Based on the foregoing embodiments, this disclosure provides an object detection device in a video stream. The device includes various units and modules included in each unit, which can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field-programmable gate array (FPGA), etc.
[0120] Figure 8 This is a schematic diagram of the composition structure of a device for detecting objects in a video stream provided in an embodiment of this disclosure, as shown below. Figure 8 As shown, the object detection device 800 in the video stream includes: a first determining module 810, a second determining module 820, a third determining module 830, a fourth determining module 840, a first updating module 850, and a fifth determining module 860, wherein:
[0121] A first determining module 810 is used to determine the initial position information of the object to be detected in the first image of the video stream; a second determining module 820 is used to determine the template image corresponding to the first image based on the initial position information, and to generate a template feature set based on the template features of the template image; a third determining module 830 is used to determine the search image corresponding to the second image of the video stream based on the initial position information; wherein the frame number of the first image is located before the frame number of the second image; a fourth determining module 840 is used to determine the similarity between the search features of the search image and each feature in the template feature set; a first updating module 850 is used to update the template features using an update method matching the similarity, to obtain updated template features; a fifth determining module 860 is used to determine the second position information of the object to be detected in the second image based on the updated template features and the search features.
[0122] In some embodiments, the first updating module is further configured to: update the template features based on the search features when all the similarities are greater than or equal to the similarity threshold, so as to obtain the updated template features.
[0123] In some embodiments, the first updating module is further configured to: if at least one of the similarities is less than the similarity threshold, obtain the feature with the highest similarity to the search feature from the template feature set; update the template feature based on the feature with the highest similarity to the search feature to obtain the updated template feature.
[0124] In some embodiments, the apparatus further includes: a second updating module, configured to add the search features to the template feature set to obtain an updated template feature set when all the similarities are greater than or equal to a similarity threshold; and a sixth determining module, configured to determine the third location information of the object to be detected in the third image of the video stream based on the updated template feature set and the second location information; wherein the frame number of the second image is located before the frame number of the third image.
[0125] In some embodiments, the second updating module is further configured to: when the number of features in the template feature set is greater than a threshold, obtain the feature with the lowest similarity to the search feature from the template feature set; remove the feature with the lowest similarity to the search feature from the template feature set, and add the search feature to the template feature set to obtain the updated template feature set.
[0126] In some embodiments, the first updating module is further configured to: perform a weighted summation on the search features and the template features to obtain the updated template features.
[0127] In some embodiments, the apparatus further includes: an extraction module, configured to extract features from the template image and the search image using a Siamese network in a preset Siamese model to obtain the template features and the search features; the fifth determination module is further configured to: perform convolution processing on the updated template features and the search features using a region generation network in the Siamese model to obtain the center coordinates and boundary coordinates of the object to be detected in the second image; wherein the Siamese model is trained using a spatial attention mechanism based on an expanded sample set; the expanded sample set is at least obtained by performing data augmentation processing on data from a selected target tracking dataset and a target detection dataset.
[0128] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this disclosure can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.
[0129] It should be noted that, in the embodiments of this disclosure, if the above-described method for detecting objects in a video stream is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this disclosure, or the part that contributes to related technologies, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, external hard drive, read-only memory (ROM), magnetic disk, or optical disk. Thus, the embodiments of this disclosure are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0130] This disclosure provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.
[0131] This disclosure provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium may be transient or non-transient.
[0132] This disclosure provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.
[0133] This disclosure provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0134] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referenced interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.
[0135] It should be noted that, Figure 9 This is a schematic diagram of a hardware entity of a computer device in an embodiment of this disclosure, such as... Figure 9 As shown, the hardware entity of the computer device 900 includes: a processor 901, a communication interface 902, and a memory 903, wherein:
[0136] Processor 901 typically controls the overall operation of computer device 900.
[0137] Communication interface 902 enables computer devices to communicate with other terminals or servers over a network.
[0138] The memory 903 is configured to store instructions and applications executable by the processor 901, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 901 and various modules in the computer device 900. It can be implemented using flash memory or random access memory (RAM). Data transfer between the processor 901, the communication interface 902, and the memory 903 can be performed via bus 904.
[0139] It should be understood that the phrase "an embodiment" or "one embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this disclosure. Therefore, "in one embodiment" or "one embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this disclosure, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure. The sequence numbers of the above embodiments of this disclosure are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0140] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0141] In the several embodiments provided in this disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0142] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0143] In addition, each functional unit in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0144] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0145] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.
[0146] The methods disclosed in the several method embodiments provided in this disclosure can be arbitrarily combined without conflict to obtain new method embodiments.
[0147] If the embodiments of this disclosure involve personal information, the products using these embodiments have clearly informed the users of the personal information processing rules and obtained their voluntary consent before processing the personal information. If the embodiments of this disclosure involve sensitive personal information, the products using these embodiments have obtained the individual's separate consent before processing the sensitive personal information, and the requirement of "express consent" is also met.
[0148] The above description is merely an embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for detecting objects in a video stream, characterized in that, include: Determine the initial position information of the object to be detected in the first image of the video stream; Based on the initial position information, a template image corresponding to the first image is determined, and a template feature set is generated based on the template features of the template image; The search image corresponding to the second image of the video stream is determined based on the initial position information; wherein the frame number of the first image is located before the frame number of the second image; Determine the similarity between the search features of the search image and each feature in the template feature set; If all the similarities are greater than or equal to the similarity threshold, the template features are updated based on the search features to obtain the updated template features; If at least one of the similarities is less than the similarity threshold, the feature with the highest similarity to the search feature is obtained from the template feature set; the template feature is updated based on the feature with the highest similarity to the search feature to obtain the updated template feature; Based on the updated template features and the search features, the second location information of the object to be detected in the second image is determined.
2. The method according to claim 1, characterized in that, The method further includes: If all the similarities are greater than or equal to the similarity threshold, the search feature is added to the template feature set to obtain an updated template feature set; Based on the updated template feature set and the second location information, the third location information of the object to be detected in the third image of the video stream is determined; wherein the frame number of the second image is located before the frame number of the third image.
3. The method according to claim 2, characterized in that, The step of adding the search features to the template feature set to obtain the updated template feature set includes: If the number of features in the template feature set is greater than a threshold, the feature with the lowest similarity to the search feature is obtained from the template feature set. Remove the feature with the lowest similarity to the search feature from the template feature set, and add the search feature to the template feature set to obtain the updated template feature set.
4. The method according to claim 1, characterized in that, The step of updating the template features based on the search features to obtain the updated template features includes: The search features and the template features are weighted and summed to obtain the updated template features.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Using the Siamese network in the preset Siamese model, feature extraction is performed on the template image and the search image to obtain the template features and the search features; The step of determining the second location information of the object to be detected in the second image based on the updated template features and the search features includes: Using the region generation network in the Siamese model, the updated template features and the search features are convolved to obtain the center coordinates and boundary coordinates of the object to be detected in the second image; The twin model is trained using a spatial attention mechanism based on an expanded sample set; the expanded sample set is at least obtained by performing data augmentation processing on data from selected target tracking datasets and target detection datasets.
6. A device for detecting objects in a video stream, characterized in that, include: The first determining module is used to determine the initial position information of the object to be detected in the first image of the video stream; The second determining module is used to determine the template image corresponding to the first image based on the initial position information, and to generate a template feature set based on the template features of the template image; The third determining module is used to determine the search image corresponding to the second image of the video stream based on the initial position information; wherein the frame number of the first image is located before the frame number of the second image; The fourth determining module is used to determine the similarity between the search features of the search image and each feature in the template feature set; The first update module is configured to update the template features based on the search features when all the similarities are greater than or equal to a similarity threshold, to obtain the updated template features; and when at least one of the similarities is less than the similarity threshold, to obtain the feature with the highest similarity to the search features from the template feature set; and to update the template features based on the feature with the highest similarity to the search features, to obtain the updated template features. The fifth determining module is used to determine the second location information of the object to be detected in the second image based on the updated template features and the search features.
7. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Tracking method for target in video and tracking device thereof
CN105931269A
Single-target tracking method for dynamic double-template updating and storage medium
CN114387459A