Systems, methods, and media for retraining a pre-trained object classifier

By retraining the pre-trained object classifier and using the instances of the tracked object for annotation and training, the problem of high annotation cost and difficult data acquisition of the object classifier training data set is solved, and the performance and adaptability of the classifier are improved.

CN116263812BActive Publication Date: 2025-05-23AXIS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211586377.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-12-14
Filing Date
2022-12-09
Publication Date
2025-05-23
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

In the prior art, the training data set annotation cost of object classifiers is high, and it is difficult to obtain training data, resulting in insufficient performance of object classifiers.

Method used

By retraining the pre-trained object classifier, annotation and training is performed using the instances of the tracked object, the annotated instance is generated to improve the classifier performance.

Benefits of technology

Improve the performance of object classifiers, especially in difficult-to-classify areas in the scene, and enhance the universality and adaptability of classifiers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116263812B_ABST
    Figure CN116263812B_ABST
Patent Text Reader

Abstract

The present invention relates to systems, methods and media for retraining a pretrained object classifier. A mechanism for retraining a pretrained object classifier is provided. The method includes obtaining a stream of image frames of a scene. Each of the image frames depicts an instance of a tracked object. The method includes classifying each instance of the tracked object as belonging to an object category with a confidence level. The method includes verifying that the confidence level of at least one instance of the tracked object is higher than a threshold confidence value. When the confidence level of at least one instance of the tracked object is higher than the threshold confidence value, the method includes annotating all instances of the tracked object in the stream of image frames as belonging to the object category, generating annotated instances of the tracked object. The method includes retraining the pretrained object classifier with at least some of the annotated instances of the tracked object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments presented herein relate to methods, systems, computer programs, and computer program products for retraining a pre-trained object classifier. Background Art

[0002] Object recognition is an umbrella term used to describe a range of related computer vision tasks that involve identifying objects in an image frame. One such task is object classification, which involves predicting the class of an object in an image frame. Another such task is object localization, which refers to identifying the location of one or more objects in an image frame and optionally providing a bounding box surrounding such located objects. Object detection can be viewed as a combination of these two tasks, and can therefore both locate and classify one or more objects in an image frame.

[0003] One way to perform object classification is to compare the object to be classified with a set of template objects indicating different object categories, and then classify the object into the object category of the template object that is most similar to the object based on some metrics. Classification can be based on a model trained with a dataset of annotated objects. A model for object classification can be developed using deep learning (DL) techniques. The generated model can be referred to as an object classification DL model, or simply a model. Such a model can be trained based on a constantly improving dataset. During a firmware update or similar update process, the trained and updated model needs to be sent to a system using an old version of the model.

[0004] Generally speaking, a model is only as good as it is trained. Therefore, a certain amount of further training of the model will improve the classification. In other words, the more training data there is, the better the performance of the object classifier (in terms of accuracy of classification, computational speed of classification that can be performed, etc.). In order to capture as much training data as possible and be as general as possible, models are traditionally statically pre-trained general models that apply to the training dataset, regardless of the conditions under which the image frames in the training data were captured. This typically requires annotating the training data. However, the cost of annotation, both manually and by computer computation, as well as privacy regulations, can limit the possibility of annotating large training datasets and the availability of the actual training data itself.

[0005] US 2017 / 039455A1 relates to a method for protecting the environment.

[0006] Foroughi Homa et al., “Robust People Counting Using Sparse Representations and Random Projections,” published in Pattern Recognition, Vol. 48, No. 10, 1 October 2015, pp. 3038-3052, relates to methods for estimating the number of people present in an image for practical applications including visual surveillance and public resource management.

[0007] US 2021 / 042530A1 relates to aspects of artificial intelligence (AI)-driven ground truth generation for object detection and tracking of image sequences.

[0008] Therefore, there still exists a need for the training of improved object classifiers. Summary of the invention

[0009] The purpose of the embodiments herein is to solve the above problems.

[0010] In general, according to the inventive concepts disclosed herein, improved training of an image classifier is achieved by retraining a pre-trained object classifier. Then, according to the inventive concepts disclosed herein, this is achieved by annotating an instance of a tracked object as belonging to the same object category based on the fact that another instance of the same tracked object has been verified to belong to an object category.

[0011] According to a first aspect, the inventive concept is defined by a method for retraining a pre-trained object classifier. The method is performed by a system including a processing circuit. The method includes obtaining a stream of image frames of a scene. Each of the image frames depicts an instance of a tracked object. The tracked object is the same object that is tracked when moving in the scene. The method includes classifying each instance of the tracked object as belonging to an object category with a confidence level. The method includes verifying that the confidence level of at least one instance of the tracked object for only one object category is higher than a threshold confidence value. It can therefore be ensured that at least one instance of the tracked object is classified to only one object category with high confidence. When at least one instance of the tracked object is classified to only one object category with high confidence, the method includes annotating all instances of the tracked object in the stream of image frames as belonging to only one object category with high confidence (i.e., the confidence level of at least one instance of the tracked object is higher than the threshold confidence value of the object category), generating annotated instances of the tracked object. The method includes retraining the pre-trained object classifier with at least some of the annotated instances of the tracked object.

[0012] According to a second aspect, the inventive concept is defined by a system for retraining a pre-trained object classifier. The system includes a processing circuit. The processing circuit is configured to cause the system to obtain a stream of image frames of a scene. Each of the image frames depicts an instance of a tracked object. The tracked object is the same object that is tracked when moving in the scene. The processing circuit is configured to cause the system to classify each instance of the tracked object as belonging to an object category with a confidence level. The processing circuit is configured to cause the system to verify that the confidence level of at least one instance of the tracked object for only one object category is higher than a threshold confidence value. It can therefore be ensured that at least one instance of the tracked object is classified into only one object category with high confidence. When at least one instance of the tracked object is classified into only one object category with high confidence, the processing circuit is configured to cause the system to annotate all instances of the tracked object in the stream of image frames as belonging to only one object category with high confidence (i.e., the confidence level of at least one instance of the tracked object is higher than the object category of the threshold confidence value), generating annotated instances of the tracked object. The processing circuitry is configured to cause the system to retrain the pre-trained object classifier with at least some of the annotated instances of the tracked objects.

[0013] According to a third aspect, the inventive concept is defined by a system for retraining a pre-trained object classifier. The system includes an acquirer module configured to acquire a stream of image frames of a scene. Each of the image frames depicts an instance of a tracked object. The tracked object is the same object that is tracked when moving in the scene. The system includes a classifier module configured to classify each instance of the tracked object as belonging to an object category with a confidence level. The system includes a verifier module configured to verify that the confidence level of at least one instance of the tracked object for only one object category is higher than a threshold confidence value. It can therefore be ensured that at least one instance of the tracked object is classified into only one object category with high confidence. The system includes an annotator module configured to annotate all instances of the tracked object in the stream of image frames as belonging to only one object category with high confidence (i.e., the confidence level of at least one instance of the tracked object is higher than the threshold confidence value of the object category), generating annotated instances of the tracked object. The system includes a retrainer module configured to retrain a pre-trained object classifier with at least some of the annotated instances of tracked objects.

[0014] According to a fourth aspect, the inventive concept is defined by a computer program for retraining a pre-trained object classifier, the computer program comprising computer program code which, when run on a system, causes the system to perform the method according to the first aspect.

[0015] According to a fifth aspect, the inventive concept is defined by a computer program product comprising a computer program according to the fourth aspect and a computer readable storage medium on which the computer program is stored. The computer readable storage medium may be a non-transitory computer readable storage medium.

[0016] Advantageously, these aspects provide improved training of pre-trained object classifiers.

[0017] In addition to the pre-trained object classifiers mentioned in the above aspects, annotated instances of tracked objects can also be used to train or retrain other object classifiers. Therefore, advantageously, these aspects provide a means for automatically generating a large amount of annotated training data from a collection of only partially annotated training data.

[0018] Other objectives, features and advantages of the enclosed embodiments will be apparent from the following detailed disclosure, from the attached dependent claims as well as from the drawings.

[0019] Generally, all terms used in the claims should be interpreted according to their ordinary meaning in the technical field, unless otherwise explicitly defined herein. All references to "an element, device, component, means, module, step, etc." should be openly interpreted as referring to at least one instance of an element, device, component, means, module, step, etc., unless otherwise explicitly stated. Unless explicitly stated, the steps of any method disclosed herein do not have to be performed in the exact order disclosed. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The inventive concept will now be described by way of example with reference to the accompanying drawings, in which:

[0021] Figure 1 schematically illustrates a system according to an embodiment;

[0022] Figure 2 is a block diagram of a system according to an embodiment;

[0023] Figure 3 is a flow chart of a method according to an embodiment;

[0024] Figure 4 schematically illustrating annotation of a stream of image frames according to an embodiment;

[0025] Figure 5 is a schematic diagram showing functional units of a system according to an embodiment; and

[0026] Figure 6 An example of a computer program product including a computer-readable storage medium according to an embodiment is shown. DETAILED DESCRIPTION

[0027] The inventive concept will now be described more fully below with reference to the accompanying drawings in which certain embodiments of the inventive concept are shown. However, the inventive concept may be embodied in many different forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided by way of example so that the disclosure will be thorough and complete and will fully convey the scope of the inventive concept to those skilled in the art. Throughout the specification, the same reference numerals refer to the same elements. Any steps or features illustrated by dashed lines should be considered optional.

[0028] As mentioned above, there still exists a need for the training of improved object classifiers.

[0029] In general, as recognized by the inventors of the inventive concepts disclosed herein, the classification results of an object classifier of an object may depend on the angle at which an image frame of the object is captured, the lighting conditions at which the image frame of the object is captured, the image resolution of the image frame, etc. Therefore, some angles, lighting conditions, image resolutions, etc. bring higher confidence levels than other angles, lighting conditions, image resolutions, etc.

[0030] This means that as a tracked object moves along a path through a scene, at some parts of the path it is likely that the tracked object will be relatively easily classified by the object classifier used, and at some other parts of the path it will be relatively difficult to classify the object using the same object classifier due to varying angles, lighting conditions, etc. However, this fact cannot be exploited because the object tracker and the object classifier typically run independently of each other.

[0031] A tracked object that is relatively easily classified will generally result in a high confidence value, while a tracked object that is relatively difficult to classify will generally result in a low confidence value for the object in a particular image frame. Since an object is unlikely to change from one object category to another during its movement through a scene, e.g., from a "person" to a "truck," this means that if an object is classified with a confidence level in any one instance, e.g., as a "person" with high confidence, the underlying physical object indicated by the tracked object will remain a "person," even if the tracked object is in a later stage of its path and therefore more difficult to classify in a later image frame, e.g., the tracked object is still classified as a "person," but only with a low confidence. The inventive concept disclosed herein exploits this recognition. It is to be noted here that the task of an object classifier is to classify a tracked object as belonging to any given object category to some level of confidence. The embodiments disclosed herein are based on the assumption that the object classifier behaves correctly and does not provide false positives, and therefore indicate the normal behavior of any object classifier.

[0032] Specifically, according to the inventive concepts disclosed herein, this recognition is exploited to generate new and classified training data for a pre-trained object classifier. Therefore, at least some of the embodiments disclosed herein relate to the automatic collection and annotation of training data for general refined object classification or for a specific camera installation. The embodiments disclosed herein specifically relate to a mechanism for retraining a pre-trained object classifier. To obtain such a mechanism, a system, a method performed by the system, and a computer program product including code, for example in the form of a computer program, are provided, which, when run on the system, causes the system to perform the method.

[0033] Figure 1 1 is a schematic diagram illustrating a system 110 to which embodiments presented herein may be applied. The system 110 includes at least one camera device 120a, 120b. Each camera device 120a, 120b is configured to capture image frames depicting a scene 160 within a respective field of view 150a, 150b. Figure 1 , scene 160 includes a walking male. The walking male indicates a tracked object 170 as surrounded by a bounding box. Tracked object 170 is a target object that is tracked using object tracking. In general, object tracking refers to a technique for estimating or predicting the position (and possible other properties) of a target object in each consecutive frame in a video clip as defined by a stream consisting of a finite number of image frames once the initial position of the target object is defined. Therefore, object tracking can be viewed as a deep learning process in which the motion of the target object is tracked. Object tracking generally involves object detection, in which an algorithm detects, classifies, and annotates the target object. A (unique) identification can be assigned to each detected object. The detected object is then tracked (while storing relevant information) as it moves through the image frames. Object detection only works when the target object is visible in the image frames. However, as will be further disclosed below, once the target object is detected forward and backward in the stream of image frames, the target object can still be tracked even if the target object is partially obscured in some of the image frames. The position of the tracked target object in two or more image frames may be considered to be indicative of the trajectory of the tracked object 170. Thus, a trajectory may be constructed by following the position of the tracked object 170 from image frame to image frame. However, since the tracked object 170 may be partially obscured in some of the image frames, the confidence level may vary from image frame to image frame.

[0034] In some examples, each camera device 120a, 120b is a digital camera device and / or is capable of pan, tilt and zoom (PTZ) and may therefore be considered a (digital) PTZ camera device. The system 110 is configured to communicate with a user interface 130 to display captured image frames. Further, the system 110 is configured to encode the images so that the encoded images can be decoded using any known video coding standard, such as: High Efficiency Video Coding (HEVC), also known as H.265 and MPEG-H Part 2; Advanced Video Coding (AVC), also known as H.264 and MPEG-4 Part 10; Versatile Video Coding (VVC), also known as H.266, MPEG-I Part 3 and Future Video Coding (FVC); any of VP9, ​​VP10 and AOMedia Video 1 (AV1), to name just a few examples. In this regard, the encoding may be performed directly in conjunction with the camera devices 120a, 120b capturing the image frames, or performed at another entity and then at least temporarily stored in the database 116. The system 110 further includes a first entity 114a and a second entity 114b. Each of the first entity 114a and the second entity 114b can be implemented in the same computing device or in separate computing devices. The first entity 114a can, for example, be configured to obtain a stream of image frames of the scene 160 and provide an instance of the tracked object 170 to the second entity 114b. Further details of the first entity 114a will be referred to Figure 2 The second entity 114b may be an object classifier 114b. The object classifier 114b is assumed to be pre-trained and is therefore referred to as pre-trained object classifier 114b hereinafter. The system 110 may include further entities, functions, nodes, and devices, as indicated by device 118. Examples of such further entities, functions, nodes, and devices are computing devices, communication devices, servers, sensors, etc.

[0035] In some aspects, the first entity 114a, the second entity 114b, and the database 116 define a video management system (VMS). Therefore, the first entity 114a, the second entity 114b, and the database 116 are considered to be part of the system 110. The first entity 114a, the second entity 114b, and the database 116 are operably connected to the camera devices 120a, 120b via the network 112. The network 112 can be wired, wireless, or partially wired and partially wireless. In some examples, the system 110 includes a communication interface 520 (e.g., Figure 5), the communication interface 520 is configured to communicate with the user interface 130 to display the image frames. An example connection between the system 110 and the user interface 130 is illustrated at reference numeral 140. The connection 140 can be wired, wireless, or partially wired and partially wireless. The user 180 can interact with the user interface 130. It should be understood that because the user interface 130 is configured to display the image frames captured by the camera devices 120a, 120b, the user interface 130 is at least partially a visual user interface.

[0036] Now the parallel reference Figure 2 and Figure 3 Embodiments for retraining a pre-trained object classifier 114b are disclosed. Figure 2 is a block diagram of a system 110 according to an embodiment. Figure 3 1 is a flow chart illustrating an embodiment of a method for retraining a pre-trained object classifier 114b. The method is performed by the system 110. The method is preferably provided as a computer program 620 (see Figure 6 ).

[0037] S102: Obtain a stream of image frames of the scene 160. The stream of image frames may be obtained by an obtainer module 410. Each of the image frames depicts an instance of a tracked object 170. The tracked object 170 is the same object 170 that is tracked as it moves in the scene 160. The stream of image frames may be obtained by the obtainer module 410 from at least one camera device 120a, 120b, such as Figure 2 as shown in .

[0038] S104 : Each instance of the tracked object 170 is classified as belonging to an object category with a confidence level. The classification may be performed by the classifier module 420 .

[0039] S106: Verify that the confidence level of at least one of the instances of the tracked object 170 for only one object category is higher than the threshold confidence value. The verification can be performed by the verifier module 430. Thereby, it can be ensured that at least one of the instances of the tracked object 170 is classified into only one object category with high confidence. Therefore, for only one object category, the confidence level of the instance of the tracked object 170 used as a reference needs to be higher than the threshold confidence value. This threshold confidence value can be adjusted up or down as desired to change the accuracy of the retraining of the pre-trained object classifier 114b. For example, by setting a very high threshold confidence value (such as 0.95 in the range of 0.0 to 1.0), the accuracy will be high. In addition, if it can be observed that the amount of wrong labels generated is higher than expected, the threshold confidence value can be adjusted upward to increase the accuracy of the retraining of the pre-trained object classifier 114b.

[0040] S118: All instances of the tracked object 170 in the stream of image frames are annotated as belonging to only one object category with a high confidence value (i.e., the confidence level for the object category of at least one of the instances of the tracked object 170 is above a threshold confidence value). This annotation generates annotated instances of the tracked object 170. The annotation may be performed by the annotator module 440.

[0041] If it can be successfully verified that the confidence level of at least one of the instances of the tracked object 170 is above a threshold confidence value, one or more instances of the same object that were classified only with a medium confidence level or a low confidence level in other image frames can therefore be extracted and annotated as if they were classified with a high confidence level.

[0042] Therefore, in this regard, S118 is entered only when it can be successfully verified that the confidence level of at least one of the instances of the tracked object 170 is higher than the threshold confidence value in S106. As described above, through S106, it can be ensured that the tracked object 170 is classified into only one object category with high confidence before entering S118.

[0043] S122 : The pre-trained object classifier 114 b is retrained with at least some of the annotated instances of the tracked object 170 . The retraining may be performed by the retrainer module 460 .

[0044] Thus, the pre-trained object classifier 114b can be retrained with additional annotated instances of the tracked object 170. If this is done for the object classification model, then over time the object classification model will be better able to provide improved object classification in areas of the scene 160 that were originally more difficult to classify (i.e., areas where the classification was assigned a low or medium confidence). Figure 4 Provide such an example.

[0045] Therefore, if it is verified in S106 that the tracked object 170 is classified as belonging to only one object category with high confidence in at least one image frame, the same tracked object 170 from other image frames is used for retraining of the pre-trained object classifier 114b. Similarly, if the tracked object 170 has never been classified as belonging to one object category with high confidence, the tracked object 170 in other image frames will not be used for retraining of the pre-trained object classifier 114b.

[0046] Embodiments involving further details of retraining the pre-trained object classifier 114b as performed by the system 110 will now be disclosed.

[0047] like Figure 2As indicated in , classification and retraining can be performed at separate devices. That is, in some embodiments, classification is performed at a first entity 114a, and retraining is performed at a second entity 114b that is physically separated from the first entity 114a. The second entity can be an object classifier 114b (pre-trained object classifier 114b).

[0048] Before all instances of tracked objects 170 in the stream of image frames are annotated as belonging to an object class, and thus before entering S118, further verification may be made. Different embodiments related thereto will now be described in turn.

[0049] It is possible that some tracked objects 170 are classified as belonging to two or more object categories, each classification having its own confidence level. Thus, in some aspects, only one of these object categories is verified with a high confidence. Thus, in some embodiments, at least some of the instances of tracked objects 170 are classified as also belonging to a further object category with a further confidence level, and the method further comprises:

[0050] S108 : Verifying that the further confidence level is lower than the threshold confidence value of at least some of the instances of the tracked object 170 . The verification may be performed by the verifier module 430 .

[0051] Therefore, if the tracked object 170 is classified into two or more different object categories with high confidence, the tracked object 170 will not be used for retraining of the pre-trained object classifier 114b. Likewise, if the tracked object 170 has never been classified into any object category with high confidence, the tracked object 170 will not be used for retraining.

[0052] In some aspects, the path along which the tracked object 170 moves from one image frame to the next is also tracked. If the tracked object 170 is classified as belonging to only one object category at least once with high confidence, the path can be used for retraining of the pre-trained object classifier 114b. Thus, in some embodiments, the tracked object 170 moves along a path in the stream of image frames, and the path is tracked as the tracked object 170 is tracked.

[0053] In some aspects, it is then verified that the path itself can be tracked with a high degree of accuracy. That is, in some embodiments, the path is tracked with a level of accuracy, and the method further comprises:

[0054] S110 : Verify that the accuracy level is above a threshold accuracy value. The verification may be performed by the verifier module 430 .

[0055] In some aspects, a verification is performed that the path has neither split nor merged. If the path has split and / or merged at least once, this may be an indication that the level of accuracy with which the path is tracked is not above a threshold accuracy value. This may be the case if it is suspected that the path has merged from two or more other paths, or has split into two or more other paths. If so, the path is determined to be of low accuracy, and the path will not be used to retrain the pre-trained object classifier 114b. Therefore, in some embodiments, the method further includes:

[0056] S112: Verify that the path is neither split into at least two paths within the stream of the image frame nor that the stream of the image frame is merged from at least two paths. In other words, verify that the path is not split or merged and thus constitutes a single path. The verification can be performed by the verifier module 430.

[0057] The same principle may also be applied if the tracked object 170 is suspected to have strange size behavior that may be suspected to be caused by shadows, mirror effects, or the like. If so, the tracked object 170 is assumed to be classified with low confidence, and the tracked object 170 will not be used to retrain the pre-trained object classifier 114b. Therefore, in some embodiments, the tracked object 170 has a size in the image frame, and the method further includes:

[0058] S114 : Verify that the size of the tracked object 170 does not vary by more than a threshold size value within the stream of image frames. The verification may be performed by the verifier module 430 .

[0059] Since the apparent size of the tracked object 170 depends on the distance between the tracked object 170 and the camera devices 120a, 120b, the size of the tracked object 170 relative to this distance can be compensated as part of the verification in S114. Specifically, in some embodiments, when verifying that the size of the tracked object 170 does not vary by more than a threshold size value within the stream of image frames 210a:210c, the size of the tracked object 170 is adjusted by a distance-dependent compensation factor, which is determined as a function of the distance between the tracked object 170 and the camera devices 120a, 120b of the stream of image frames capturing the scene 160.

[0060] As in Figure 1 As shown in the figure and as reference Figure 1As disclosed, the scene 160 may be captured by one or more camera devices 120a, 120b. In the case where the scene 160 is captured by two or more camera devices 120a, 120b, there is a stream of image frames of the tracked object 170 may therefore also be captured by two or more different camera devices 120a, 120b. Therefore, in some embodiments, the stream of image frames originates from image frames captured by at least two camera devices 120a, 120b. If the tracking of the tracked object 170 is performed locally at each of the camera devices 120a, 120b, this may require the information of the tracked object 170 to be transferred between the at least two camera devices 120a, 120b. In other examples, the tracking of the tracked object 170 is performed centrally on the image frames received from all of the at least two camera devices 120a, 120b. The latter does not require any information about the tracked object 170 to be exchanged between the at least two camera devices 120a, 120b.

[0061] In some aspects, such as when the stream of image frames originates from image frames captured by at least two camera devices 120a, 120b, but also in other examples, there is a risk of losing tracking of the tracked object 170 and / or a risk of the classification of the tracked object 170 changing from one image frame to the next. Therefore, in some embodiments, the method further comprises:

[0062] S116 : Verify that the object category of the instance of the tracked object 170 has not changed within the stream of image frames. The verification may be performed by the verifier module 430 .

[0063] In some aspects, in order to avoid training bias (e.g., machine learning bias, algorithmic bias, or artificial intelligence bias), the pre-trained object classifier 114b is not retrained using any tracked object 170 that has been classified with a high confidence level. This can be achieved in different ways. The first way is to explicitly exclude tracked objects 170 that have been classified with a high confidence level from retraining. Specifically, in some embodiments, the pre-trained object classifier 114b is retrained only with annotated instances of tracked objects 170 whose confidence level is verified to be no higher than a threshold confidence value. The second way is to set a low weight value for tracked objects 170 that have been classified with a high confidence level during retraining. In this way, tracked objects 170 that have been classified with a high confidence level can be implicitly excluded from retraining. That is, in some embodiments, each of the annotated instances of the tracked object 170 is assigned a corresponding weight value, and when the pre-trained object classifier 114b is retrained, the annotated instances of the tracked object 170 are weighted according to the weight value, and the weight value of the annotated instances of the tracked object 170 whose confidence level is verified to be higher than the threshold confidence value is lower than the weight value of the annotated instances of the tracked object 170 whose confidence level is verified to be not higher than the threshold confidence value. Thus, by means of the weight value, the object classifier 114b will be less affected by the tracked objects 170 that have been classified with a high confidence level during its retraining.

[0064] In addition to retraining the pre-trained object classifier 114b, the annotated instances of the tracked object 170 may have further different uses. In some aspects, the annotated instances of the tracked object 170 may be collected at the database 116 and / or provided to the further device 118. Therefore, in some embodiments, the method further includes:

[0065] S120: The annotated instances of the tracked object 170 are provided to the database 116 and / or the further device 118. The annotated instances may be provided to the database 116 and / or the further device 118 by the provider module 450.

[0066] This also enables other pre-trained object classifiers to benefit from the annotated instances of the tracked object 170 .

[0067] Next reference Figure 4 , Figure 4 Two streams 200, 200' of image frames 210a: 210c are schematically illustrated. Each stream 200, 200' consists of a finite sequence of image frames. Each such stream may correspond to a video segment captured in a scene 160 in which one or more tracked objects 170 are tracked. For illustrative purposes, each stream 200, 200' is Figure 4 The image is composed of three image frames 210a and 210c. Figure 4 2 illustrates a first stream 200 of image frames 210a:210c, wherein only the tracked object 170 in image frame 210c is classified with a confidence level above a threshold confidence value. Figure 4 In the example, for illustrative purposes, the tracked object 170 is assumed to be classified as belonging to the object class "walking male" with a confidence level above a threshold confidence value only in image frame 210c. This is because in image frames 210a and 210b, the lower half of the tracked object 170 is covered by leaves, which makes it difficult for the object classifier to determine whether the person is walking. Figure 4 2 also illustrates a second stream 200' of the same image frames 210a:210c after application of the embodiments disclosed herein, as indicated by the arrows labeled "Annotation". Through the application of the embodiments disclosed herein, the instances of the tracked image 170 in the image frames 210a, 210b are also annotated as belonging to the same object category as the instances of the tracked image 170 in the image frame 210c. This is shown in FIG. Figure 4 , illustrated by arrows 220a, 220b. Thus, despite the presence of partially occluding foliage, the instances of the tracked object 170 in the image frames 210a, 210b are annotated as being classified as belonging to the object class "walking male" with a confidence level above a threshold confidence value. In other words, the instances of the tracked object 170 in the image frames 210a, 210b inherit the annotations of the instance of the tracked object 170 in the image frame 210c. The instances of the tracked object 170 in the image frames 210a, 210b can then be used to retrain the pre-trained object classifier 114b. Thus, in order for the method to work, it is only necessary that the classification of the object 170 in one image frame 210c of the stream 200 is above a threshold value (e.g., Figure 4 200′). As a result of running the method, object 170 will also be annotated in image frames 210a and 210b in stream 200′. In turn, because object 170 has also been annotated in image frames 210a and 210b in stream 200′, this means that retraining pre-trained object classifier 114b with stream 200′ will indeed improve pre-trained object classifier 114b. More precisely, when pre-trained object classifier 114b is run on a new stream with image frames of partially obscured objects belonging to the object class “walking male”, such objects may also be tracked and annotated as being classified with a high confidence level.

[0068] Figure 5 The components of the system 110 according to an embodiment are schematically illustrated in terms of a number of functional units. Figure 6The processing circuit 510 is provided by any combination of one or more of a suitable central processing unit (CPU), a multiprocessor, a microcontroller, a digital signal processor (DSP), etc., in accordance with software instructions in a computer program product 610 (e.g., in the form of a storage medium 530) as shown in FIG. The processing circuit 510 may further be provided as at least one application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).

[0069] Specifically, the processing circuit 510 is configured to cause the system 110 to perform a set of operations or steps as described above. For example, the storage medium 530 may store the set of operations, and the processing circuit 510 may be configured to obtain the set of operations from the storage medium 530 to cause the system 110 to perform the set of operations. The set of operations may be provided as a set of executable instructions.

[0070] Thus, the processing circuit 510 is thereby arranged to perform the method as disclosed herein. The storage medium 530 may also include a permanent memory, for example, any single one or combination of a magnetic memory, an optical memory, a solid-state memory, or even a remotely mounted memory. The system 110 may further include a communication interface 520, which is at least configured to communicate with further devices, functions, nodes, and devices. As such, the communication interface 520 may include one or more transmitters and receivers including analog and digital components. The processing circuit 510 controls the general operation of the system 110, for example, by sending data and control signals to the communication interface 520 and the storage medium 530, by receiving data and reports from the communication interface 520, and by obtaining data and instructions from the storage medium 530. Other components of the system 110 and related functions are omitted to avoid confusing the concepts proposed herein.

[0071] Figure 6 An example of a computer program product 610 comprising a computer-readable storage medium 630 is shown. On this computer-readable storage medium 630, a computer program 620 may be stored that may enable the processing circuit 210 and entities and devices operatively coupled to the processing circuit 210, such as the communication interface 220 and the storage medium 230, to perform methods according to the embodiments described herein. The computer program 620 and / or computer program product 610 may thus provide a means for performing any of the steps disclosed herein.

[0072] exist Figure 6In the example of , computer program product 610 is illustrated as an optical disc such as a CD (compact disc) or a DVD (digital versatile disc) or a Blu-ray disc. Computer program product 610 may also be embodied as a memory such as a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or an electrically erasable programmable read-only memory (EEPROM), and more specifically, as a non-volatile storage medium of a device in an external memory such as a USB (universal serial bus) memory or in a flash memory such as a compact flash. Thus, although computer program 620 is schematically shown here as a track on the depicted optical disc, computer program 620 may be stored in any manner suitable for computer program product 610.

[0073] The inventive concept has mainly been described above with reference to a few embodiments. However, as readily appreciated by a person skilled in the art, other embodiments than the ones disclosed above are equally possible within the scope of the inventive concept as defined in the appended claims.

Claims

1. A method for retraining a pre-trained object classifier, the method being performed by a system including a processing circuit, the method comprises: Obtaining a stream of image frames of a scene, wherein each image frame in the image frames depicts an instance of a tracked object, and wherein the tracked object is the same object being tracked while moving in the scene; Classifying each instance of the tracked object at a confidence level as belonging to an object category; Verifying that the confidence level for at least one of the instances of the tracked object for only one object category is higher than a threshold confidence value, for thereby ensuring that at least one of the instances of the tracked object is classified to the only one object category with a high confidence; and when at least one of the instances of the tracked object is classified to the only one object category with the high confidence: Annotating all instances of the tracked object in the stream of image frames as belonging to the only one object category with a high confidence, generating annotated instances of the tracked object; and Retraining the pre-trained object classifier with at least some of the annotated instances of the tracked object.

2. The method according to claim 1, wherein, At least some of the instances of the tracked object are classified at a further confidence level as also belonging to a further object category, and wherein the method further comprises: Verifying that the further confidence level is lower than the threshold confidence value for at least some of the instances of the tracked object.

3. The method according to claim 1, wherein, The method further comprises: Verifying that the object category of the instances of the tracked object does not change in the stream of image frames.

4. The method according to claim 1, wherein, The tracked object moves along a path in the stream of image frames, and wherein the path is tracked while the tracked object is being tracked.

5. The method according to claim 4, wherein, The path is tracked at an accuracy level, and wherein the method further comprises: Verifying that the accuracy level is higher than a threshold accuracy value.

6. The method according to claim 4, wherein, The method further comprises: Verifying that the path does not split into at least two paths nor merge from at least two paths within the stream of image frames.

7. The method according to claim 1, wherein, The tracked object has a size in the image frame, and wherein the method further comprises: Verifying that the size of the tracked object does not change by more than a threshold size value within the stream of image frames.

8. The method according to claim 7, wherein, When verifying that the size of the tracked object does not change by more than the threshold size value within the stream of image frames, the size of the tracked object is adjusted by a distance-dependent compensation factor, the compensation factor being determined as a function of the distance between the tracked object and the camera device capturing the stream of image frames of the scene.

9. The method according to claim 1, in, The pre-trained object classifier is retrained using only the annotated instances of the tracked object for which the confidence level is verified to be no higher than the threshold confidence value.

10. The method according to claim 1, in, Each of the annotated instances of the tracked object is assigned a corresponding weight value, and when the pre-trained object classifier is retrained, the annotated instances of the tracked object are weighted according to the weight values, and wherein the weight value of the annotated instance of the tracked object whose confidence level is verified to be higher than the threshold confidence value is lower than the weight value of the annotated instance of the tracked object whose confidence level is verified to be not higher than the threshold confidence value.

11. The method according to claim 1, in, The method further comprises: The annotated instances of the tracked objects are provided to a database and / or to a further device.

12. The method according to claim 1, in, The stream of image frames is derived from image frames captured by at least two camera devices.

13. The method according to claim 1, in, The classification is performed at a first entity and the retraining is performed at a second entity that is physically separated from the first entity.

14. A system for retraining a pre-trained object classifier, the system comprising a processing circuit configured to cause the system to: Get a stream of image frames of the scene, in, Each of the image frames depicts an instance of a tracked object, and wherein the tracked object is the same object that is tracked while moving in the scene; classifying each instance of the tracked object as belonging to an object class with a confidence level; verifying that the confidence level for only one object class of at least one of the instances of the tracked object is above a threshold confidence value, for thereby ensuring that at least one of the instances of the tracked object is classified with high confidence into the only one object class; and when the at least one of the instances of the tracked object is classified with the high confidence into the only one object class: annotating all instances of the tracked object in the stream of image frames as belonging to the only one object category with high confidence, generating annotated instances of the tracked object; and The pre-trained object classifier is retrained using at least some of the annotated instances of the tracked object.

15. A non-transitory computer readable storage medium having a computer program stored thereon for retraining a pre-trained object classifier, the computer program comprising computer code that, when executed on a processing circuit of a system, causes the system to: Get a stream of image frames of the scene, in, Each of the image frames depicts an instance of a tracked object, and wherein the tracked object is the same object that is tracked while moving in the scene; classifying each instance of the tracked object as belonging to an object class with a confidence level; verifying that the confidence level for only one object class of at least one of the instances of the tracked object is above a threshold confidence value, for thereby ensuring that at least one of the instances of the tracked object is classified with high confidence into the only one object class; and when the at least one of the instances of the tracked object is classified with the high confidence into the only one object class: annotating all instances of the tracked object in the stream of image frames as belonging to the only one object category with high confidence, generating annotated instances of the tracked object; and The pre-trained object classifier is retrained using at least some of the annotated instances of the tracked object.

Citation Information

Patent Citations

  • Computer-vision based security system using a depth camera

    US20170039455A1

  • Artificial-intelligence powered ground truth generation for object detection and tracking on image sequences

    US20210042530A1

  • Target tracking method and device, computing equipment and storage medium

    CN111667501A

  • Method and device for clustering content samples

    CN111898704A