System, method, and computer program for retraining pre-trained object classifier

JP2023088294A5Active Publication Date: 2025-07-31AXIS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022196797
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-14
Filing Date
2022-12-09
Publication Date
2025-07-31
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

Existing object classifiers face challenges in improving training data annotation efficiency and accuracy due to high costs and privacy regulations, limiting the acquisition of large datasets, and the need for improved training methods to enhance classification performance.

Method used

Retraining a pre-trained object classifier by annotating instances of a tracked object with high confidence levels, ensuring they belong to only one object class, and using these annotated instances to improve the classifier's performance on difficult classification areas.

Benefits of technology

Automatically generates a large set of annotated training data from partially annotated data, enhancing the classifier's accuracy and ability to handle challenging classification scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a mechanism for retraining a pre-trained object classifier.SOLUTION: A method comprises obtaining a stream of image frames of a scene. Each of the image frames depicts an instance of a tracked object. The method comprises classifying, with a level of confidence, each instance of the tracked object so as to belong to an object class. The method comprises verifying that the level of confidence for at least one of the instances of the tracked object is higher than a threshold confidence value. The method comprises, when it is verified being higher than the threshold confidence value, annotating all instances of the tracked object in the stream of image frames as belonging to the object class, and yielding annotated instances of the tracked object. The method comprises retraining a pre-trained object classifier with at least some of the annotated instances of the tracked object.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments presented in this specification relate to a method, a system, a computer program, and a computer program product for retraining a pre-trained object classifier.

Background Art

[0002] Object recognition is a general term for describing a set of related computer vision tasks involving the identification of objects within an image frame. One such task is object classification, which involves predicting the class of an object within an image frame. Another such task is object position estimation, which refers to identifying the location of one or more objects within an image frame and optionally providing a bounding box surrounding the thus localized object. Object detection can be considered a combination of these two tasks and thus localizes and classifies one or more objects within an image frame.

[0003] One way to perform object classification is to compare the object to be classified to a set of template objects representing different object classes and then classify the object according to a metric into the object class of the template object that is most similar to the object. Classification can be based on a model trained on a dataset of annotated objects. A model for object classification can be developed using deep learning (DL) techniques. The resulting model can be referred to as an object classification DL model or simply a model. Such a model can be trained based on a continuously improving dataset. The trained and updated model needs to be sent to a system using an older version of the model during a firmware update or a similar update process.

[0004] Generally speaking, a model is only as good as it is when it is trained. Therefore, a certain amount of further training on the model may improve classification. In other words, the more training data there is, the better the object classifier may perform (in terms of classification accuracy, computational speed, etc.). To incorporate as much training data as possible and to be as general as possible, the model is a conventional static pre-trained general-purpose model that applies to the set of training data regardless of the conditions under which the image frames in the training data were taken. This generally requires annotating the training data. However, the cost of annotating, whether manually or computationally, and privacy regulations may limit the possibility of annotating large sets of training data, and the acquisition of actual training data itself.

[0005] U.S. Patent Application Publication No. 2017 / 039455 relates to methods for protecting the environment.

[0006] "Robust people counting using sparse representation and random projection" by Foroughi Homa et al., Pattern Recognition, Vol. 48, No. 10, October 1, 2015, pp. 3038-3052, describes a method for estimating the number of people present in an image for practical applications including visual surveillance and public resource management.

[0007] U.S. Patent Application Publication No. 2021 / 042530 relates to an artificial intelligence (AI)-based ground truth generation method for object detection and tracking in image sequences.

[0008] Therefore, improved training of object classifiers is still needed. [Overview of the project]

[0009] The object of the embodiments described herein is to address the above-mentioned problems.

[0010] Generally speaking, according to the inventive concept disclosed herein, improved training of an image classifier is achieved by retraining a pre-trained object classifier. Then, according to the inventive concept disclosed herein, this is achieved by annotating instances of the tracked object to belong to an object class, based on the fact that it has already been verified that another instance of the same tracked object belongs to the same object class.

[0011] According to a first aspect, the concept of the present invention is defined by a method for retraining a pre-trained object classifier. The method is performed by a system comprising a processing circuit. The method includes taking a stream of image frames of a scene. Each image frame depicts an instance of a tracked object. The tracked object is the exact same object that is tracked as it moves through the scene. The method includes classifying each instance of the tracked object by a confidence level so that it belongs to an object class. The method includes verifying that the confidence levels of at least one object class and just one object class of the instances of the tracked object are higher than a threshold confidence value. This ensures that at least one instance of the tracked object is classified with high confidence to just one object class. If the method verifies that it is higher than the threshold confidence value, it includes annotating all instances of the tracked object in the stream of image frames as belonging with high confidence to just one object class (i.e., the object class in which the confidence level of at least one instance of the tracked object is higher than the threshold confidence value), thereby yielding annotated instances of the tracked object. The method involves retraining a pre-trained object classifier using at least some of the annotated instances of the objects being tracked.

[0012] According to a second aspect, the concept of the present invention is defined by a system for retraining a pre-trained object classifier. The system comprises a processing circuit, which is configured to cause the system to acquire a stream of image frames of a scene, each of which depicts an instance of a tracked object, the tracked object being the exact same object being tracked as it moves through the scene. The processing circuit is configured to cause the system to classify each instance of the tracked object to belong to an object class at a confidence level. The processing circuit is configured to cause the system to verify that the confidence level of at least one instance of the tracked object and only one object class is higher than a threshold confidence value, thereby ensuring that at least one instance of the tracked object is classified with high confidence to only one object class. If it is verified that the confidence level is higher than the threshold confidence value, the processing circuit is configured to cause the system to annotate all instances of the tracked object in the stream of image frames with high confidence that they belong to only one object class (i.e., the object class in which at least one confidence level of the tracked object instance is higher than the threshold confidence value), thereby yielding annotated instances of the tracked object. The processing circuit is configured to cause the system to retrain a pre-trained object classifier using at least some of the annotated instances of the tracked object.

[0013] According to a third aspect, the concept of the present invention is defined by a system for retraining a pre-trained object classifier. The system comprises an acquirer module configured to acquire a stream of image frames of a scene. Each image frame depicts an instance of a tracked object. The tracked object is the exact same object that is tracked as it moves through the scene. The system comprises a classifier module configured to classify each instance of the tracked object to a level of confidence so that it belongs to an object class. The system comprises a validator module configured to verify that the confidence level of at least one instance of the tracked object and just one object class is higher than a threshold confidence value. This ensures that at least one instance of the tracked object is classified with high confidence to just one object class. The system comprises a processing circuit configured to annotate all instances of the object being tracked in a stream of image frames, resulting in annotated instances of the object being tracked, by assuming with high confidence that they belong to only one object class (i.e., an object class in which at least one confidence level of the instance of the object being tracked is higher than a threshold confidence value). The system comprises a retrainer module configured to retrain an object classifier that has been pre-trained on at least some of the annotated instances of the object being tracked.

[0014] According to a fourth aspect, the concept of the present invention is defined by a computer program for retraining a pre-trained object classifier, the computer program including computer program code that, when executed on a system, causes the system to perform the method according to the first aspect.

[0015] According to the fifth aspect, the concept of the present invention is defined by a computer program product including a computer program according to the fourth aspect, and a computer-readable storage medium in which the computer program is stored. The computer-readable storage medium may be a non-temporary computer-readable storage medium.

[0016] Advantageously, these embodiments provide improved training for pre-trained object classifiers.

[0017] Annotated instances of tracked objects can also be used to train or retrain object classifiers other than the pre-trained object classifiers described in the embodiments above. Therefore, advantageously, these embodiments provide a means for automatically generating a large set of annotated training data from a set containing only partially annotated training data.

[0018] Other purposes, features, and advantages of the attached embodiments will become apparent from the following detailed disclosure, from the attached dependent claims, and from the drawings.

[0019] In general, all terms used in the claims should be interpreted according to their ordinary meanings in the art unless specifically defined herein. All references to “a / an / the element, apparatus, component, means, module, step, etc.” should be broadly interpreted as referring to at least one example of an element, apparatus, component, means, module, step, etc. unless specifically stated otherwise. The steps of any method disclosed herein do not need to be performed in the exact order disclosed unless expressly stated otherwise.

[0020] Next, the concept of the present invention will be described with reference to the attached drawings as an example. [Brief explanation of the drawing]

[0021] [Figure 1] A schematic diagram of the system according to the embodiment is shown below. [Figure 2] It is a block diagram of a system according to an embodiment. [Figure 3] It is a flowchart of a method according to an embodiment. [Figure 4] It schematically shows the annotation of a stream of image frames according to an embodiment. [Figure 5] It is a schematic diagram showing functional units of a system according to an embodiment. [Figure 6] It shows an example of a computer program product including a computer-readable storage medium according to an embodiment. **Embodiments for Carrying Out the Invention**

[0022] Next, the inventive concept will be more fully described below with reference to the accompanying drawings that illustrate certain embodiments of the inventive concept. However, the inventive concept may be embodied in many different forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete and will fully convey the scope of the inventive concept to those skilled in the art. Throughout the description, like numbers refer to like elements. Any step or feature shown in dashed lines should be considered optional.

[0023] As described above, there is still a need for improved training of object classifiers.

[0024] Generally speaking, as realized by the inventors of the inventive concept disclosed herein, the classification result of an object classifier of an object may depend on the angle at which an image frame of the object is captured, the illumination condition under which the image frame of the object is captured, the image resolution of the image frame, etc. Therefore, some angles, illumination conditions, image resolutions, etc. result in a higher level of reliability than other angles, illumination conditions, image resolutions, etc.

[0025] This means that when the tracked object is moving along a path through the scene, between some portions of the path, the tracked object is likely to be relatively easy to classify by the object classifier used, and between some other portions of the path, due to changes such as angles, lighting conditions, etc., the object is likely to be relatively difficult to classify using the same object classifier. However, this fact cannot be utilized because object trackers and object classifiers typically operate independently of each other.

[0026] Tracked objects that are relatively easy to classify generally result in high confidence values, while tracked objects that are relatively difficult to classify generally result in low confidence values for the object within that particular image frame. Since it is unlikely that an object will change from one object class to another, e.g., from "person" to "truck", while moving through the scene, this means that if an instance of an object is classified as a "person" with a high confidence level, for example, the physical object represented by the tracked object will remain a "person" at later stages of its path and thus for the tracked object in later image frames, even if it is more difficult to classify, e.g., the tracked object will still be classified as a "person" with a low confidence. This realization is utilized by the inventive concept disclosed herein. Note that it is the task of the object classifier to classify the tracked object into a certain confidence level of belonging to any given object class. The embodiments disclosed herein are based on the assumption that the object classifier is operating correctly and does not provide misclassifications in this regard and thus represents the normal behavior of any object classifier.

[0027] In particular, according to the inventive concept disclosed herein, this implementation is used to generate new classified training data for a pre-trained object classifier. Accordingly, at least some of the embodiments disclosed herein relate to the automated storage and annotation of training data for sophisticated object classification or specific camera installations. Embodiments disclosed herein relate, in particular, to a mechanism for retraining a pre-trained object classifier. To obtain such a mechanism, a system, a method performed by the system, and a computer program product are provided which includes code in the form of a computer program that, when run on the system, causes the system to perform the method.

[0028] Figure 1 is a schematic diagram showing a system 110 to which embodiments of the embodiments presented herein may be applied. The system 110 comprises at least one camera device 120a, 120b. Each camera device 120a, 120b is configured to capture image frames depicting a scene 160 within their respective fields of view 150a, 150b. According to the exemplary example in Figure 1, the scene 160 includes a man walking. The walking man represents a tracked object 170, enclosed by a bounding box. The tracked object 170 is a target object being tracked using object tracking. Generally speaking, object tracking refers to a technique for estimating or predicting the position (and possible other properties) of a target object in each consecutive frame within a video segment, defined by a stream consisting of a limited number of image frames, given an initial position of the target object. Object tracking can be viewed as a deep learning process to which the movement of a target object is tracked. Object tracking typically involves object detection, where an algorithm detects, classifies, and annotates the target object. Each detected object may be assigned (unique) identification information. The detected object is then tracked as it moves through the image frames (while remembering the relevant information). Object detection will only work if the target object is visible within the image frames. However, as will be further disclosed below, if the target object is partially obscured in some of the image frames, once the target object is detected, it may still be possible to track the target object both forward and backward in the stream of image frames. The position of the target object tracked within two or more image frames can be considered to represent the tracking of the tracked object 170. Thus, a trajectory can be constructed by tracking the position of the tracked object 170 frame by frame. However, the confidence level may vary from image frame to image frame because the tracked object 170 may be partially obscured within some of the image frames.

[0029] In some examples, each camera device 120a, 120b is a digital camera device and / or capable of pan, tilt, and zoom (PTZ), and can therefore be considered a (digital) PTZ camera device. System 110 is configured to communicate with a user interface 130 for displaying captured image frames. Furthermore, System 110 is configured to encode images so that they can be decoded using any known video coding standard, such as, to name just a few examples, High Efficiency Video Coding (HEVC), also known as H.265 and MPEG-H Part 2; Advanced Video Coding (AVC), also known as H.264 and MPEG-4 Part 10; Versatile Video Coding (VVC), also known as H.266, MPEG-I Part 3, and Future Video Coding (FVC); VP9, ​​VP10, and AOMedia Video 1 (AV1). In this regard, encoding may be performed in direct cooperation with the camera devices 120a, 120b that capture image frames, or it may be performed by another entity and then stored, at least temporarily, in the database 116. System 110 further comprises a first entity 114a and a second entity 114b. Each of the first entity 114a and the second entity 114b may be implemented on either the same computing device or separate computing devices. The first entity 114a may be configured, for example, to acquire a stream of image frames of scene 160 and provide instances of the tracked object 170 to the second entity 114b. Further details of the first entity 114a are disclosed with reference to Figure 2. The second entity 114b may be an object classifier 114b. It is assumed that the object classifier 114b is pre-trained, and therefore hereafter referred to as the pre-trained object classifier 114b.System 110 may comprise further entities, functions, nodes, and devices, as represented by device 118. Examples of such further entities, functions, nodes, and devices include computing devices, communication devices, servers, sensors, and the like.

[0030] In some embodiments, the first entity 114a, the second entity 114b, and the database 116 define a video management system (VMS). Thus, the first entity 114a, the second entity 114b, and the database 116 are considered part of system 110. The first entity 114a, the second entity 114b, and the database 116 are operably connected to camera devices 120a, 120b via network 112. Network 112 may be wired, wireless, or partially wired and partially wireless. In some examples, system 110 includes a communication interface 520 (as shown in Figure 5) configured to communicate with a user interface 130 to display image frames. Reference numeral 140 indicates an exemplary connection between system 110 and the user interface 130. Connection 140 may be wired, wireless, or partially wired and partially wireless. User 180 may interact with the user interface 130. Since the user interface 130 is configured to display image frames captured by the camera devices 120a and 120b, it is understood to be at least partially a visual user interface.

[0031] Next, embodiments for retraining a pre-trained object classifier 114b are disclosed with parallel reference to Figures 2 and 3. Figure 2 is a block diagram of the system 110 according to the embodiment. Figure 3 is a flowchart of an embodiment of the method for retraining the pre-trained object classifier 114b. The method is performed by the system 110. The method is advantageously provided as a computer program 620 (see Figure 6).

[0032] S102: A stream of image frames from scene 160 is acquired. The stream of image frames can be acquired by the acquirer module 410. Each image frame depicts an instance of the tracked object 170. The tracked object 170 is the exact same object 170 that is tracked as it moves within scene 160. The stream of image frames can be acquired by the acquirer module 410 from at least one camera device 120a, 120b, as shown in Figure 2.

[0033] S104: Each instance of the tracked object 170 is classified to belong to an object class at a confidence level. Classification can be performed by the classifier module 420.

[0034] S106: It is verified that the confidence level of at least one instance of the tracked object 170, and of only one object class, is higher than the threshold confidence value. Verification may be performed by the verifier module 430. This ensures that at least one instance of the tracked object 170 is classified with high confidence into only one object class. Therefore, the confidence level of the instance of the tracked object 170 used as a criterion must be higher than the threshold confidence value of only one object class. This threshold confidence value may be adjusted upward or downward as needed to change the accuracy of retraining the pre-trained object classifier 114b. For example, accuracy can be increased by setting a very high threshold confidence value (e.g., 0.95 on a scale of 0.0 to 1.0). Furthermore, if it is observed that the amount of misrepresentation produced is higher than desired, the threshold confidence value can be adjusted upward to increase the accuracy of retraining the pre-trained object classifier 114b.

[0035] S118: All instances of the tracked object 170 in the stream of image frames are annotated with high confidence as belonging to a single object class (i.e., an object class in which at least one instance of the tracked object 170 has a confidence level higher than the threshold confidence value). This annotation results in annotated instances of the tracked object 170. The annotation may be performed by the annotator module 440.

[0036] If it can be successfully verified that the confidence level of at least one instance of the tracked object 170 is higher than the threshold confidence value, then one or more instances of the same object in other image frames that are only classified at a medium or low confidence level can be extracted and annotated as if they were classified at a high confidence level.

[0037] Therefore, in this respect, we proceed to S118 only if it can be successfully verified that the confidence level of at least one instance of the tracked object 170 is higher than the threshold confidence value in S106. As described above, S106 ensures that the tracked object 170 is classified with high confidence into only one object class before proceeding to S118.

[0038] S122: The pre-trained object classifier 114b is retrained on at least some of the annotated instances of the tracked object 170. Retraining may be performed by the retrainer module 460.

[0039] Therefore, the pre-trained object classifier 114b can be retrained with additional annotated instances of the tracked object 170. When this is done to the object classification model, the object classification model can therefore provide improved classification of objects over time in areas of scene 160 where classification was originally difficult, i.e., areas where classification was given low or moderate confidence. An example of this is provided below with reference to Figure 4.

[0040] Therefore, if it is verified that the tracked object 170 in S106 is classified with high confidence as belonging to only one object class in at least one image frame, the same tracked object 170 from other image frames will be used to retrain the pre-trained object classifier 114b. Similarly, if the tracked object 170 is not classified with high confidence as belonging to one object class, the tracked object 170 in other image frames will not be used to retrain the pre-trained object classifier 114b.

[0041] Next, embodiments relating to further details of the retraining of the pre-trained object classifier 114b performed by the system 110 are disclosed.

[0042] As shown in Figure 2, classification and retraining can be performed on separate devices. That is, in some embodiments, classification is performed on a first entity 114a, and retraining is performed on a second entity 114b that is physically separated from the first entity 114a. The second entity may be a (pre-trained) object classifier 114b.

[0043] Further verification may be performed before all instances of the tracked object 170 in the stream of image frames are annotated as belonging to an object class, and therefore before entering S118. Next, different embodiments related to this will be described.

[0044] Some of the tracked objects 170 may be classified to belong to two or more object classes, and each classification may have its own level of confidence. Thus, in some embodiments, it is verified that only one of these object classes has a high level of confidence. Thus, in some embodiments, at least some of the instances of the tracked objects 170 may also be classified to belong to further object classes with further levels of confidence, and the method S108: Further includes verifying that the level of confidence is lower than the threshold confidence value for at least some of the instances of the tracked object 170. Verification may be performed by the verifier module 430.

[0045] Therefore, if the tracked object 170 is classified with high confidence into two or more different object classes, the tracked object 170 will not be used to retrain the pre-trained object classifier 114b. The same applies if the tracked object 170 is not classified with high confidence into any object class, then the tracked object 170 will not be used for retraining.

[0046] In some embodiments, the path that the tracked object 170 takes as it moves from one image frame to the next is also tracked. If the tracked object 170 is classified with high confidence at least once to belong to only one object class, the path can be used to retrain the pre-trained object classifier 114b. Thus, in some embodiments, the tracked object 170 moves along a path in a stream of image frames, and the path is tracked as the tracked object 170 is tracked.

[0047] In some embodiments, it is verified that the path itself can be traced with high accuracy. That is, in some embodiments, the path is traced at the level of accuracy, and the method is S110: Further includes verifying that the level of accuracy is higher than the threshold accuracy value. Verification can be performed by the verifier module 430.

[0048] In some embodiments, it is verified that the path has neither split nor merged. If the path has split and / or merged at least once, this may indicate that the level of certainty the path has been tracked is below a threshold certainty value. This may be the case if it is suspected that the path has merged from two or more other paths, or split into two or more other paths. In such cases, the path will be determined to have low certainty and will not be used to retrain the pre-trained object classifier 114b. Therefore, in some embodiments, the method is S112: Further includes verifying that the path is not divided into at least two paths and is not merged from at least two paths in the stream of image frames. In other words, it is verified that the path is not divided or merged in any way and therefore constitutes a single path. Verification may be performed by the verifier module 430.

[0049] The same principle can also be applied if the object being tracked 170 is suspected to have odd size behavior that may be caused by shadowing, mirroring, or something similar. In that case, the object being tracked 170 is assumed to be classified with low confidence and will not be used to retrain the pre-trained object classifier 114b. Thus, in some embodiments, the object being tracked 170 has a size within the image frame, and the method is S114: Further includes verifying that the size of the tracked object 170 does not change beyond a threshold size value within the stream of image frames. Verification may be performed by the verifier module 430.

[0050] Since the apparent size of the tracked object 170 depends on the distance between the tracked object 170 and the camera devices 120a and 120b, compensation for the size of the tracked object 170 with respect to this distance can be performed as part of the verification in S114. In particular, in some embodiments, the size of the tracked object 170 is adjusted by a distance-dependent compensation coefficient determined as a function of the distance between the tracked object 170 and the camera devices 120a and 120b that captured the stream of image frames from scene 160, after verifying that the size of the tracked object 170 does not change beyond a threshold size value within the stream of image frames 210a:210c.

[0051] As shown in Figure 1 and disclosed with reference to Figure 1, scene 160 may be captured by one or more camera devices 120a, 120b. If scene 160 is captured by two or more camera devices 120a, 120b, then the stream of image frames in which the tracked object 170 resides may also be captured by two or more different camera devices 120a, 120b. Thus, in some embodiments, the stream of image frames consists of image frames captured by at least two camera devices 120a, 120b. If tracking of the tracked object 170 is performed locally on each of the camera devices 120a, 120b, this may require that information about the tracked object 170 be communicated between at least two camera devices 120a, 120b. In other examples, tracking of the tracked object 170 is performed centrally on image frames received from all of at least two camera devices 120a, 120b. The latter does not require any information about the tracked object 170 that is exchanged between at least two camera devices 120a, 120b.

[0052] In some embodiments, for example, if the stream of image frames originates from image frames captured by at least two camera devices 120a, 120b, in other embodiments as well, there is a risk that tracking of the tracked object 170 will be lost and / or that the classification of the tracked object 170 will change from image frame to image. Therefore, in some embodiments, the method is S116: Further includes verifying that the object class of the instance of the tracked object 170 does not change within the stream of image frames. Verification may be performed by the verifier module 430.

[0053] In some embodiments, to avoid training bias (e.g., machine learning bias, algorithmic bias, or artificial intelligence bias), the pre-trained object classifier 114b is not retrained using tracked objects 170 that have already been classified at a high confidence level. There are various ways to achieve this. The first method is to explicitly exclude tracked objects 170 that have already been classified at a high confidence level from retraining. In particular, in some embodiments, the pre-trained object classifier 114b is retrained only with annotated instances of tracked objects 170 whose confidence level has been verified not to be higher than a threshold confidence value. The second method is to assign a low weight to tracked objects 170 that have already been classified at a high confidence level during retraining. In this way, tracked objects 170 that have already been classified at a high confidence level can be implicitly excluded from retraining. In other words, in some embodiments, each annotated instance of the tracked object 170 is assigned a weight value to which the annotated instance of the tracked object 170 is weighted when the pre-trained object classifier 114b is retrained, and the weight value of an annotated instance of the tracked object 170 that has been verified to have a confidence level higher than a threshold confidence value is lower than the weight value of an annotated instance of the tracked object 170 that has been verified to have a confidence level lower than a threshold confidence value. This weight value makes the object classifier 114b less susceptible to the influence of tracked objects 170 that have already been classified at a high confidence level during retraining.

[0054] Aside from retraining the pre-trained object classifier 114b, there may be further different uses for annotated instances of the tracked object 170. In some embodiments, annotated instances of the tracked object 170 may be collected in a database 116 and / or provided to a further device 118. Thus, in some embodiments, the method is S120: Further includes the fact that an annotated instance of the tracked object 170 is provided to database 116 and / or further device 118. The annotated instance may be provided to database 116 and / or further device 118 by provider module 450.

[0055] This also allows other pre-trained object classifiers to benefit from annotated instances of the tracked object 170.

[0056] Next, refer to Figure 4, which schematically shows two streams 200, 200' of image frames 210a:210c. Each stream 200, 200' consists of a limited sequence of image frames. Each such stream may correspond to an captured video segment of a scene 160 in which one or more objects 170 are tracked. For illustrative purposes, each stream 200, 200' consists of three image frames 210a;210c in Figure 4. In particular, Figure 4 shows the first stream 200 of image frames 210a:210c, in which the tracked object 170 is classified with a confidence level higher than the threshold confidence value only within image frame 210c. In Figure 4, for illustrative purposes, it is assumed that the tracked object 170 is classified with a confidence level higher than the threshold confidence value so that it belongs to the object class "walking man" only within image frame 210c. This is because, in image frames 210a and 210b, the lower half of the tracked object 170 is obscured by vegetation, making it difficult for the object classification unit to determine whether or not a person is walking. Figure 4 also shows a second stream 200' of the same image frames 210a:210c, but after the application of the embodiments disclosed herein, as indicated by the arrow labeled “Annotation”. By applying the embodiments disclosed herein, instances of the tracked image 170 in image frames 210a and 210b are also annotated as belonging to the same object class as the instance of the tracked image 170 in image frame 210c. This is indicated by arrows 220a and 220b in Figure 4. Thus, instances of the tracked object 170 in image frames 210a and 210b are also annotated as belonging to the object class “walking man” at a confidence level higher than the threshold confidence value, despite the presence of partially obscuring vegetation. In other words, instances of the tracked object 170 in image frames 210a and 210b inherit the annotations of the instance of the tracked object 170 in image frame 210c.Next, the pre-trained object classifier 114b can be retrained using instances of the tracked object 170 in image frames 210a and 210b. Thus, for this method to work, it is only necessary that the classification of object 170 in one image frame 210c of stream 200 is above the threshold (as shown in Figure 4 as "Confidence: High"). As a result of performing the method, object 170 will also be annotated in image frames 210a and 210b in stream 200'. Next, since object 170 is also annotated in image frames 210a and 210b in stream 200', this means that retraining the pre-trained object classifier 114b on stream 200' will actually improve the pre-trained object classifier 114b. More precisely, when a pre-trained object classifier 114b is run on a new stream of image frames containing partially obscured objects belonging to the object class "walking man," such objects can also be tracked and annotated as being classified at a high level of confidence.

[0057] Figure 5 schematically shows the components of system 110 according to one embodiment in terms of several functional units. The processing circuit 510 is provided using one or more of any combination of suitable central processing units (CPUs), multiprocessors, microcontrollers, digital signal processors (DSPs), etc., which can execute software instructions stored in a computer program product 610 (as shown in Figure 6) in the form of a storage medium 530. The processing circuit 510 may be further provided as at least one application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA).

[0058] In particular, the processing circuit 510 is configured to cause the system 110 to execute a set of operations or steps, as described above. For example, the storage medium 530 may store a set of operations, and the processing circuit 510 may be configured to retrieve a set of operations from the storage medium 530 in order to cause the system 110 to execute the set of operations. The set of operations may be provided as a set of executable instructions.

[0059] Accordingly, the processing circuit 510 is arranged to perform the methods disclosed herein. The storage medium 530 may also comprise a persistent storage device which can be any single or combination of, for example, magnetic memory, optical memory, solid-state memory, or even remotely mounted memory. The system 110 may further comprise a communication interface 520 configured to communicate with further devices, functions, nodes, and other devices. Thus, the communication interface 520 may comprise one or more transmitters and receivers comprising analog and digital components. The processing circuit 510 controls the general operation of the system 110, for example, by transmitting data and control signals to the communication interface 520 and the storage medium 530, by receiving data and reports from the communication interface 520, and by retrieving data and instructions from the storage medium 530. Other components and associated functions of the system 110 are omitted in order not to obscure the concepts presented herein.

[0060] Figure 6 shows an example of a computer program product 610 comprising a computer-readable storage medium 630. This computer-readable storage medium 630 can store a computer program 620, which can cause a processing circuit 210 and entities and devices operably coupled thereto, such as a communication interface 220 and the storage medium 230, to perform the methods according to the embodiments described herein. Thus, the computer program 620 and / or the computer program product 610 can provide means for performing any of the steps disclosed herein.

[0061] In the example in Figure 6, the computer program product 610 is shown as an optical disc such as a CD (Compact Disc), DVD (Digital Multipurpose Disc), or Blu-ray disc. The computer program product 610 can also be embodied as a memory such as Random Access Memory (RAM), Read-Only Memory (ROM), Erasable Programmable Read-Only Memory (EPROM), or Electrically Erasable Programmable Read-Only Memory (EEPROM), and more specifically as a non-volatile storage medium in a device within external memory such as USB (Universal Serial Bus) memory or flash memory such as CompactFlash memory. Thus, although the computer program 620 is schematically shown as a track on the optical disc depicted herein, the computer program 620 can be stored in any way suitable for the computer program product 610.

[0062] The concept of the present invention has been described above primarily with reference to several embodiments. However, as will be readily apparent to those skilled in the art, other embodiments not disclosed above are equally possible within the scope of the concept of the invention, as defined by the appended claims.

Claims

1. A method for retraining a pre-trained object classifier (114b), the method being executed by a system (110), the system (110) comprising a processing circuit (510), the method comprising: Obtaining (S102) a stream (200, 200') of image frames (210a: 210c) of a scene (160), each of the image frames (210a: 210c) depicting an instance of a tracked object (170), the tracked object (170) being the same object (170) tracked as it moves within the scene (160), obtaining (S102) the stream (200, 200'); Classifying (S104) each instance of the tracked object (170) at a level of confidence as belonging to an object class; Verifying (S106) that the level of confidence of at least one of the instances of the tracked object (170) and of only one object class is higher than a threshold confidence value, thereby ensuring that at least one of the instances of the tracked object (170) is classified with high confidence into the only one object class, verifying (S106); and if it is verified that it is higher than the threshold confidence value, Annotating (S118) all instances of the tracked object (170) in the stream (200, 200') of the image frames (210a: 210c) as belonging with high confidence to the only one object class, resulting in annotated instances of the tracked object (170); Retraining (S122) the pre-trained object classifier (114b) using at least some of the annotated instances of the tracked object (170); A method, comprising.

2. At least some of the instances of the tracked object (170) are classified as also belonging to a further object class at a further level of confidence, the method comprising: Verifying (S108) that for at least some of the instances of the tracked object (170), the further level of confidence is lower than the threshold confidence value The method according to claim 1, further comprising

3. The method being verifying that the object class of the instance of the tracked object (170) does not change within the stream (200, 200') of the image frames (210a: 210c) (S116) The method according to claim 1 or 2, further comprising

4. The tracked object (170) moves along a path within the stream (200, 200') of the image frames (210a: 210c), the method according to claim 1 or 2, wherein the path is tracked when the tracked object (170) is tracked.

5. The path being tracked at a level of accuracy, the method being verifying that the level of accuracy is higher than a threshold accuracy value (S110) The method according to claim 4, further comprising

6. The method being verifying that the path is not divided into at least two paths and that there is no confluence from at least two paths within the stream (200, 200') of the image frames (210a: 210c) (S112) The method according to claim 4, further comprising

7. The tracked object (170) having a size within the image frame (210a: 210c), the method being verifying that the size of the tracked object (170) does not change beyond a threshold size value within the stream (200, 200') of the image frames (210a: 210c) (S114) The method according to claim 1 or 2, further comprising

8. The size of the tracked object (170) is adjusted by a distance-dependent compensation factor determined as a function of the distance between the tracked object (170) and the camera device (120a, 120b) that captured the stream (200, 200') of the image frames (210a: 210c) of the scene (160) when it is verified that the size of the tracked object (170) does not change beyond the threshold size value within the stream (200, 200') of the image frames (210a: 210c) of the tracked object (170). The method according to claim 7.

9. The method according to claim 1 or 2, wherein the pre-trained object classifier (114b) is retrained only with the annotated instances of the tracked object (170) for which it has been verified that the level of confidence is not higher than the threshold confidence value.

10. The method according to claim 1 or 2, wherein each of the annotated instances of the tracked object (170) is assigned a respective weighting value according to which the annotated instance of the tracked object (170) is weighted when the pre-trained object classifier (114b) is retrained, and the weighting value of the annotated instance of the tracked object (170) for which it has been verified that the level of confidence is higher than the threshold confidence value is lower than the weighting value of the annotated instance of the tracked object (170) for which it has been verified that the level of confidence is not higher than the threshold confidence value.

11. The method further includes providing the annotated instances of the tracked object (170) to a database (116) and / or a further device (118) (S120) The method according to claim 1 or 2.

12. The method according to claim 1 or 2, wherein the stream (200, 200') of the image frames (210a: 210c) is derived from the image frames (210a: 210c) captured by at least two camera devices (120a, 120b).

13. The method according to claim 1 or 2, wherein the classifying is performed by a first entity (114a) and the retraining is performed by a second entity (114b) physically separated from the first entity (114a).

14. A system (110) for retraining a pre-trained object classifier (114b), the system (110) comprising a processing circuit (510), the processing circuit causing the system (110) to Obtaining a stream (200, 200') of image frames (210a:210c) of a scene (160), each of the image frames (210a:210c) depicting an instance of a tracked object (170), and the tracked object (170) being the same object (170) that is tracked as it moves within the scene (160), obtaining the stream (200, 200'); Classifying each instance of the tracked object (170) at a level of confidence so as to belong to an object class; Verifying that at least one of the instances of the tracked object (170) and the level of confidence of exactly one object class is higher than a threshold confidence value, thereby ensuring that at least one of the instances of the tracked object (170) is classified with high confidence to the exactly one object class. When it is verified that it is higher than the threshold confidence value; Annotating all instances of the tracked object (170) in the stream (200, 200') of the image frames (210a:210c) as belonging with high confidence to the exactly one object class, resulting in annotated instances of the tracked object (170); Retraining the pre-trained object classifier (114b) using at least some of the annotated instances of the tracked object (170); A system (110) configured to cause.

15. A non-transitory computer-readable storage medium (630) storing a computer program (620) for retraining a pre-trained object classifier (114b), the computer program including computer code, and when the computer code is executed on a processing circuit (510) of a system (110), the computer code causes the system (110) to Obtaining (S102) a stream (200, 200') of image frames (210a: 210c) of a scene (160), wherein each of the image frames (210a: 210c) depicts an instance of a tracked object (170), and the tracked object (170) is the same object (170) that is tracked when moving within the scene (160), obtaining (S102) the stream (200, 200') Classifying (S104) each instance of the tracked object (170) at a level of confidence so as to belong to an object class Verifying (S106) that the level of confidence of at least one of the instances of the tracked object (170) and of only one object class is higher than a threshold confidence value, thereby ensuring that at least one of the instances of the tracked object (170) is classified with high confidence to the only one object class, verifying (S106); and when it is verified that it is higher than the threshold confidence value Annotating (S118) all instances of the tracked object (170) within the stream (200, 200') of the image frames (210a: 210c) as belonging with high confidence to the only one object class, resulting in annotated instances of the tracked object (170) Retraining (S122) the pre-trained object classifier (114b) using at least some of the annotated instances of the tracked object (170) A non-transitory computer-readable storage medium (630) causing the above