Object masking in video stream
Patent Information
- Application Number
- JP2022112221
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-07-19
- Filing Date
- 2022-07-13
- Publication Date
- 2025-07-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing camera monitoring systems struggle to reliably mask moving objects, particularly vehicles at high speeds, in video streams, leading to potential failures in object detection and classification.
A method that segments video streams into foreground and background segments, using different classifications and thresholds based on object movement to accurately mask objects, with lower complexity and higher speed for foreground objects and higher accuracy for background objects.
Enhances the reliability and efficiency of object masking in video streams by reducing misclassification and ensuring privacy protection, even for fast-moving objects.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of camera surveillance, and more particularly to a method and a control unit for masking an object in a video stream.
Background Art
[0002] In various camera surveillance applications, there is sometimes a need to mask an object in a video stream captured by a camera. Some important reasons for masking an object are to protect the privacy of people appearing in the video stream and to protect other types of personal information that may be captured in the video stream.
[0003] As an example, object detection may be used to detect vehicles. Masking the vehicle, or perhaps more importantly, the vehicle's license plate or the people inside the vehicle, may be done by extracting the image coordinates of the relevant parts of the object. Once the image coordinates are known, the relevant parts can be masked or pixelated in the video stream.
[0004] However, moving objects such as vehicles can be difficult to detect with high reliability, especially at high speeds, such as on a highway. In such cases, there is a risk of failure in object detection and classification respectively, and as a result, there is a risk that the relevant object is not masked.
[0005] As a result, there is room for improvement with respect to masking an object in a video stream.
Summary of the Invention
[0006] In view of the above and other drawbacks of the prior art, an object of the present invention is to provide an improved method for masking an object in a video stream that alleviates at least some of those drawbacks of the prior art.
[0007] According to a first aspect of the present invention, a method for masking an object in a video stream is provided.
[0008] This method includes the following steps: acquiring a video stream; detecting objects in the video stream; determining whether the detected objects belong to the foreground of the video stream, which indicates moving objects, or the background of the video stream, which indicates stationary objects; classifying the detected objects as a specific type using a first classifier if they belong to the foreground, and using a second classifier, different from the first classifier, if they belong to the background; and masking the detected objects in the video stream if they belong to a specific type.
[0009] The present invention is based on the realization of using different classifiers depending on whether an object belongs to the foreground or background of a video stream. The foreground is a portion or segment of the video stream containing moving objects, and the background is a portion or segment of the video stream containing stationary objects. By providing two different classifiers, it is possible to tailor the classifiers to different segments, i.e., different types of features of objects in the background or foreground. More precisely, the first classifier may be specifically configured to classify objects in the foreground, and the second classifier may be specifically configured to classify objects in the background. Thus, these classifiers do not need to be efficient or accurate in the background and the other in the foreground.
[0010] This method may include the step of segmenting the video stream into a background and a foreground, i.e., a background segment and a foreground segment. Such segmentation may be performed by extracting pixel data from pixels in each frame of the video stream showing moving objects, and this data may be used to construct the foreground. Pixel data from pixels in each frame of the video stream showing stationary objects is used to construct the background.
[0011] Following segmentation, it may be determined whether the object belongs to a background segment or a foreground segment, that is, whether the object was found or detected in a background segment or a foreground segment.
[0012] Therefore, determining whether an object belongs to the background or the foreground may involve segmenting the video stream into background and foreground, and subsequently determining whether the object belongs to the background or the foreground.
[0013] Masking an object in a video stream involves pixelating the pixels that are to be masked, or in other words, covering them or filling them in black.
[0014] In one embodiment, the computational complexity of the first classifier is lower than that of the second classifier.
[0015] This preferably provides a first classifier that is faster than a second classifier. The faster classifier may be less accurate than the slower classifier, however, for detecting fast-moving objects in the foreground, a faster classifier is preferable even if it is somewhat less accurate. In particular, for privacy purposes, a certain amount of misclassification is acceptable. If these classifiers are neural networks, the size of the neural network of the first classifier may be smaller than the size of the neural network of the second classifier.
[0016] In one embodiment, the second classifier may be configured to perform classification for a smaller number of frames per unit of time in the video stream than the first classifier. The second classifier is used to classify objects in the background segment of the video stream where stationary objects are detected, and therefore a high frame rate is not required. In other words, stationary objects do not move, and therefore a lower frame rate is sufficient to detect objects in the background. Conversely, a higher frame rate is required in the foreground segment where moving objects are expected. Thus, the second classifier can perform slower but more accurate classification, while the first classifier can perform faster but less accurate classification. This is one preferred way to tailor these classifiers to those specific tasks of object classification in different segments of the same video stream.
[0017] In one embodiment, the second classifier may be configured to classify only discontinuous timeframes of the video stream. Preferably, it is not necessary to classify every frame, as the second classifier is expected to classify only stationary objects. To enable the use of a more accurate classifier that requires processing time, preferably, only discontinuous timeframes occurring at regular time intervals in the video stream may be classified by the second classifier. Such discontinuous timeframes may be, for example, every 5th, 10th, 15th, 20th, or 25th timeframe.
[0018] In one embodiment, the first classifier may be configured to classify moving objects, and the second classifier may be configured to classify stationary objects. Thus, the first classifier may be trained using only video streams containing moving objects, and the second classifier may be trained using video streams containing only stationary objects.
[0019] Neural networks provide an efficient tool for classification. Various types of neural networks adapted for classification are known. Suitable examples of neural networks include recurrent neural networks and convolutional neural networks. Recurrent neural networks are particularly efficient at capturing temporal evolution.
[0020] Furthermore, other suitable classifiers may be decision tree classifiers such as random forest classifiers, which are efficient for classification. In addition, classifiers such as support vector machine classifiers and logistic regression classifiers are also possible.
[0021] Furthermore, the classifier may be a statistical classifier, a heuristic classifier, or a fuzzy logic classifier. Additionally, the use of tables, i.e., lookup tables containing combinations of data, is also feasible.
[0022] A second aspect of the present invention provides a method for masking an object in a video stream. This method includes the steps of: acquiring a video stream; detecting an object in the video stream; determining whether the detected object belongs to the foreground of the video stream, which indicates a moving object, or to the background of the video stream, which indicates a stationary object; if the detected object is determined to belong to the foreground, classifying the detected object as a specific type using a lower classification threshold than that used for the object belonging to the background; and if the detected object is classified as a specific type of object, masking the object in the video stream.
[0023] This second aspect of the present invention is based on the implementation of using different thresholds in the classifier depending on whether an object belongs to the foreground or background of a video stream. The foreground is the portion or segment of the video stream that contains moving objects, and the background is the portion or segment of the video stream that contains stationary objects. Providing a lower threshold for what is acceptable as a particular type of object reduces the amount of certain types of objects that are not masked, even if the amount of misclassification increases. For scenes where privacy is important, this is not an issue. More importantly, all objects of a particular type are masked. For background segments, this threshold is kept higher because the background area of stationary objects is easy to classify within this range. Lowering the classification threshold in the foreground segment increases the likelihood of correctly classifying all objects of a particular type, even if they are moving quickly.
[0024] This method may include the step of segmenting the video stream into a background and a foreground, i.e., a background segment and a foreground segment. Such segmentation may be performed by extracting pixel data from pixels in each frame of the video stream showing moving objects, and this data may be used to construct the foreground. Pixel data from pixels in each frame of the video stream showing stationary objects is used to construct the background.
[0025] Following segmentation, it may be determined whether the object belongs to a background segment or a foreground segment, that is, whether the object was found or detected in a background segment or a foreground segment.
[0026] In each of the embodiments, this method may include determining the speed of the detected object and selecting a classification threshold according to the speed of the detected object when determining that the detected object belongs to the foreground, and the detected object is classified using the selected classification threshold. In this way, the threshold can be adapted to the speed of the object so that an appropriate threshold is used. The threshold should be selected to ensure that a particular type of object is correctly classified while keeping the level of classification omission, i.e., detection omission, low, and that the particular type of object is detected and classified even when the object is moving very fast.
[0027] In one exemplary embodiment, the threshold may be selected according to the following procedure. When the speed of the object exceeds the speed threshold, the object is classified using a classifier with a first classification threshold, and when the speed of the object is less than the speed threshold, the object is classified using the classifier with a second classification threshold higher than the first classification threshold. This preferably provides an easy-to-understand speed threshold for selecting the classification threshold. The speed threshold may be fixed.
[0028] In one possible embodiment, a speed exceeding the speed threshold indicates that the object is moving, and a speed less than the speed threshold indicates a stationary object. In other words, the speed threshold may be zero.
[0029] In some embodiments, the classification threshold is according to the speed of the object. In other words, instead of a fixed threshold, the threshold is selected as being according to the speed of the object, for example, as a sliding threshold set adaptively. This preferably provides a more accurately set threshold, and as a result, preferably provides an improved classification result.
[0030] In each of the embodiments, the method may include classifying an object into a foreground set of frames and a background set of frames, and masking each of the objects classified as being a particular type of object. In other words, it should be ensured that all of the classified, particular type of objects in both the foreground and the background are masked to protect privacy.
[0031] In some possible implementations, the particular type of object may be a vehicle.
[0032] According to a third aspect of the present invention, there is provided a control unit configured to perform any one of the steps of the aspects and embodiments described herein.
[0033] Each further embodiment of this third aspect of the present invention and the effects obtained through them are very similar to those described above for the first and second aspects of the present invention.
[0034] According to a fourth aspect of the present invention, there is provided a system including an imaging device configured to capture a video stream and a control unit according to the third aspect.
[0035] Each further embodiment of this fourth aspect of the present invention and the effects obtained through them are very similar to those described above for the first, second, and third aspects of the present invention.
[0036] According to a fifth aspect of the present invention, there is provided a computer program including instructions that, when executed by a computer, cause the computer to execute any one of the methods of the embodiments described herein.
[0037] Each further embodiment of this fifth aspect of the present invention and the effects obtained through them are very similar to those described above for the other aspects of the present invention.
[0038] Further features of the present invention and the advantages thereof will become apparent when considering the appended claims and the following description. A person skilled in the art will recognize that various features of the present invention can be combined to create other embodiments described below, each without departing from the scope of the invention.
[0039] Various aspects of the present invention, including their specific features and advantages, will be readily apparent from the following embodiments for carrying out the invention and the accompanying drawings shown below. [Brief explanation of the drawing]
[0040] [Figure 1] As one application example of each embodiment of the present invention, a scene monitored by an imaging device is conceptually shown. [Figure 2] The segmentation into background and foreground segments according to each embodiment of the present invention is conceptually shown. [Figure 3] This is a flowchart of the method steps according to each embodiment of the present invention. [Figure 4] This conceptually illustrates the stream of frames in a video stream. [Figure 5] This is a flowchart of the method steps according to each embodiment of the present invention. [Figure 6] This is a flowchart of the method steps according to each embodiment of the present invention. [Figure 7] This is a flowchart of the method steps according to each embodiment of the present invention. [Figure 8] This is a block diagram of the system according to each embodiment of the present invention. [Modes for carrying out the invention]
[0041] The present invention will be described in further detail below with reference to the accompanying drawings. Herein are some currently preferred embodiments of the invention. However, the invention may be embodied in many different forms and should not be understood as being limited to the embodiments shown below. Rather, these embodiments are provided for the sake of perfection and completeness, and to fully convey the scope of the invention to those skilled in the art. Similar reference numerals indicate similar components throughout these drawings.
[0042] These drawings, particularly Figure 1, show a scene 1 being monitored by an imaging device 100, such as a camera or, more specifically, a surveillance camera. In scene 1, there is a moving object 102 and a set of stationary objects 104a, 104b, and 104c. The moving object 102 may be a vehicle driving on a road 106, and at least one of the stationary objects, for example 104a, is a vehicle parked next to the road 106.
[0043] The imaging device 100 continuously monitors scene 1 by capturing a video stream of the scene and the objects within it. It is desirable to mask certain objects, such as the vehicle 104a and especially their license plates, so that they are not visible in the output video stream. To this end, it is necessary to detect objects in the scene, classify them as belonging to a specific type, such as a vehicle, and then mask them.
[0044] Detecting and classifying stationary objects presents different challenges compared to classifying moving objects. High-speed moving objects 102 can be difficult to detect and classify accurately because they are only present in Scene 1 for a limited time. To mitigate this problem, the inventors propose using different properties for classification depending on whether the background of stationary objects or the foreground of moving objects is considered.
[0045] Moving on to Figure 2, Scene 1 is represented by an example where two or more frames 202 have been accumulated showing a moving object 102 in motion. Each frame of the video stream is divided or segmented into a foreground 204 showing a moving object such as a vehicle 102, and a background 206 showing stationary objects 104a to c.
[0046] Foreground and background segmentation can be performed, for example, by analyzing the captured scenes in a video stream. If motion is detected between two frames, the relevant region of the frame is tagged as the foreground. Regions of the frame where no motion is detected are tagged as the background. Segmentation can be performed in various ways and is generally based on many image features, such as color, intensity, and region-based segmentation. A few examples are briefly described below.
[0047] For example, segmentation can be based on motion vectors. In this case, areas in a video stream frame with motion vectors larger than a certain size become foreground segments, and the remaining areas become background segments.
[0048] Segmentation may be based on differences in pixel values. For example, differences in pixel intensity and color may be used for segmentation.
[0049] Figure 3 is a flowchart of method steps according to a first embodiment of the present invention for detecting, classifying, and masking objects in a video stream.
[0050] In step S102, a video stream is acquired. This may be done by the imaging device 100. In other words, the imaging device acquires a video stream of scene 1, which includes the moving object 102 and the stationary objects 104a to c.
[0051] In step S104, the moving object 102 and objects 104a through c are detected in the video stream. This detection may be performed before or after segmenting or dividing the video stream 202 into background 206 and foreground 204.
[0052] In step S106, it is determined whether the detected object belongs to the foreground 204 of the video stream 202 showing the moving object 102 or to the background 206 of the video stream 202 showing the stationary objects 104a to c. Such determination may be made by determining whether the object is moving or not. The moving object 102 generally belongs to the foreground, while the stationary objects 104a to c belong to the background.
[0053] Furthermore, the video stream may be segmented into background and foreground. Here, after segmentation, it is determined whether an object belongs to the background or the foreground. Thus, the video stream may be divided into background 206 and foreground 204, and it can be concluded whether an object belongs to background 206 or foreground 204 depending on which segment the object is detected in.
[0054] In steps S108a to S108b, if the detected object is determined to belong to the foreground, the first classifier is used in step S108a to classify the detected object as belonging to a specific type. If the detected object is determined to belong to the background, the second classifier is used in step S108b to classify the detected object as belonging to a specific type. The first classifier is different from the second classifier.
[0055] A specific type may be defined as an object being a vehicle such as a car, truck, boat, or motorcycle. Additionally, or alternatively, a specific type may be defined as an object being a person.
[0056] The differences between classifiers may be different properties, but the main principle is that the first classifier should classify moving objects, and the second classifier should classify stationary objects.
[0057] In step S110, if an object detected in either the foreground or background is classified as a specific type of object, this object is masked in the video stream. Therefore, in the output video stream, objects classified as a specific type are masked. For example, if object 104a is classified as a vehicle by the second classifier, vehicle 104a is masked in step S110, and if a moving object 102 is classified as a vehicle by the first classifier, vehicle 102 is masked in step S110.
[0058] The first classifier differs from the second classifier. For example, the computational complexity of the first classifier is lower than that of the second classifier. This offers the advantage that the first classifier may process data from the video stream faster and can classify moving objects, at the cost of less accuracy.
[0059] A video stream is a stream of frames, as is common in video processing. Figure 4 conceptually illustrates a video stream as a set of 400 conceptual frames. To take advantage of the different requirements of these classifiers, a second classifier may be configured to classify a smaller number of frames per unit of time in the video stream than the first classifier. For example, the second classifier may be configured to classify only discontinuous timeframes 402 of the video stream. In other words, objects in the background 206 are classified only in discontinuous timeframes 402, while objects in the foreground are classified using the full set of frames 400. Discontinuous timeframes 402 occur at fixed time intervals divided by a fixed number of time minutes or a fixed number of intermediate timeframes, for example, discontinuous timeframes 402 are the 5th, 10th, 15th, 20th, or 25th timeframes.
[0060] The second classifier is expected to classify only stationary objects, so there is no need for temporally dense data, such as a full set of 400 frames. In contrast, the first classifier is expected to classify moving objects, so it is preferable to use temporally dense video data, such as a full set of 400 frames, rather than for the second classifier, which only requires discontinuous timeframes 402.
[0061] Furthermore, the first classifier may be configured to specifically classify moving objects, and the second classifier may be configured to specifically classify stationary objects. For example, the first classifier may be trained using image data of moving objects, while the second classifier may be trained using image data of stationary objects.
[0062] Figure 5 is a flowchart of method steps according to a second embodiment of the present invention for detecting, classifying, and masking objects in a video stream. This second embodiment of the method is based on the same insights as that of the first embodiment, which is described with reference to Figure 3, and therefore belongs to the same concept relating to the present invention.
[0063] In step S102, a video stream is acquired. This may be done by the imaging device 100. In other words, the imaging device acquires a video stream of scene 1, which includes the moving object 102 and the stationary objects 104a to c. Step S102 in Figure 5 is performed in the same way as step S102 in Figure 3.
[0064] In step S104, the moving object 102 and objects 104a through c are detected in the video stream. This detection may be performed before or after segmenting or dividing the video stream 202 into background 206 and foreground 204.
[0065] In step S106, it is determined whether the detected object belongs to the foreground 204 of the video stream 202 showing the moving object 102 or to the background 206 of the video stream 202 showing the stationary objects 104a to c. Such determination may be made by determining whether the object is moving or not. Moving objects generally belong to the foreground, and stationary objects belong to the background. Step S106 in Figure 5 is performed in the same way as step S106 in Figure 3.
[0066] In steps S208a to S208b, if the detected object in step S208a is determined to belong to the foreground, then in step S208b, a second classification threshold is used. If the object is determined to belong to the background, a lower first classification threshold is used to classify the detected object as belonging to a specific type. The first classification threshold is lower than the second classification threshold.
[0067] In step S110, if an object detected in either the foreground or background is classified as a specific type of object, this object is masked in the video stream. Therefore, in the output video stream, objects classified as a specific type are masked. For example, if object 104a is classified as a vehicle using a second classification threshold, vehicle 104a is masked in step S110, and if moving object 102 is classified as a vehicle using a first classification threshold, vehicle 102 is masked in step S110. Step S110 in Figure 5 is performed in the same way as step S110 in Figure 3.
[0068] Moving objects 102 are more difficult to detect and classify, or require higher computational resources. Therefore, the first classification threshold for detected objects of a specific type in the foreground 204 is set relatively low. Stationary objects 104a to c are easy to detect and classify. Therefore, if it is known that the objects are stationary, a relatively high classification threshold can be used to determine if the detected objects belong to a specific type. Based on this insight, the first classification threshold used for the foreground 204 is lower than the second classification threshold used for the background 206 of the video stream.
[0069] Referring to Figure 6, if the detected object is determined to belong to the foreground in step S106, the velocity of the detected object is determined in step S204. The velocity of the object may be determined by analyzing consecutive frames in the video stream or by other known means in the prior art.
[0070] In step S206, a classification threshold is selected according to the velocity of the detected object. The detected object is classified using the classification threshold selected in step S208a, as explained in relation to Figure 5.
[0071] Similarly, in the case of different classifiers, as explained with reference to Figure 3 and here with reference to Figure 7, if the detected object is determined to belong to the foreground in step S106, the velocity of the detected object is determined in step S204.
[0072] The velocity of an object may be determined by analyzing consecutive frames in a video stream, or by other known means in the prior art.
[0073] In step S206, a classification threshold is selected according to the velocity of the detected object. The detected object is then classified in step S108a using the selected classification threshold and the first classifier, as explained in relation to Figure 3.
[0074] The classification threshold may be set and adjusted according to the application currently being targeted. For example, a fixed threshold may be used. Therefore, if the velocity of a moving object 102 exceeds the velocity threshold, the object is classified using a classifier with a first classification threshold. However, if the velocity of a moving object 102 is less than the velocity threshold, the object is classified using a classifier with a second classification threshold that is higher than the first classification threshold.
[0075] In some possible implementations, a velocity above the velocity threshold indicates that object 102 is moving, while a velocity below the velocity threshold indicates that objects 104a through c are stationary.
[0076] The classification threshold may be a sliding threshold, which is dependent on the object's velocity. Therefore, a predetermined function may be used. Here, the object's velocity is used as input, and its output is the selected classification threshold. For example, a change in the object's velocity may result in a proportional change in the selected classification threshold.
[0077] Any object classified as a specific type, such as a vehicle with a license plate, will be masked, whether it is within the frame of the foreground set or the background set.
[0078] The classifiers described herein may perform operations on different types of classifiers. For example, a classifier neural network adapted to perform the classification step may be used. Various types of neural networks adapted to perform classification are conceivable and known. An example of a suitable neural network is a convolutional neural network, which may or may not have recursive features. Other suitable classifiers may be decision tree classifiers such as random forest classifiers. Furthermore, classifiers such as support vector machine classifiers, logistic regression classifiers, heuristic classifiers, fuzzy logic classifiers, or statistical classifiers, or lookup tables, may also be used in the classifier.
[0079] The classifier provides output indicating the result of the classification step, such as whether the object is of a particular type, which may be provided to some extent by the classification threshold.
[0080] The classification threshold may be adjustable for a particular classifier. The classification threshold may be an experimentally determined threshold.
[0081] Figure 8 is a block diagram of a system 800 according to each embodiment of the present invention. The system 800 includes an imaging device 100 configured to capture a video stream and a control unit 802 configured to perform one of the methods described in relation to Figures 3 to 7. The output, which is a video stream, is masked if a particular type of object is present.
[0082] The control unit includes a microprocessor, microcontroller, programmable digital signal processor, or other programmable device. The control unit also, or alternatively, includes an application-specific integrated circuit, programmable gate array or programmable array logic, programmable logic device, or digital signal processor. The control unit includes programmable devices such as the aforementioned microprocessor, microcontroller, or programmable digital signal processor. The processor may further include computer-executable code that controls the operation of the programmable device.
[0083] The control functions of the Disclosure may be implemented using an existing computer processor, or by a dedicated computer processor incorporated for this or another purpose for an appropriate system, or by a hardwired system. Embodiments within the scope of the Disclosure include program products including machine-readable media for executing or having machine-executable instructions or data structures stored therein. Such machine-readable media may be any available media accessible by a general-purpose or dedicated computer or other machine having a processor. For illustrative purposes, such machine-readable media may include RAM, ROM, EPROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other media accessible by a general-purpose or dedicated computer or other machine having a processor that can be used to execute or store desired program code in the form of machine-executable instructions or data structures. When information is transferred to or provided to a machine via a network or another communication connection (either hardwired, wireless, or a combination of hardwired and wireless), the machine correctly views that connection as machine-readable media. Thus, any such connection is correctly referred to as machine-readable media. The above combinations also fall within the scope of machine-readable media. Machine-executable instructions include, for example, instructions and data that cause a general-purpose computer, a dedicated computer, or a dedicated processing machine to perform a specific function or group of functions.
[0084] While drawings may show sequences, the order of steps may differ from those depicted. Furthermore, two or more steps may be performed simultaneously, or partially simultaneously. Such variations will depend on the selected software and hardware systems, as well as the designer's choices. All such variations are within the scope of this disclosure. Similarly, software implementations can be achieved using standard programming techniques involving rule-based logic and other logic to accomplish various connection, processing, comparison, and decision steps. Additionally, while the present invention has been described with reference to its specifically illustrated embodiments, many different changes and modifications will be apparent to those skilled in the art.
[0085] Furthermore, variations to the disclosed embodiments can be understood and achieved by an addressee skilled in the art in the practice of the patented invention by examining the drawings, the disclosure, and the appended claims. Furthermore, in the claims, the term “comprising” does not exclude other elements or steps. The indefinite articles “a” or “an” do not exclude the plural.
Claims
1. A method for masking an object in a video stream, comprising: obtaining a video stream; detecting an object in the video stream; determining that one detected object belongs to the foreground of the video stream indicating a moving object and that one detected object belongs to the background of the video stream indicating a stationary object; classifying the detected object belonging to the foreground as being of a specific type using a first classifier configured to classify moving objects of a specific type, and classifying the detected object belonging to the background as being of the specific type using a second classifier configured to classify stationary objects of the same specific type as the first classifier; determining that the detected objects in the background and the foreground are classified as being of the specific type; masking only the objects classified as being of the specific type in the foreground and the background of the video stream A method comprising the above steps.
2. The method according to claim 1, wherein the computational complexity of the first classifier is lower than that of the second classifier, and the first classifier is configured to process data from the video stream at a faster speed with lower accuracy than the second classifier.
3. The method according to claim 1, wherein the second classifier is configured to perform classification on a smaller number of frames per time unit of the video stream than the first classifier.
4. The method according to claim 1, wherein the second classifier is configured to perform classification only on non-consecutive time frames of the video stream.
5. When it is determined that the detected object belongs to the foreground, determining the speed of the detected object, and selecting a classification threshold according to the speed of the detected object, wherein the detected object is classified using the selected classification threshold, and when the speed of the object exceeds a speed threshold, classifying the object using a classifier having a first classification threshold, and when the speed of the object is less than the speed threshold, classifying the object using a classifier having a second classification threshold higher than the first classification threshold. The method according to claim 1. **Claim 6**: The method according to claim 5, wherein a speed exceeding the speed threshold indicates that the object is moving, and a speed less than the speed threshold indicates a stationary object. **Claim 7**: The method according to claim 5, wherein the classification threshold is a function of the speed of the object. **Claim 8**: The method according to claim 1, comprising classifying objects in the set of foreground frames and the set of background frames, and masking each object classified as being of the specific type of object. **Claim 9**: The method according to claim 1, comprising segmenting the video stream into a background and a foreground, and subsequently determining whether the object belongs to the background or the foreground following the segmentation. **Claim 10**: The method according to claim 1, wherein the specific type of object is a vehicle. **Claim 11** A method for masking an object in a video stream, comprising: obtaining a video stream; detecting an object in the video stream; determining that one detected object belongs to a first portion of the video stream indicating a moving object and that one detected object belongs to a second portion of the video stream indicating a stationary object; classifying the detected objects belonging to the first portion as being of a specific type, the classifying of the detected objects belonging to the first portion as being of a specific type being performed using a classification threshold lower than a classification threshold used for the detected objects determined to belong to the second portion, the classifying including classifying the detected objects belonging to the second portion as being of the specific type; determining that the detected objects in the first portion and the second portion are classified as being of the specific type; masking only the objects classified as being of the specific type in the first portion and the second portion of the video stream and including. **Claim 12** When determining that the detected object belongs to the first portion, determining the speed of the detected object; selecting a classification threshold according to the speed of the detected object and including, wherein the detected object is classified using the selected classification threshold. When the speed of the object exceeds a speed threshold, classify the object using a classifier having a first classification threshold; when the speed of the object is less than the speed threshold, classify the object using a classifier having a second classification threshold higher than the first classification threshold. The method according to claim 11.
13. A speed exceeding the speed threshold indicates that the object is moving, and a speed less than the speed threshold indicates a stationary object. The method according to claim 12.
14. The classification threshold is a function of the speed of the object. The method according to claim 11.
15. The method according to claim 11, comprising classifying an object in a set of frames of the first portion and a set of frames of the second portion, and masking each object classified as being of the particular type of object.
16. The method according to claim 11, comprising segmenting the video stream into a second portion and a first portion, and subsequently determining whether the object belongs to the second portion or the first portion.
17. A non-transitory computer-readable storage medium storing instructions for performing the method according to any one of claims 1 to 16 when executed on a device having processing capabilities.