Apparatus and method for determining whether there is moveable object located in scene of at least predetermined portion in given period of time

A method using machine learning-based feature vector similarity updates in image frames iteratively determines the presence of movable objects, addressing the unreliability of existing tracking and re-identification methods by enhancing robustness and accuracy in loitering detection.

JP2023165641A5Pending Publication Date: 2025-12-23AXIS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023072883
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-05
Filing Date
2023-04-27
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing methods for loitering detection in image analysis, such as tracking and re-identification, are prone to errors due to temporary interruptions and increased risk of erroneous re-identification, especially in scenarios with frequent object entry and exit, leading to unreliable determination of movable objects in a scene.

Method used

A method utilizing feature vectors from a machine learning module to iteratively update an indicator based on similarity, determining the presence of a movable object by assigning an initial value to feature vectors or clusters and updating it frame by frame, ensuring robustness against temporary interruptions.

Benefits of technology

The method provides a more reliable and robust determination of movable objects in a scene by minimizing false positives and negatives, even with temporary object movements, by leveraging machine learning for feature vector similarity updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method and an apparatus for determining whether there is a movable object located in a captured scene of a predetermined portion in a given period of time.SOLUTION: A method includes the steps of: receiving a plurality of feature vectors for a movable object in a first sequence; assigning an indicator having an initial value to a cluster of feature vectors identified among first feature vectors extracted for the movable object in the first image frame of the first sequence, or feature vectors extracted for the movable object in a second sequence of image frames captured before a given period of time preceding the first image frame; updating a value of the indicator based on similarity between the feature vector and the first feature vector or similarity between the feature vector and the cluster of feature vectors; and determining that there is a movable object located in the captured scene based on the value of the indicator.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to analyzing the position of a movable object in a captured scene over time, and in particular to: In the captured scene for a given period home At least a certain portion Between It relates to determining whether there is a movable object located. [Background technology]

[0002] Identifying when a movable object, such as a person, vehicle, or bag, is located within a scene captured by a camera for a given period of time is a task of automated image analysis. This task, sometimes referred to as loitering detection, can be used to detect when a movable object is located anywhere within a captured scene for longer than expected from typical behavior within the captured scene in order to trigger an alarm, further analysis of the captured scene's image data, and / or further action. For example, a bag located within a captured scene of an airport for a given period of time, a person located within a captured scene containing an automated teller machine (ATM) for a given period of time, etc., may constitute a situation requiring further analysis and / or further action. Various approaches have been used to address this task. For example, one approach uses tracking of movable objects within a scene, measures the amount of time the object is tracked within the scene, and determines whether the tracked object is located within the scene for a given period of time. However, in such an approach, tracking of the movable object may be temporarily suspended. For example, a movable object may be temporarily hidden, or the movable object may temporarily leave (or be removed from) a scene and then return (or be reintroduced into) the scene. When a movable object is identified and tracked again after an interruption, there may be no correlation between the tracking of the movable object before the interruption and the tracking of the movable object after the interruption. To overcome this problem, a re-identification algorithm may be used to identify whether an object has been previously identified in a scene. However, such re-identification introduces other problems. For example, in a scenario where many movable objects enter (introduce) and exit (remove) a scene, the risk of a movable object being erroneously re-identified increases over time. Summary of the Invention

[0003] The object of the present invention is to In the captured scene for a given period home At least a certain portion BetweenThe object is to facilitate improved determination of whether there is a located movable object.

[0004] According to a first aspect, In the captured scene for a given period home At least a certain portion Between A method is provided for determining whether there is a movable object located. The method includes receiving a plurality of feature vectors of a movable object in a first sequence of image frames captured during a given time period, the feature vectors being received from a machine learning module trained to extract similar feature vectors of the same movable object in different image frames. The method further includes extracting a first feature vector extracted for the movable object in a first image frame of the first sequence of image frames or a feature vector of a movable object located in a first image frame preceding the first image frame. do The method further includes assigning an indicator having an initial value to a cluster of feature vectors identified among the extracted feature vectors for the movable object in the second sequence of image frames captured prior to the given time period. Whether the first vector or cluster of feature vectors is alive can be determined based on the value of the indicator. The method further includes iteratively updating the indicator for each image frame subsequent to the first image frame in the first sequence of image frames and captured within the given time period from the first image frame by updating the value of the indicator for each feature vector of the plurality of feature vectors for the movable object in the image frame based on the similarity between the feature vector and the first feature vector or the similarity between the feature vector and the cluster of feature vectors, respectively. The method further includes iteratively updating the indicator by updating the value of the indicator for each image frame subsequent to the first image frame in the first sequence of image frames and captured within the given time period from the first image frame based on the similarity between the feature vector and the first feature vector or the similarity between the feature vector and the cluster of feature vectors, respectively. of completion Sometimes and, conditional on the first feature vector or cluster of feature vectors respectively indicating that the first feature vector or cluster of feature vectors is alive, In the captured scene for a given period home At least a certain portion Between The method further includes determining that there is a located movable object.

[0005] The present disclosure utilizes the recognition that the similarity between a feature vector and a first feature vector, or between a feature vector and a cluster of feature vectors, is a measure of the likelihood that the feature vector is associated with the same movable object as the first feature vector, or that one or more feature vectors in the cluster of feature vectors are associated with the same movable object, respectively. This is because the feature vectors are received from a machine learning module trained to extract similar feature vectors for the same movable object in various image frames. Because similarity is determined and the value of the indicator is updated frame by frame during a given period, each update contributes only a portion of the total iterative update of the indicator's value. Even if the feature vector is not associated with the same movable object, cases in which the feature vector is similar to the first feature vector or cluster of feature vectors can be relatively rare. Therefore, these cases contribute only a small amount to the total iterative update of the indicator's value.

[0006] The method of the first aspect is more robust than, for example, tracking or re-identification-based methods because it is less susceptible to single instances where the feature vector is similar to the first feature vector or a cluster of feature vectors, even if the feature vectors are not associated with the same movable object. These instances contribute only a fraction of all iterative updates of the indicator value. Furthermore, these instances can be made relatively rare by training a machine learning module.

[0007] By "movable object" is meant an object that is or can be moved, such as a person, a vehicle, a bag, or a box.

[0008] " In the captured scene for a given period home At least a certain portion Between "Located movable object" refers to a movable object that leaves (is removed from) the captured scene. Re-entering (re-introducing) a captured scene One or more times It's okay to have however The movable object is, for a given period home At least a certain portion During Located within the captured scene Ruko The predetermined portion can be any portion greater than 0 and less than or equal to the given period. Furthermore, a movable object does not have to be located at the same place in the scene, but can move around in the scene.

[0009] With respect to a particular point in time, "the value of the indicator indicates that the first feature vector or cluster of feature vectors, respectively, is alive" means that the value of the indicator is alive at that particular point in time. The method If it finishes, In the captured scene for a given period home At least a certain portion Between It means that the value is such that it can be concluded that there is a movable object located. For example, the condition for being considered alive can be based on the indicator being greater than a threshold value.

[0010] In the captured scene for a given period home At least a certain portion The movable object It should be noted that the determination of location is based on the similarity of the feature vector in the image frame and the first feature vector or cluster of feature vectors, and since this similarity is based on a machine learning module, the determination can only be made to a certain degree of confidence.

[0011] The method according to the first aspect comprises: In the captured scene for a given period home At least a certain portion Between The method can further include triggering an alarm upon determining that there is a located movable object.

[0012] In one embodiment, an indicator is assigned to the first feature vector, with an initial value greater than a first threshold, where a value of the indicator greater than the first threshold indicates that the first feature vector is alive and a value of the indicator less than or equal to the first threshold indicates that the first feature vector is not alive. Furthermore, the indicator is iteratively updated by: i) decreasing the value of the indicator by a first amount; and ii) for each feature vector of a plurality of feature vectors for movable objects in the image frames, determining a second amount based on the similarity between the feature vector and the first feature vector, and increasing the value of the indicator by the second amount. The iterative updating is performed for each image frame subsequent to the first image frame in the first sequence of image frames, captured within a given period of time from the first image frame.

[0013] The feature vectors can be received from a machine learning module trained to minimize the distance between feature vectors extracted for the same movable object in different image frames, and the second amount can be determined based on the distance between the further feature vector and the first feature vector. The second amount can then be determined to be larger for smaller distances between the feature vector and the first feature vector.

[0014] Alternatively, the second quantity may be determined to be a fixed, non-zero quantity if the distance between the feature vector and the first feature vector is less than a threshold distance, and may be determined to be zero if the distance between the feature vector and the first feature vector is greater than or equal to the threshold distance.

[0015] In one embodiment, an indicator is assigned to the cluster of feature vectors, the initial value of which is greater than a second threshold, a value of the indicator indicating that the number of feature vectors determined to belong to the cluster of feature vectors per time unit is greater than the second threshold indicates that the cluster of feature vectors is alive, and a value of the indicator indicating that the number of feature vectors determined to belong to the cluster of feature vectors per time unit is equal to or less than the second threshold indicates that the first feature vector is not alive. Furthermore, the indicator is iteratively updated by: (i) decreasing the value of the indicator by a third amount; and (ii) for each feature vector of the plurality of feature vectors for movable objects in the image frames, determining whether the feature vector belongs to the cluster of feature vectors based on the similarity between the feature vector and the cluster of feature vectors, and if the feature vector belongs to the cluster of feature vectors, increasing the value of the indicator by a fourth amount. The iterative updating is performed for each image frame captured within a given period of time following the first image frame in the first sequence of image frames.

[0016] The feature vectors can be received from a machine learning module trained to extract feature vectors that maximize the probability that the same movable object in different image frames is located in the same cluster according to a cluster algorithm, and the feature vectors are determined to belong to a cluster of feature vectors according to the cluster algorithm.

[0017] The feature vectors can be received from a machine learning module trained to extract feature vectors for the same movable object in different image frames that have a mutual distance that is shorter than the distance to feature vectors extracted for other movable objects, and the value of the indicator is updated based on the distance between the feature vector and the first feature vector or the distance between the feature vector and a cluster of feature vectors, respectively.

[0018] The machine learning module may comprise a neural network.

[0019] According to a second aspect, there is provided a non-transitory computer-readable storage medium storing instructions for performing a method according to the first aspect or a method according to the first aspect when executed by an apparatus having a processing and a receiver.

[0020] According to a third aspect, In the captured scene for a given period home At least a certain portion Between A method for determining whether there is a movable object located is provided. The apparatus includes a circuit configured to perform a receiving function, an assigning function, an updating function, and a determining function. The receiving function is configured to receive a plurality of feature vectors for the movable object in a first sequence of image frames captured during a given time period, the feature vectors being received from a machine learning module trained to extract similar feature vectors for the same movable object in different image frames. The assigning function assigns the first feature vector extracted for the movable object in a first image frame of the first sequence of image frames or a feature vector for a movable object preceding the first image frame. do The updating function is configured to iteratively update the indicator by, for each image frame subsequent to a first image frame in the first sequence of image frames and captured within the given time period from the first image frame, updating the value of the indicator for each feature vector of the plurality of feature vectors for the movable object in the image frame based on a similarity between the feature vector and the first feature vector or a similarity between the feature vector and the cluster of feature vectors, respectively. The determining function determines whether the value of the indicator is greater than the iterative updating value. of completion Sometimesand, conditional on the first feature vector or cluster of feature vectors respectively indicating that the first feature vector or cluster of feature vectors is alive, In the captured scene for a given period home At least a certain portion Between The system is configured to determine that there is a movable object located.

[0021] In the assigning function, an indicator may be assigned to the first feature vector, with an initial value greater than a first threshold, where a value of the indicator greater than the first threshold indicates that the first feature vector is alive and a value of the indicator less than or equal to the first threshold indicates that the first feature vector is not alive. In the updating function, the indicator may then be iteratively updated by decreasing the value of the indicator by a first amount, and for each feature vector of a plurality of feature vectors of the movable object in the image frame, determining a second amount based on the similarity between the feature vector and the first feature vector, and increasing the value of the indicator by the second amount.

[0022] Alternatively, in the assigning function, an indicator may be assigned to the cluster of feature vectors, the initial value being greater than a second threshold, a value of the indicator indicating that the number of feature vectors determined to belong to the cluster of feature vectors per time unit is greater than the second threshold indicating that the cluster of feature vectors is alive, and a value of the indicator indicating that the number of feature vectors determined to belong to the cluster of feature vectors per time unit is equal to or less than the second threshold indicating that the first feature vector is not alive. In the updating function, the indicator may then be iteratively updated by decreasing the value of the indicator by a third amount, and for each feature vector of the plurality of feature vectors for the movable object in the image frame, determining whether the feature vector belongs to the cluster of feature vectors based on the similarity between the feature vector and the cluster of feature vectors, and if the feature vector belongs to the cluster of feature vectors, increasing the value of the indicator by a fourth amount.

[0023] Further scope of applicability of the present invention will become apparent from the detailed description given hereinafter. It should be understood, however, that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the scope of the invention will become apparent to those skilled in the art from this detailed description.

[0024] Therefore, it is to be understood that the present invention is not limited to the particular components of the apparatus described or the operation of the methods described, and that such apparatus and methods may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It should be noted that, as used in this specification and the appended claims, the articles "a," "an," "the," and "said" are intended to mean that there are one or more elements, unless the context clearly dictates otherwise. Furthermore, the words "comprising," "including," "containing," and similar expressions do not exclude other elements or steps.

[0025] These and other aspects of the present invention will be explained in more detail with reference to the accompanying drawings, which should not be considered limiting but are used for explanation and understanding. [Brief explanation of the drawings]

[0026] [Figure 1] 1 shows a flowchart relating to an embodiment of the method of the present disclosure. [Figure 2] 10 shows a flowchart including substeps of an operation for iteratively updating an indicator in accordance with an embodiment of a method of the present disclosure. [Figure 3] 10 shows a flowchart including substeps of an operation for iteratively updating an indicator in connection with an alternative embodiment of the method of the present disclosure. [Figure 4] 1 shows a schematic diagram relating to an embodiment of an apparatus of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0027] The present invention will now be described with reference to the accompanying drawings, in which a presently preferred embodiment of the invention is shown. This invention may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein.

[0028] The present invention is applicable to scenarios where a movable object enters (is introduced) and leaves (is removed) from a scene captured over time in a sequence of image frames. The sequence of image frames may form part of a video. Such a scenario occurs, for example, when a scene is captured in a sequence of image frames by a surveillance camera. The movable object has been detected using an object detection module using any kind of object detection. The object of the present invention is to detect when a movable object, such as a person, a vehicle, a bag, etc., enters (is introduced) and leaves (is removed) from a captured scene. Stay Or, The relevant Movable objects In the captured scene The goal is to determine whether an object repeatedly moves in and out of the captured scene, such that it is located for more than a predetermined portion of a given period of time. This is sometimes called loitering detection, but it can be extended to cover other movable objects besides people, and also to cover movable objects that temporarily move out of the captured scene one or more times over a given period of time. The captured scene can be, for example, , possible Moving objects pass only once and / or for the expected period of time. In the scene Stay and then move from the captured scene and are not expected to return for a significant period of time. Such a scene may be an airport scene, a scene including an ATM.

[0029] In the following, with reference to Figure 1, In the captured scene for a given period home At least a certain portion Between An embodiment of a method 100 for determining whether there is a located movable object is described.

[0030] The predetermined portion can be any portion greater than 0 and less than or equal to the given period. The predetermined portion and the given period can be set based on a typical movement pattern of a movable object in the captured scene. For example, the predetermined portion and the given period can be set based on a typical movement pattern of a movable object in the captured scene. In the captured scene for a given period home At least a certain portion Between The located movable object may be configured to constitute a deviation from an expected movement pattern and / or constitute a movement pattern that should be further analyzed, for example, this may constitute an undesirable or suspicious movement pattern from a surveillance point of view.

[0031] Setting a predetermined portion greater than 0 and less than a given period is Movable objects for at least a predetermined portion of a given period before leaving (being removed from) the captured scene. Stay, Does not return (reintroduce) to the captured scene for the remainder of a given period case Suitable for identifying the given of Setting the fraction greater than 0 and less than the given period means that the movable object will move for a total of 1000 times the given period. home At least a certain portion Between As such, it is also suitable for identifying when a movable object enters (is introduced) and leaves (is removed) a captured scene more than once.

[0032] specified of Setting the portions equal for a given period ensures that the movable object is present in the captured scene for all of the given period. Stay It is only suitable for identifying when

[0033] The method 100 includes receiving S110 a plurality of feature vectors for a movable object in a first sequence of image frames captured during a given time period, because the feature vectors are received from a machine learning module trained to extract similar feature vectors for the same movable object in different image frames.

[0034] The first sequence of image frames may be all image frames captured during a given period of time, or it may be a subset of all image frames captured during a given period of time. If the first sequence of image frames is a subset of all image frames captured during a given period of time, it may be, for example, every third image frame, every fourth image frame, etc. If it is a subset of all image frames, it preferably consists of image frames uniformly distributed over the given period of time.

[0035] The characteristics of the feature vectors depend on the type of machine learning module that receives them. For example, the machine learning module can include a support vector machine (SVM), a neural network, or other types of machine learning. See, for example, Z. Ming et al., "Deep learning-based person re-identification methods: A survey and outlook of recent works," College of Computer Science, Sichuan University, Chendu 610065, China, 2022 (https: / / arxiv.org / abs / 2110.04764). Instead of feature vectors from a machine learning module, it is also possible to use handcrafted features such as color histograms. However, the latter may result in less accurate determinations.

[0036] In some embodiments, only one feature vector is extracted for each movable object in each image frame of the first sequence of image frames. Alternatively, two or more feature vectors can be extracted for each movable object in each image frame of the first sequence of image frames. For example, if two or more feature vectors are extracted, they may relate to different portions of the movable object. In this alternative, the method may be performed on one or more of the feature vectors. For example, if a separate feature vector is extracted for a person's face and one or more other feature vectors are extracted for other portions of the person in the image frames, the method may be performed on only the feature vector associated with the face, or on some or all of the one or more other feature vectors.

[0037] The method 100 further includes assigning S120 an indicator having an initial value to a first feature vector extracted for the movable object in a first image frame of the first sequence of image frames, or assigning S120 an indicator to a cluster of feature vectors extracted for the movable object in a second sequence of image frames captured before a given period, i.e., before the first image frame.

[0038] Once the indicators are assigned to clusters of feature vectors, the method is preceded by, or implicitly includes, receiving feature vectors extracted for the movable object in the second sequence of image frames and identifying clusters of feature vectors among the feature vectors extracted for the movable object in the second sequence of image frames.

[0039] The second sequence of image frames may be any sequence of image frames captured before the first sequence of image frames, but is preferably a sequence that ends immediately before the start of the first sequence of image frames.

[0040] The second sequence of image frames may be all image frames captured during a further given period of time before the start of the first sequence of image frames, or it may be a subset of all image frames captured during the further given period of time. If the second sequence of image frames is a subset of all image frames captured during the further given period of time, it may be, for example, every second image frame, every third image frame, etc. If it is a subset of all image frames, it preferably consists of image frames uniformly distributed over the further given period of time.

[0041] Clusters of feature vectors may be identified among the feature vectors extracted for the movable objects in the second sequence of image frames using any suitable clustering algorithm, such as k-means, expectation-maximization (EM), agglomerative clustering, etc. The clusters of feature vectors may be defined such that the feature vectors of a cluster of feature vectors are likely to be extracted for the same movable object. This may be based, for example, on the similarity of the feature vectors of the cluster of feature vectors.

[0042] The method 100 further includes iteratively updating S130 the indicator for each image frame subsequent to the first image frame in the first sequence of image frames and captured within a given time period from the first image frame. At each iteration, an update is performed for each feature vector of a plurality of feature vectors of the movable object in the image frame. The update includes updating a value of the indicator based on a similarity between the feature vector and the first feature vector or a similarity between the feature vector and a cluster of feature vectors, respectively.

[0043] The indicator value is updated based on the similarity between the feature vector and the first feature vector, or between the feature vector and a cluster of feature vectors, meaning that a higher similarity results in a larger contribution to the total update of the value in the iterative update.

[0044] The indicator value is updated using only the similarity between each feature vector in each image frame and the first feature vector, or by using the similarity between each feature vector in each image frame and a cluster of feature vectors. Further analysis is not required to determine whether different feature vectors are extracted for the same movable object in different image frames, or whether all feature vectors in a cluster of feature vectors are extracted for the same movable object in different image frames. Therefore, feature vectors in image frames extracted for movable objects different from the first feature vector may contribute to updating the indicator value. However, as a machine learning module trained to extract similar feature vectors for the same movable object in different image frames, feature vectors extracted for the same movable object as the first feature vector have a higher contribution than feature vectors extracted for movable objects different from the first feature vector.

[0045] The same measure of similarity between feature vectors and clusters of feature vectors is used to measure the similarity between the first image frame and the second image frame. do It can be used for updating as used when identifying clusters of feature vectors among feature vectors extracted for moving objects in a second sequence of image frames captured before a given period of time.

[0046] If the feature vector extracted for the movable object by the machine learning module includes a large number of dimensions (e.g., 256), for example, if the machine learning module includes a neural network, the number of dimensions can be reduced before determining the similarity between the feature vector and a first feature vector or before determining the similarity between the feature vector and a cluster of feature vectors. For example, the dimensionality of the feature vector can be reduced by projecting the feature vector into a principal component analysis (PCA) space. The dimensionality reduction can be performed after the feature vector is received from the machine learning module or before it is received. In the former case, the method includes the further operation of reducing the dimensionality of each feature vector of the plurality of feature vectors. In the latter case, the dimensionality of each feature vector of the plurality of feature vectors has already been reduced when it is received from the machine learning module.

[0047] The method 100 includes repeatedly updating the value of the indicator. of completion Sometimes Based on a condition C140 that the first feature vector or cluster of feature vectors respectively indicates alive, In the captured scene At least a predetermined portion of a given period Between The method further includes determining S150 that there is a movable object located.

[0048] With respect to a particular point in time, "the value of the indicator indicates that the first feature vector or cluster of feature vectors, respectively, is alive" means that the value of the indicator indicates that, if the method were to terminate at that particular point in time, In the captured scene At least a predetermined portion of a given period Between This means that the value is such that it can be concluded that there is a movable object located.

[0049] In the captured scene At least a predetermined portion of a given period BetweenDetermining that there is a movable object located S150 may be further conditioned C140 on the value of the indicator indicating that the first vector or cluster of feature vectors, respectively, is alive through repeated updates of the indicator.

[0050] Method 100 is In the captured scene At least a predetermined portion of a given period Between The method may further include triggering an alarm S160 upon determining that a movable object is present.

[0051] The indicator value is updated repeatedly. of completion Sometimes If the first feature vector or cluster of feature vectors indicates not alive, this is based on the indicator assigned to the first feature vector or cluster of feature vectors, respectively: In the captured scene At least a predetermined portion of a given period Between This means that it cannot be said that there is a movable object located. This situation does not usually lead to any particular further action being taken in the method. Specifically, there will typically be several other feature vectors, or several other clusters of feature vectors, corresponding to the first feature vector being monitored during a given period, and respective indicators will be assigned and iteratively updated so that one of these several other feature vectors or several other clusters of feature vectors In the captured scene At least a predetermined portion of a given period Between It can be related to other movable objects that are located, In the captured scene At least a predetermined portion of a given period Between It cannot be said that there are no movable objects in position.

[0052] Method 100 may be performed on either the first feature vector or a cluster of feature vectors, however, method 100 may also be performed on the first feature vector or on a cluster of feature vectors in parallel.

[0053] The feature vectors may be received from a machine learning module trained to extract feature vectors for the same movable object in different image frames that have a shorter mutual distance than distances to feature vectors extracted for other movable objects S110. The value of the indicator may then be updated based on the distance between the feature vector and the first feature vector or the distance between the feature vector and a cluster of feature vectors, respectively.

[0054] Method 100 is described for a first feature vector extracted for a movable object in a first image frame of a first sequence of image frames, or for a cluster of feature vectors identified among feature vectors extracted for a movable object in a second sequence of image frames. Method 100 may then be performed for each vector extracted for a movable object in a first image frame of the first sequence of image frames, or for each cluster of feature vectors identified among feature vectors extracted for a movable object in a second sequence of image frames. Furthermore, method 100 may be performed for each vector extracted for a movable object in each image frame subsequent to the first image frame. Clusters of feature vectors may be identified among the feature vectors extracted for a movable object at regular intervals, for example, intervals corresponding to the second sequence of image frames.

[0055] The substeps of the operation of iteratively updating indicators in connection with the method embodiment of the present disclosure are described below with reference to FIGS.

[0056] For S120 embodiments in which an indicator is assigned to the first feature vector, the initial value of the indicator may be set to a value greater than a first threshold, where a value of the indicator greater than the first threshold indicates that the first feature vector is alive and a value of the indicator less than or equal to the first threshold indicates that the first feature vector is not alive.

[0057] 2, the number of frames subsequent to the first image frame in the first sequence of image frames that are captured within a given period of time from the first image frame is n, and therefore the iterative updating of the indicator S130 is performed for each image frame i=1→n. Furthermore, the number of feature vectors extracted for each frame j is m i , and therefore, for each image frame i, the respective extracted feature vectors j=1→m for that image frame i i The update is performed on

[0058] Starting with the first image frame i=1 S1305 following the first image frame in the first sequence of image frames, the value of the indicator is decreased S1310 by a first amount. The first amount is typically the same for all iterations for all image frames i=1→n. Thus, for iterations across all image frames i=1→n, the value of the indicator is decreased by n·first amount. Starting with the first extracted feature vector j=1 S1315 for the first image frame i=1, a second amount is determined S1320 based on the similarity between feature vector j=1 and the first feature vector. The value of the indicator is then increased S1325 by a second amount. The feature vector index j is incremented S1330 so that the vector index j is incremented for image frame i=1 j>m. i For S1335, the number of feature vectors is m i The operations of determining the second quantity S1320, incrementing S1325, and incrementing the feature vector index j S1330 are repeated as long as: Next, the image frame index i is incremented i=i+1 S1340. Iterative updates are then performed for the remaining image frames i=2→n in a similar manner as for the first image frame i=1.

[0059] In the captured scene At least a predetermined portion of a given period Between Determining that there is a movable object located S150 includes iteratively updating of completion Sometimes , the condition is that the value of the indicator is greater than the first threshold, i.e., the first feature vector is alive.

[0060] The first amount, the second amount, and the first threshold value are preferably determined based on whether the movable object associated with the first feature vector is: In the captured scene At least a predetermined portion of a given period Between If the indicator value is greater than the first threshold after the iterative update of Figure 2, the movable object associated with the first feature vector is located. , in the captured scene At least a predetermined portion of a given period Between If not located, the indicator value is set such that after the iterative update of Figure 2, the indicator value must be less than or equal to a first threshold value. The value is used to determine whether the movable object associated with the first feature vector is located within the first threshold value. In the captured scene At least a predetermined portion of a given period Between The value may be set by an iterative process involving evaluation of the method's ability to accurately determine whether a moving object is located. The value depends on the type of moving object and the type of machine learning module used to extract the feature vector of the moving object. The value also depends on the number of image frames in the first sequence of image frames, which depends on the length of a given period. The value also depends on the predetermined portion.

[0061] The feature vector may be received from a machine learning module trained to minimize the distance between feature vectors extracted for the same movable object in different image frames S110. A second quantity may then be determined based on the distance between the feature vector and the first feature vector S1320. The second quantity may then be determined to be larger for smaller distances between the feature vector and the first feature vector S1320.

[0062] Alternatively, the second quantity may be determined to be a fixed, non-zero quantity if the distance between the feature vector and the first feature vector is less than a threshold distance S1320, and may be determined to be zero if the distance between the feature vector and the first feature vector is greater than or equal to the threshold distance S1320, such that any feature vector whose distance to the first feature vector is greater than the threshold distance does not contribute to updating the indicator value.

[0063] Method 100 is described for a first feature vector extracted for a movable object in a first image frame of a first sequence of image frames. Method 100 can then be performed for each vector extracted for a movable object in the first image frame of the first sequence of image frames. Furthermore, method 100 can be performed for each vector extracted for a movable object in each image frame following the first image frame. Thus, the method can be performed for each extracted feature vector in each image frame, and thus each extracted feature vector has an indicator assigned to it. When a given time period expires for the extracted feature vector, it is determined whether the assigned indicator indicates that the feature vector is alive. In this case, In the captured scene for a given period home At least a certain portion BetweenIf the feature vector is not present in the scene captured for more than a predetermined portion of the given time period, the method terminates, no further comparisons are performed with extracted feature vectors in subsequent image frames associated with the feature vector, and the indicator assigned to the feature vector can be discarded. Robustness is achieved by performing the method for each extracted feature vector in each image frame. If the indicator assigned to a first feature vector extracted in a first image frame and associated with a movable object does not indicate that the first feature vector is alive after the given time period from the first image frame, a second feature vector extracted in a second image frame following the first image frame and associated with the same movable object can have an assigned indicator indicating that the second feature vector is alive after the given time period from the second image frame.

[0064] In the following, the substeps of the operation of iteratively updating the indicators in connection with an alternative embodiment of the method of the present disclosure will be described with reference to FIGS.

[0065] For S120 embodiments in which an indicator is assigned to a cluster of feature vectors, the initial value of the indicator may be set to a value greater than a second threshold. A value of the indicator indicating that the number of feature vectors determined to belong to the cluster of feature vectors per time unit is greater than the second threshold indicates that the cluster of feature vectors is alive, and a value of the indicator indicating that the number of feature vectors determined to belong to the cluster of feature vectors per time unit is equal to or less than the second threshold indicates that the first feature vector is not alive.

[0066] 3, the number of frames captured within a given time period following the first image frame in the first sequence of image frames is n, and therefore the iterative updating of the indicator S130 is performed for each image frame i=1→n. Furthermore, the number of feature vectors extracted for each frame j is m i , and therefore, for each image frame i, the respective extracted feature vectors j=1→m for that image frame i i The update is performed on

[0067] Starting with the first image frame i=1 S13505 following the first image frame in the first sequence of image frames, the value of the indicator is decreased S1355 by a third amount. The third amount is typically the same for all iterations for all image frames i=1→n. Thus, for iterations across all image frames i=1→n, the value of the indicator is decreased by n·first amount. Starting with the first extracted feature vector j=1 S1360 for the first image frame i=1, it is determined S1365 whether the feature vector belongs to a cluster of feature vectors based on the similarity between the feature vector and the cluster of feature vectors, and if the feature vector belongs to the cluster of feature vectors, the value of the indicator is increased S1370 by a fourth amount. The feature vector index j is incremented at j=j+1 S1375, and vector index j is incremented S1376 for image frame i=1 j>m. i For S1380, the number of feature vectors is m i The operations of determining whether feature vector j belongs to a cluster of feature vectors S1365, incrementing S1370, and incrementing feature vector index j S1375 are repeated as long as: Next, image frame index i is incremented i=i+1 S1385. Iterative updates are then performed for the remaining image frames i=2→n in a similar manner as for the first image frame i=1.

[0068] In the captured scene At least a predetermined portion of a given period BetweenDetermining that there is a movable object located S150 includes iteratively updating of completion Sometimes ,Condition is that the value of the indicator is greater than the second threshold,,i.e., the cluster of feature vectors is alive.

[0069] The third amount, the fourth amount, and the second threshold value are used to determine whether a moving object associated with a cluster of feature vectors is moving within a given period of time. Our At least a certain portion Between ,If located within the captured scene, the value of the indicator after ,iterative update in Fig. 3 is greater than the second threshold, and the ,movable object associated with the cluster feature vector is located within the ,captured scene for a given period. Our At least a certain portion Between , the value of the indicator after iterative updating of FIG. 2 is preferably set to be less than or equal to a second threshold value if the movable object associated with the first feature vector is not located within the captured scene. In the captured scene for at least a predetermined portion of a given period between The value may be set by an iterative process involving evaluation of the method's ability to accurately determine whether a moving object is located. The value depends on the type of moving object and the type of machine learning module used to extract the feature vector of the moving object. The value also depends on the number of image frames in the first sequence of image frames, which depends on the length of a given period. The value also depends on the predetermined portion.

[0070] Method 100 is described for clusters of feature vectors identified among feature vectors extracted for movable objects in a second sequence of image frames. Method 100 can be performed for each cluster of feature vectors identified among feature vectors extracted for movable objects in the second sequence of image frames. Clusters of feature vectors can be identified among feature vectors extracted for movable objects at regular intervals, for example, intervals corresponding to the second sequence of image frames. Method 100 can then be performed for each cluster of feature vectors identified among feature vectors extracted for movable objects at regular intervals. The identification of clusters should be such that each cluster of feature vectors is likely to correspond to a movable object. By periodically identifying clusters of feature vectors, the method can be performed for new movable objects entering the scene. After a given period associated with each cluster, the method terminates, no further comparisons with extracted feature vectors in subsequent image frames are performed for that cluster, and the indicator assigned to that cluster can be discarded.

[0071] The feature vectors may be received from a machine learning module trained to extract feature vectors that maximize the probability that the same movable object in different image frames is located in the same cluster according to a clustering algorithm S110. The feature vectors may then be determined to belong to a cluster of feature vectors according to the clustering algorithm S1370.

[0072] For example, density-based spatial clustering for applications with noise (DBSCAN) can be used as a clustering algorithm, as disclosed, for example, in Ester, M. et al., "A density-based algorithm for discovering clusters in large spatial databases with noise," Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (KDD-96). AAAI Press. pp. 226-231. CiteSeerX 10.1.1.121.9220. ISBN 1-57735-004-9.

[0073] Alternatively, the feature vectors may be received from a machine learning module trained to minimize the distance between feature vectors extracted for the same movable object in different image frames S110. Instead of having a fixed fourth quantity, the fourth quantity may be determined based on the distance between the feature vector and the cluster. The fourth quantity may then be determined to be larger the smaller the distance between the feature vector and the cluster, and may be added to the indicator without the requirement that the feature vector belong to the cluster.

[0074] Figure 4 shows In the captured scene for a given period home At least a certain portion Between 2 shows a schematic diagram of an embodiment of an apparatus 200 of the present disclosure for determining whether there is a located movable object. The apparatus 200 comprises a circuit 210. The circuit 210 is configured to perform the functions of the apparatus 200. The circuit 210 may include a processor 212, such as a central processing unit (CPU), a microcontroller, or a microprocessor. The processor 212 is configured to execute program code. The program code may, for example, be configured to perform the functions of the apparatus 200.

[0075] The device 200 may further include a memory 230. The memory 230 may be one or more of a buffer, flash memory, a hard drive, removable media, volatile memory, non-volatile memory, random access memory (RAM), or another suitable device. In a typical arrangement, the memory 230 may include non-volatile memory for long-term data storage and volatile memory that serves as device memory for the circuit 210. The memory 230 may exchange data with the circuit 210 over a data bus. Associated control lines and address buses may also be present between the memory 230 and the circuit 210.

[0076] The functionality of device 200 may be embodied in the form of executable logic routines (e.g., lines of code, software programs, etc.) stored on a non-transitory computer-readable medium (e.g., memory 230) of device 200 and executed by circuit 210 (e.g., using processor 212). Furthermore, the functionality of device 200 may be a standalone software application or may form part of a software application that performs additional tasks related to device 200. The described functionality may be thought of as methods that a processing unit, e.g., processor 212 of circuit 210, is configured to execute. Also, while the described functionality may be implemented in software, such functionality may also be performed via dedicated hardware or firmware, or some combination of hardware, firmware, and / or software.

[0077] The circuit 210 is configured to perform a receiving function 231 , an allocating function 232 , an updating function 233 , and a determining function 234 .

[0078] The receiving function 231 is configured to receive a plurality of feature vectors for a movable object in a first sequence of image frames captured during a given period, the feature vectors being received from a machine learning module trained to extract similar feature vectors for the same movable object in different image frames.

[0079] The assignment function 232 assigns an indicator having an initial value to a first feature vector extracted for a movable object in a first image frame of a first sequence of image frames, or to a feature vector preceding the first image frame. do and assigning to a cluster of feature vectors identified among the feature vectors extracted for the movable object in the second sequence of image frames captured before the given period of time.

[0080] The update function 233 is configured to iteratively update the indicator by, for each image frame subsequent to the first image frame in the first sequence of image frames and captured within a given period from the first image frame, updating the value of the indicator for each feature vector of the plurality of feature vectors for the movable object in the image frame based on the similarity between the feature vector and the first feature vector or the similarity between the feature vector and a cluster of feature vectors, respectively.

[0081] The decision function 234 determines whether the value of the indicator is updated by the iterative update function 233. of completion Sometimes If the first feature vector or cluster of feature vectors indicates alive, respectively, In the captured scene for a given period home At least a certain portion Between The system is configured to determine that there is a movable object located.

[0082] Device 200 may be implemented in one or more physical locations, with one or more of its functions performed at one physical location and one or more other functions performed at another physical location. Alternatively, all of the functions of device 200 may be implemented as a single device in a single physical location. The single device may be, for example, a surveillance camera.

[0083] In an assignment function 232, an indicator may be assigned to the first feature vector, with an initial value greater than a first threshold, where a value of the indicator greater than the first threshold indicates that the first feature vector is alive and a value of the indicator less than or equal to the first threshold indicates that the first feature vector is not alive. In an update function 233, the indicator may then be iteratively updated by decreasing the value of the indicator by a first amount, determining, for each feature vector of the plurality of feature vectors for a movable object in the image frame, a second amount based on the similarity between the feature vector and the first feature vector, and increasing the value of the indicator by the second amount. For further details and possible adaptations, see FIG. 2 and the description that references it.

[0084] Alternatively, the assigning function 232 can assign an indicator to the cluster of feature vectors, with an initial value of 0, where a value of the indicator indicating that the number of feature vectors determined to belong to the cluster of feature vectors per unit time is greater than a second threshold indicates that the cluster of feature vectors is alive, and a value of the indicator indicating that the number of feature vectors determined to belong to the cluster of feature vectors per unit time is equal to or less than the second threshold indicates that the first feature vector is not alive. The updating function 233 can then iteratively update the indicator by determining, for each feature vector of a plurality of feature vectors for movable objects in the image frame, whether the feature vector belongs to the cluster of feature vectors based on the similarity between the feature vector and the cluster of feature vectors, and increasing the value of the indicator by a fourth amount if the feature vector belongs to the cluster of feature vectors. For further details and possible adaptations, see FIG. 3 and the description referencing it.

[0085] The functions performed by circuit 210 may further be adapted as corresponding steps of the method embodiments described in connection with FIGS.

[0086] Those skilled in the art will understand that the present invention is not limited to the above-described embodiments. Rather, many modifications and variations are possible within the scope of the appended claims. Such modifications and variations can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.

Claims

1. A computer-implemented method for determining whether there is a movable object located within a captured scene for at least a predetermined portion of a given period of time, comprising: receiving a plurality of feature vectors for a movable object in a first sequence of image frames captured during the given time period, the plurality of feature vectors being received from a machine learning module trained to extract similar feature vectors for the same movable object in different image frames; assigning an indicator having an initial value to a first feature vector extracted for a movable object in a first image frame of said first sequence of image frames; the initial value is greater than a first threshold; a value of the indicator greater than the first threshold indicates that the first feature vector is alive; assigning an indicator, a value of the indicator being less than or equal to the first threshold indicating that the first feature vector is not alive; performing a recursive update of the indicator for each subsequent image frame captured within the given time period from the first image frame in the first sequence of image frames, the recursive update of the indicator comprising: decreasing the value of the indicator by a first amount; For each feature vector of the plurality of feature vectors for a movable object in the image frame: determining a second amount based on the similarity between the feature vector and the first feature vector; increasing the value of the indicator by the second amount; performing a repeated update of said indicator, performed by determining that there is a movable object located within the captured scene for at least the predetermined portion of the given time period, provided that the value of the indicator indicates that the first feature vector is alive upon completion of the iterative updates; 20. A computer-implemented method comprising:

2. the feature vectors are received from a machine learning module trained to minimize the distance between feature vectors extracted for the same movable object in different image frames; The computer-implemented method of claim 1 , wherein the second amount is determined based on a distance between a further feature vector and the first feature vector.

3. The computer-implemented method of claim 2, wherein the second quantity increases as the distance between the feature vector and the first feature vector decreases.

4. 3. The computer-implemented method of claim 2, wherein the second quantity is determined to be a fixed non-zero quantity if the distance between the feature vector and the first feature vector is less than a threshold distance, and is determined to be zero if the distance between the feature vector and the first feature vector is greater than or equal to the threshold distance.

5. the feature vectors are received from a machine learning module trained to extract feature vectors for the same movable object in different image frames that have a mutual distance that is shorter than a distance to feature vectors extracted for other movable objects; The computer-implemented method of claim 1 , wherein the value of the indicator is updated based on a distance between the feature vector and the first feature vector or a distance between the feature vector and a cluster of the feature vector, respectively.

6. The computer-implemented method of claim 1 , wherein the machine learning module comprises a neural network.

7. A non-transitory computer-readable storage medium storing instructions for performing the method of any one of claims 1 to 6 when executed by a device having a processor and a receiver.

8. An apparatus for determining whether there is a movable object located within a captured scene for at least a predetermined portion of a given period of time, said apparatus comprising: a receiving function configured to receive a plurality of feature vectors for a movable object in a first sequence of image frames captured during the given time period, the plurality of feature vectors being received from a machine learning module trained to extract similar feature vectors for the same movable object in different image frames; an assignment function configured to assign an indicator having an initial value to a first feature vector extracted for a movable object in a first image frame of said first sequence of image frames, the initial value is greater than a first threshold; a value of the indicator greater than the first threshold indicates that the first feature vector is alive; an assignment function, wherein a value of the indicator less than or equal to the first threshold indicates that the first feature vector is not alive; an update function configured to perform a recursive update of the indicator for each subsequent image frame captured within the given time period from the first image frame in the first sequence of image frames, the recursive update of the indicator comprising: decreasing the value of the indicator by a first amount; For each feature vector of the plurality of feature vectors for a movable object in the image frame: Determining a second amount; increasing the value of the indicator by the second amount; an update function, performed by a determining function configured to determine, on condition that the value of the indicator indicates that the first feature vector is alive upon completion of the iterative updating, that there is a movable object within the captured scene that has been located for at least the predetermined portion of the given period of time; 11. An apparatus comprising: a circuit configured to perform 9. A computer-implemented method for determining whether there is a movable object located within a captured scene for at least a predetermined portion of a given period of time, comprising: receiving a plurality of feature vectors for a movable object in a first sequence of image frames captured during the given time period, the plurality of feature vectors being received from a machine learning module trained to extract similar feature vectors for the same movable object in different image frames; assigning an indicator having an initial value to a cluster of feature vectors identified from among feature vectors extracted for a movable object in a second sequence of image frames, the second sequence of image frames being captured prior to the given time period preceding a first image frame in the first sequence of image frames; the initial value is greater than a second threshold value; the value of the indicator is the number of feature vectors determined to belong to the cluster of feature vectors; a value of the indicator greater than the second threshold indicates that the cluster of feature vectors is alive; assigning an indicator, a value of the indicator less than or equal to the second threshold indicating that the cluster of feature vectors is not alive; performing a recursive update of the indicator for each subsequent image frame captured within the given time period from the first image frame in the first sequence of image frames, the recursive update of the indicator comprising: decreasing the value of the indicator by a third amount; For each feature vector of the plurality of feature vectors for a movable object in the image frame: determining whether the feature vector belongs to the cluster of feature vectors based on the similarity between the feature vector and the cluster of feature vectors; increasing the value of the indicator by a fourth amount if the feature vector belongs to the cluster of feature vectors; performing a repeated update of said indicator, performed by determining that there is a movable object located within the captured scene for at least the predetermined portion of the given time period, provided that the value of the indicator indicates that the cluster of feature vectors is alive upon completion of the iterative updating; 20. A computer-implemented method comprising:

10. A non-transitory computer-readable storage medium storing instructions for performing the method of claim 9 when executed by a device having a processor and a receiver.

11. The feature vectors are received from a machine learning module trained to extract feature vectors that maximize the probability that the same movable object in different image frames is located in the same cluster according to a clustering algorithm; The method of claim 9 , wherein the feature vector is determined to belong to a cluster of feature vectors according to the cluster algorithm.

12. An apparatus for determining whether there is a movable object located within a captured scene for at least a predetermined portion of a given period of time, said apparatus comprising: a receiving function configured to receive a plurality of feature vectors for a movable object in a first sequence of image frames captured during the given time period, the plurality of feature vectors being received from a machine learning module trained to extract similar feature vectors for the same movable object in different image frames; an assigning function configured to assign an indicator having an initial value to a cluster of feature vectors identified from among feature vectors extracted for a movable object in a second sequence of image frames, the second sequence of image frames being captured prior to a first image frame of the first sequence of image frames and before the given period of time; the initial value is greater than a second threshold value; the value of the indicator is the number of feature vectors determined to belong to the cluster of feature vectors; a value of the indicator greater than the second threshold indicates that the cluster of feature vectors is alive; an assignment function, wherein a value of the indicator less than or equal to the second threshold indicates that the cluster of feature vectors is not alive; an update function configured to perform a recursive update of the indicator for each subsequent image frame captured within the given time period from the first image frame in the first sequence of image frames, the recursive update of the indicator comprising: decreasing the value of the indicator by a third amount; For each feature vector of the plurality of feature vectors for a movable object in the image frame: determining whether the feature vector belongs to the cluster of feature vectors based on a similarity between the feature vector and the cluster of feature vectors; increasing the value of the indicator by a fourth amount if the feature vector belongs to the cluster of feature vectors; an update function, implemented by a determining function configured to determine, on condition that the value of the indicator indicates that the cluster of feature vectors is alive upon completion of the iterative updating, that there is a movable object located within the captured scene for at least the predetermined portion of the given time period; 11. An apparatus comprising: a circuit configured to perform