Apparatus and method for determining whether there is a movable object located within a scene for at least a predetermined portion of a given period.

JP7915181B2Active Publication Date: 2026-09-03AXIS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023072883
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-05
Filing Date
2023-04-27
Publication Date
2026-09-03
Estimated Expiration
2043-04-27

Smart Images

  • Figure 0007915181000001
    Figure 0007915181000001
  • Figure 0007915181000002
    Figure 0007915181000002
  • Figure 0007915181000003
    Figure 0007915181000003
Patent Text Reader

Abstract

To provide a method and an apparatus for determining whether there is a movable object located in a captured scene of a predetermined portion in a given period of time.SOLUTION: A method includes the steps of: receiving a plurality of feature vectors for a movable object in a first sequence; assigning an indicator having an initial value to a cluster of feature vectors identified among first feature vectors extracted for the movable object in the first image frame of the first sequence, or feature vectors extracted for the movable object in a second sequence of image frames captured before a given period of time preceding the first image frame; updating a value of the indicator based on similarity between the feature vector and the first feature vector or similarity between the feature vector and the cluster of feature vectors; and determining that there is a movable object located in the captured scene based on the value of the indicator.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to analyzing the position of a movable object in a captured scene over time, and more specifically to Within the captured scene for a given period home at least a predetermined portion During relates to determining whether there is a movable object located at the position.

Background Art

[0002] Identifying when a moving object, such as a person, vehicle, or bag, is located within a scene captured by a camera over a given period of time is a task of automated image analysis. This task, sometimes called wander detection, can be used to detect when a moving object is located somewhere within a captured scene for longer than would be expected from typical behavior within the captured scene, in order to trigger an alarm, further analysis of the captured scene image data, and / or further action. For example, a bag being located in a captured scene of an airport for a given period of time, or a person being located in a captured scene including an ATM for a given period of time, may constitute a situation requiring further analysis and / or further action. Various techniques have been used to address this task. For example, one technique uses tracking of moving objects within a scene, measuring the time an object is tracked within the scene and determining whether the tracked object is located within the scene for longer than a given period during a given time period. However, in such techniques, tracking of moving objects may be temporarily interrupted. For example, a movable object may be temporarily hidden, or it may temporarily leave (or be removed from) the scene and then return to (or be reintroduced to) the scene. When a movable object is identified and then tracked again after an interruption, there may be no correlation between the tracking of the movable object before the interruption and the tracking of the movable object after the interruption. To overcome this problem, a re-identification algorithm can be used to determine whether an object has been previously identified in the scene. However, such re-identification introduces other problems. For example, in a scenario where many movable objects enter (are introduced into) and leave (are removed from) the scene, the risk of a movable object being incorrectly re-identified increases over time. [Overview of the project]

[0003] The objective of the present invention is, Within the captured scene For a given period home at least a certain portion DuringThe goal is to make it easier to determine whether or not there is a movable object in position.

[0004] According to the first embodiment, Within the captured scene For a given period home at least a certain portion During A method is provided for determining whether there is a movable object located. The method includes receiving multiple feature vectors of the movable object in a first sequence of image frames captured during a given period, the feature vectors being received from a machine learning module trained to extract similar feature vectors of the same movable object in various image frames. The method uses a first feature vector extracted for the movable object in the first image frame of the first sequence of image frames, or preceding the first image frame. do The method further includes assigning an indicator with an initial value to clusters of feature vectors identified in feature vectors extracted for movable objects in image frames of a second sequence captured prior to a given period. Whether the first vector or cluster of feature vectors is alive or not can be determined based on the value of the indicator. The method further includes iteratively updating the indicator by updating the value of the indicator for each feature vector of multiple feature vectors of movable objects in the image frame for each image frame captured from the first image frame within a given period, based on the similarity between the feature vector and the first feature vector, or the similarity between the feature vector and the cluster of feature vectors, respectively. The method updates the indicator value as iteratively updated of completion Sometimes On the condition that the first feature vector or cluster of feature vectors is alive, Within the captured scene For a given period home at least a certain portion During This further includes determining whether there is a movable object located there.

[0005] This disclosure utilizes the understanding that the similarity between a feature vector and a first feature vector, or between a feature vector and a cluster of feature vectors, is a measure of the likelihood that a feature vector is associated with the same movable object as the first feature vector, or that one or more feature vectors in a cluster of feature vectors are each associated with the same movable object. This is because the feature vectors are received from a machine learning module trained to extract similar feature vectors for the same movable object in various image frames. Since similarity is determined and the indicator value is updated frame by frame over a given period, each update contributes only to a portion of the total iterative update of the indicator value. It can be relatively rare for a feature vector to be similar to the first feature vector or cluster of feature vectors even if it is not associated with the same movable object. Therefore, these cases contribute only slightly to the total iterative update of the indicator value.

[0006] The first embodiment of the method is more robust than, for example, methods based on tracking or re-identification, because it is less susceptible to the influence of a single instance if the feature vectors are similar to the first feature vector or cluster of feature vectors, even if the feature vectors are not associated with the same movable object. These instances contribute only to a fraction of the iterative updates of the indicator value. Furthermore, these instances can be made relatively rare by training the machine learning module.

[0007] A "movable object" refers to an object that can be moved or is capable of being moved, such as a person, vehicle, bag, or box.

[0008] " Within the captured scene For a given period home at least a certain portion During A "movable object" is a movable object that leaves (is removed from) the captured scene. The captured scene can be re-entered (re-introduced). Once or multiple times It's good to have, however The movable object is, for a given period home at least a certain portion During Located within the captured scene Ruko This means that the given portion can be any portion greater than 0 and less than or equal to a given period. Furthermore, movable objects do not need to be located in the same place in the scene and can move around within the scene.

[0009] "The indicator value, in relation to a specific point in time, indicates that the first feature vector or cluster of feature vectors is alive, respectively" means that the indicator value, in relation to that specific point in time, indicates that the indicator value is alive. Method If it ends, Within the captured scene For a given period home at least a certain portion During This means the value is such that it can be concluded that a movable object is located. For example, the condition for being considered alive could be based on the indicator being greater than a threshold.

[0010] Within the captured scene For a given period home at least a certain portion During the movable object The determination of location is based on the similarity of feature vectors within the image frame and the first feature vector or cluster of those feature vectors. Since this similarity is based on a machine learning module, please note that the determination can only be performed for a specific level of confidence.

[0011] The method according to the first embodiment is Within the captured scene For a given period home at least a certain portion During This could further include triggering an alarm if it detects the presence of a moving object.

[0012] In one embodiment, an indicator is assigned to the first feature vector, an initial value of the indicator is larger than a first threshold, a value of the indicator larger than the first threshold indicates that the first feature vector is alive, and a value of the indicator equal to or less than the first threshold indicates that the first feature vector is not alive. Further, the indicator is iteratively updated by i) decreasing the value of the indicator by a first amount, and ii) for each feature vector among a plurality of feature vectors for a movable object in an image frame, determining a second amount based on a similarity between the feature vector and the first feature vector, and increasing the value of the indicator by the second amount. The iterative updating is performed for each image frame captured within a given period from a first image frame, subsequent to the first image frame in a first sequence of image frames.

[0013] The feature vectors may be received from a machine learning module trained to minimize a distance between feature vectors extracted for a same movable object in different image frames, and the second amount is determined based on a distance between a further feature vector and the first feature vector. Then, it may be determined that the smaller the distance between the feature vector and the first feature vector is, the larger the second amount is.

[0014] Alternatively, the second amount may be determined to be a fixed non-zero amount when the distance between the feature vector and the first feature vector is smaller than a threshold distance, and may be determined to be zero when the distance between the feature vector and the first feature vector is equal to or larger than the threshold distance.

[0015] In one embodiment, an indicator is assigned to a cluster of feature vectors, an initial value of the indicator is greater than a second threshold, a value of the indicator indicating that the number of feature vectors determined to belong to the cluster of feature vectors per unit time is greater than the second threshold indicates that the cluster of feature vectors is alive, and a value of the indicator indicating that the number of feature vectors determined to belong to the cluster of feature vectors per unit time is less than or equal to the second threshold indicates that the first feature vector is not alive. Further, the indicator is iteratively updated by: i) decreasing the value of the indicator by a third amount, and ii) for each feature vector among the plurality of feature vectors of a movable object in an image frame, determining whether the feature vector belongs to the cluster of feature vectors based on similarity between the feature vector and the cluster of feature vectors, and increasing the value of the indicator by a fourth amount if the feature vector belongs to the cluster of feature vectors. The iterative update is performed for each image frame captured within a given period from a first image frame, subsequent to the first image frame in a first sequence of image frames.

[0016] The feature vectors may be received from a machine learning module trained to extract feature vectors that maximize the probability that the same movable object in different image frames is located within the same cluster according to a clustering algorithm, and the feature vectors are determined to belong to the cluster of feature vectors according to the clustering algorithm.

[0017] The feature vectors may be received from a machine learning module trained to extract feature vectors for the same movable object in different image frames that have a mutual distance shorter than a distance to feature vectors extracted for other movable objects, and the value of the indicator is updated based respectively on a distance between the feature vector and the first feature vector, or a distance between the feature vector and the cluster of feature vectors.

[0018] Machine learning modules can be equipped with neural networks.

[0019] According to a second embodiment, a non-temporary computer-readable storage medium is provided which, when executed by a device having a processor and a receiver, stores an instruction for carrying out the method according to the first embodiment or the method according to the first embodiment.

[0020] According to the third aspect, Within the captured scene For a given period home at least a certain portion During A method is provided for determining whether there is a movable object located. The device comprises a circuit configured to perform a receiving function, an assigning function, an updating function, and a determination function. The receiving function is configured to receive multiple feature vectors for movable objects in a first sequence of image frames captured during a given period, the feature vectors being received from a machine learning module trained to extract similar feature vectors for the same movable object in various image frames. The assigning function extracts a first feature vector for a movable object in a first image frame of the first sequence of image frames, or preceding the first image frame. do The system is configured to assign an indicator with an initial value to clusters of feature vectors identified in feature vectors extracted for movable objects in a second sequence of image frames captured before a given period. The update function is configured to iteratively update the indicator by updating the indicator value for each feature vector of multiple feature vectors for movable objects in the image frame, for each image frame captured within a given period from the first image frame following the first image frame in the first sequence of image frames, based on the similarity between the feature vector and the first feature vector, or the similarity between the feature vector and the cluster of feature vectors, respectively. The determination function determines if the indicator value is iteratively updated of completion SometimesOn the condition that the first feature vector or cluster of feature vectors is alive, Within the captured scene For a given period home at least a certain portion During It is configured to determine if there is a movable object in position.

[0021] In the assignment function, the indicator may be assigned to a first feature vector, with an initial value greater than a first threshold. An indicator value greater than the first threshold indicates that the first feature vector is alive, and an indicator value less than or equal to the first threshold indicates that the first feature vector is not alive. In the update function, the indicator can then be iteratively updated by decreasing the indicator value by a first amount, determining a second amount for each of the multiple feature vectors of a movable object in the image frame based on the similarity between the feature vector and the first feature vector, and increasing the indicator value by a second amount.

[0022] Alternatively, in the assignment function, the indicator may be assigned to a cluster of feature vectors, with an initial value greater than the second threshold. An indicator value indicating that the number of feature vectors determined to belong to the feature vector cluster per unit of time is greater than the second threshold indicates that the feature vector cluster is alive, while an indicator value indicating that the number of feature vectors determined to belong to the feature vector cluster per unit of time is less than or equal to the second threshold indicates that the first feature vector is not alive. In the update function, the indicator can then be iteratively updated by decreasing the indicator value by a third amount, determining for each feature vector of multiple feature vectors for a movable object in an image frame whether the feature vector belongs to the feature vector cluster based on the similarity between the feature vector and the feature vector cluster, and if the feature vector belongs to the feature vector cluster, increasing the indicator value by a fourth amount.

[0023] Further applications of the present invention will become apparent from the detailed description given below. However, since various changes and modifications within the scope of the present invention will become apparent to those skilled in the art from this detailed description, it should be understood that the detailed description and specific examples, while illustrating preferred embodiments of the present invention, are given merely as examples.

[0024] Therefore, it should be understood that the present invention is not limited to the operation of specific components of the described apparatus or the described method, and such apparatus and method may vary. It should also be understood that the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit them. It should be noted that, as used herein and in the appended claims, the articles “a,” “an,” “the,” and “said” are intended to mean that there is one or more elements unless the context clearly indicates otherwise. Furthermore, “comprising,” “including,” “containing,” and similar expressions are not intended to exclude other elements or steps.

[0025] The above and other aspects of the present invention will be described in more detail with reference to the accompanying drawings. The drawings should not be considered limiting and are used for illustrative and understanding purposes. [Brief explanation of the drawing]

[0026] [Figure 1] A flowchart relating to embodiments of the method disclosed herein is shown. [Figure 2] A flowchart is shown that includes a substep of an operation to iteratively update an indicator in relation to an embodiment of the method of the present disclosure. [Figure 3] A flowchart is shown that includes a substep of an operation to iteratively update an indicator in relation to an alternative embodiment of the method of the present disclosure. [Figure 4] A schematic diagram relating to an embodiment of the apparatus of this disclosure is shown. [Modes for carrying out the invention]

[0027] The present invention will be described below with reference to the accompanying drawings illustrating currently preferred embodiments of the invention. However, the present invention may be embodied in many different forms and should not be construed as being limited to the embodiments described herein.

[0028] The present invention is applicable to scenarios in which a moving object enters (is introduced) and leaves (is removed) a scene captured over time within a sequence of image frames. A sequence of image frames can form part of a video. Such a scenario occurs, for example, when a scene is captured within a sequence of image frames by a surveillance camera. The moving object is detected using an object detection module that employs any type of object detection. The object of the present invention is to detect moving objects such as people, vehicles, bags, etc., within a captured scene. Stay Dolphin, or, The Movable objects Within the captured scene The task is to determine whether an object repeatedly moves in and out of the captured scene so that it is located for more than a predetermined portion of a given period. This is sometimes called wander detection, but it has been extended to cover other movable objects besides people, and also to cover movable objects that temporarily move out of the captured scene once or multiple times over a given period. The captured scene is, for example, , possible The moving object passes through only once and / or only for the expected duration. Within the scene Stay Furthermore, it is expected that the captured scene will move away and not return for a considerable period of time. It could be a scene. Such a scene could be an airport scene or a scene including an ATM.

[0029] In the following sections, please refer to Figure 1. Within the captured scene For a given period home at least a certain portion During An embodiment of method 100 for determining whether there is a movable object in position will be described.

[0030] The predetermined portion can be any portion greater than 0 and less than or equal to a given period. The predetermined portion and the given period can be set based on the typical movement patterns of movable objects in the captured scene. For example, the predetermined portion and the given period may be: Within the captured scene For a given period home at least a certain portion During The position of a movable object may be configured to constitute a deviation from the expected movement pattern and / or a movement pattern that should be further analyzed. For example, this may constitute a movement pattern that is undesirable or suspicious from a monitoring perspective.

[0031] Setting a predetermined portion to be greater than 0 and shorter than a given period is, Movable objects Before the captured scene leaves (is removed), at least a predetermined portion of a given period Stay, The scene will not revert to (be reintroduced) the scene captured during the remainder of the given period. case Suitable for identification. of Setting a portion greater than 0 and shorter than a given period means that the movable object, in total, over a given period home at least a certain portion During It is also suitable for identifying when a movable object enters (is introduced into) and leaves (is removed from) a scene that has been captured more than once, in order to determine its position.

[0032] prescribed of Setting a portion equal to a given period means that the movable object is within the captured scene for the entirety of the given period. Stay It is only suitable for identifying when it is present.

[0033] Method 100 includes receiving multiple feature vectors for a movable object in a first sequence of image frames captured during a given period S110. This is because the feature vectors are received from a machine learning module trained to extract similar feature vectors for the same movable object in various image frames.

[0034] The first sequence of image frames may be all image frames captured during a given period, or it may be a subset of all image frames captured during a given period. If the first sequence of image frames is a subset of all image frames captured during a given period, it may be, for example, every other image frame, every three image frames, and so on. If it is a subset of all image frames, it is preferable that it consists of image frames that are uniformly distributed over the given period.

[0035] The characteristics of feature vectors depend on the type of machine learning module that receives them. For example, machine learning modules can include support vector machines (SVMs), neural networks, or other types of machine learning. See, for example, "Deep learning-based person re-identification methods: A survey and outlook of recent works" by Z. Ming et al., College of Computer Science, Sichuan University, Chendu 610065, China, 2022 (https: / / arxiv.org / abs / 2110.04764). It is also possible to use handcrafted features such as color histograms instead of feature vectors from machine learning modules. However, the latter may result in less accurate determinations.

[0036] In some embodiments, only one feature vector is extracted for each movable object in each image frame of a first sequence of image frames. Alternatively, two or more feature vectors can be extracted for each movable object in each image frame of a first sequence of image frames. For example, if two or more feature vectors are extracted, they may relate to different parts of the movable object. In this alternative, the method may be performed on one or more of the feature vectors. For example, if a separate feature vector is extracted for a person's face, and one or more other feature vectors are extracted for other parts of the person in the image frame, the method may be performed only on the feature vector related to the face, or on some or all of the one or more other feature vectors.

[0037] Method 100 further includes assigning an indicator having an initial value to a first feature vector extracted for a movable object in a first image frame of a first sequence of image frames S120, or assigning an indicator to a cluster of feature vectors extracted for a movable object in a second sequence of image frames captured a given period earlier, i.e., before the first image frame S120.

[0038] Once an indicator is assigned to a cluster of feature vectors, the method is preceded, or implicitly includes, receiving feature vectors extracted for a movable object in a second sequence of image frames, and identifying clusters of feature vectors in the feature vectors extracted for a movable object in a second sequence of image frames.

[0039] The second sequence of image frames may be any sequence of image frames captured before the first sequence of image frames, but is preferably a sequence that ends immediately before the start of the first sequence of image frames.

[0040] The second sequence of image frames may be all image frames captured during a further given period prior to the start of the first sequence of image frames, or it may be a subset of all image frames captured during that further given period. If the second sequence of image frames is a subset of all image frames captured during that further given period, it may be, for example, each second image frame, each third image frame, and so on. If it is a subset of all image frames, it is preferable that it consists of image frames uniformly distributed over the further given period.

[0041] Feature vector clusters can be identified among the feature vectors extracted for moving objects in a second sequence of image frames using any suitable clustering algorithm, such as k-means, expectation maximization (EM), or agglomerative clustering. Feature vector clusters may be defined such that the feature vectors of the feature vector clusters are likely to be extracted for the same moving object. This can be based, for example, on the similarity of the feature vectors of the feature vector clusters.

[0042] Method 100 further includes iteratively updating the indicator S130 for each image frame captured within a given period from the first image frame, following the first image frame in a first sequence of image frames. In each iteration, the update is performed for each feature vector of multiple feature vectors of a movable object in the image frame. The update includes updating the value of the indicator based on the similarity between the feature vector and the first feature vector, or the similarity between the feature vector and the cluster of feature vectors, respectively.

[0043] The indicator's value is updated based on the similarity between the feature vector and the first feature vector, or the similarity between the feature vector and the cluster of feature vectors. This means that higher similarity leads to a greater contribution to the total value update in iterative updates.

[0044] Updating the indicator value requires either the similarity between each feature vector in each image frame and the first feature vector, or the similarity between each feature vector in each image frame and a cluster of feature vectors. Further analysis is not needed to determine whether different feature vectors are extracted for the same movable object in different image frames, or whether all feature vectors in a cluster of feature vectors are extracted for the same movable object in different image frames. Therefore, feature vectors extracted for different movable objects than the first feature vector can contribute somewhat to updating the indicator value. However, as a machine learning module trained to extract similar feature vectors for the same movable object as the first feature vector, feature vectors extracted for the same movable object have a higher contribution than feature vectors extracted for different movable objects.

[0045] The same measure of similarity between feature vectors and clusters of feature vectors precedes the first image frame. do It can be used for updating, as was used when identifying clusters of feature vectors in the feature vector extracted for movable objects in a second sequence of image frames captured before a given period.

[0046] If the feature vectors extracted for a movable object by a machine learning module contain a large number of dimensions (e.g., 256), for example, if the machine learning module includes a neural network, the number of dimensions can be reduced before determining the similarity between the feature vectors and the first feature vector, or before determining the similarity between the feature vectors and clusters of feature vectors. For example, the dimension of a feature vector can be reduced by projecting the feature vector into a principal component analysis (PCA) space. Dimensionality reduction may be performed after the feature vectors are received from the machine learning module, or before they are received. In the former case, the method includes further operations to reduce the dimension of each feature vector in a group of feature vectors. In the latter case, the number of dimensions of each feature vector in a group of feature vectors has already been reduced when received from the machine learning module.

[0047] Method 100 is a method in which the indicator value is repeatedly updated. of completion Sometimes Based on condition C140, which indicates that the first feature vector or cluster of feature vectors is alive, Within the captured scene at least a predetermined portion of a given period During The process further includes determining whether there is a movable object located there (S150).

[0048] In relation to a specific point in time, "the indicator value indicates that the first feature vector or cluster of feature vectors is alive, respectively" means that the indicator value indicates that the method has terminated at that particular point in time. Within the captured scene at least a predetermined portion of a given period During This means the value is such that it can be concluded that there is a movable object located there.

[0049] Within the captured scene at least a predetermined portion of a given period DuringDetermining that there is a movable object located can be done in S150, further provided that the indicator value indicates that the clusters of the first vector or feature vectors are alive through the iterative updating of the indicator, respectively.

[0050] Method 100 is, Within the captured scene at least a predetermined portion of a given period During The system may further include S160, which triggers an alarm if it detects the presence of a moving object.

[0051] The indicator value is repeatedly updated. of completion Sometimes If the first feature vector or cluster of feature vectors indicates that they are not alive, this is based on the indicators assigned to the first feature vector or cluster of feature vectors. Within the captured scene at least a predetermined portion of a given period During This means that it cannot be said that there is a movable object in position. This situation does not usually lead to any particular further action performed in that manner. Specifically, typically there are several other feature vectors corresponding to a first feature vector being monitored during a given period, or several other clusters of feature vectors corresponding to a cluster of feature vectors, each with an assigned indicator that is updated iteratively, so that one of these several other feature vectors or several other clusters of feature vectors is Within the captured scene at least a predetermined portion of a given period During Because it may be related to other movable objects located there, Within the captured scene at least a predetermined portion of a given period During It cannot be said that there are no movable objects in position.

[0052] Method 100 may be performed on either a first feature vector or a cluster of feature vectors. However, Method 100 may also be performed in parallel on either the first feature vector or the cluster of feature vectors.

[0053] The feature vector can be received from a machine learning module trained to extract feature vectors for the same moving object in various image frames that have a shorter mutual distance than the distance to feature vectors extracted for other moving objects S110. The indicator value can then be updated based on the distance between the feature vector and the first feature vector, or the distance between the feature vector and the cluster of feature vectors, respectively.

[0054] Method 100 describes a cluster of feature vectors identified in the first feature vectors extracted for a movable object in the first image frame of a first sequence of image frames, or in the feature vectors extracted for a movable object in a second sequence of image frames. Method 100 can then be performed on each vector extracted for a movable object in the first image frame of the first sequence of image frames, or on each cluster of feature vectors identified in the feature vectors extracted for a movable object in the second sequence of image frames. Furthermore, Method 100 can be performed on each vector extracted for a movable object in each subsequent image frame. Clusters of feature vectors can be identified in the feature vectors extracted for the movable object at regular intervals, for example, intervals corresponding to the second sequence of image frames.

[0055] The following describes substeps of an operation for iteratively updating an indicator in relation to an embodiment of the method of the present disclosure, with reference to Figures 1 and 2.

[0056] In the S120 embodiment, where the indicator is assigned to the first feature vector, the initial value of the indicator may be set to a value greater than the first threshold. An indicator value greater than the first threshold indicates that the first feature vector is alive, and an indicator value less than or equal to the first threshold indicates that the first feature vector is not alive.

[0057] In relation to Figure 2, the number of frames captured within a given period from the first image frame in the first sequence of image frames is n, and therefore the iterative update S130 of the indicator is performed for each image frame i=1→n. Furthermore, the number of feature vectors extracted for each frame j is m. i There are individual features, and therefore, for each image frame i, the respective extracted feature vector j=1→m i An update will be performed on it.

[0058] Starting from the first image frame i=1 S1305, which follows the first image frame in the first sequence of image frames, the indicator value decreases by a first amount S1310. The first amount is usually the same for all iterations for all image frames i=1→n. Thus, for all iterations across all image frames i=1→n, the indicator value decreases by n·first amount. Starting from the first extracted feature vector j=1 S1315 for the first image frame i=1, a second amount is determined based on the similarity between feature vector j=1 and the first feature vector S1320. The indicator value is then increased by the second amount S1325. The index j of the feature vector is incremented by j=j+1 S1330, and the index j of the vector is j>m for image frame i=1 i For S1335, the number of feature vectors is m i As long as the following conditions are met, the operations of determining the second quantity (S1320), increasing it (S1325), and incrementing the feature vector index j (S1330) are repeated. Next, the index i of the image frame is incremented to i=i+1 (S1340). Then, the iterative update is performed for the remaining image frames i=2→n in the same manner as for the first image frame i=1.

[0059] Within the captured scene at least a predetermined portion of a given period During S150 determines that there is a movable object located and iteratively updates of completion Sometimes The condition is that the indicator value is greater than the first threshold, i.e., the first feature vector is alive.

[0060] The first quantity, the second quantity, and the first threshold are preferably such that the movable object associated with the first feature vector is Within the captured scene at least a predetermined portion of a given period During If positioned, after the iterative update in Figure 2, the indicator value must be greater than the first threshold, and the movable object associated with the first feature vector is , within the captured scene at least a predetermined portion of a given period During If not located, after the iterative update in Figure 2, the indicator value must be less than or equal to the first threshold. The value is determined by whether the movable object associated with the first feature vector is Within the captured scene at least a predetermined portion of a given period During The value can be set by an iterative process that involves evaluating the ability to accurately determine whether or not something is located. The value depends on the type of movable object and the type of machine learning module used to extract the feature vectors of the movable object. The value also depends on the number of image frames in a first sequence of image frames, which depends on the length of a given period. The value also depends on a given portion.

[0061] Feature vectors can be received from a machine learning module trained to minimize the distance between feature vectors extracted for the same moving object in various image frames S110. Then, a second quantity can be determined based on the distance between the feature vector and the first feature vector S1320. Then, the second quantity can be determined to be larger as the distance between the feature vector and the first feature vector decreases S1320.

[0062] Alternatively, the second quantity may be determined to be a fixed non-zero quantity S1320 if the distance between the feature vector and the first feature vector is less than the threshold distance, and may be determined to be zero S1320 if the distance between the feature vector and the first feature vector is greater than or equal to the threshold distance. As a result, no feature vector whose distance to the first feature vector is longer than the threshold distance contributes to updating the indicator value.

[0063] Method 100 describes a first feature vector extracted for a movable object in the first image frame of a first sequence of image frames. Method 100 can then be performed for each vector extracted for a movable object in the first image frame of the first sequence of image frames. Furthermore, Method 100 can be performed for each vector extracted for a movable object in each subsequent image frame. Thus, the method can be performed for each extracted feature vector in each image frame, and therefore each extracted feature vector has an indicator assigned to it. When a given period has elapsed with respect to an extracted feature vector, it is determined whether its assigned indicator indicates that the feature vector is alive. In this case, Within the captured scene For a given period home at least a certain portion DuringIt is determined that there is a movable object located there. Otherwise, the feature vector does not indicate that there is a movable object located in the scene captured for more than a predetermined portion of a given period. After a given period related to the feature vector, the method terminates, no further comparison is performed with extracted feature vectors in subsequent image frames related to that feature vector, and the indicator assigned to that feature vector can be discarded. Robustness is achieved by performing the method for each extracted feature vector in each image frame. If the indicator assigned to a first feature vector extracted in a first image frame and related to a movable object does not indicate that the first feature vector is alive after a given period from the first image frame, then a second feature vector extracted in a second image frame following the first image frame and related to the same movable object may have an assigned indicator indicating that the second feature vector is alive after a given period from the second image frame.

[0064] The following describes substeps of an operation for iteratively updating an indicator in relation to an alternative embodiment of the method of this disclosure, with reference to Figures 1 and 3.

[0065] In the S120 embodiment, where an indicator is assigned to a cluster of feature vectors, the initial value of the indicator can be set to a value greater than a second threshold. An indicator value indicating that the number of feature vectors determined to belong to a cluster of feature vectors per unit of time is greater than the second threshold indicates that the cluster of feature vectors is alive, while an indicator value indicating that the number of feature vectors determined to belong to a cluster of feature vectors per unit of time is less than or equal to the second threshold indicates that the first feature vector is not alive.

[0066] In relation to Figure 3, the number of frames captured within a given period from the first image frame in the first sequence of image frames is n, and therefore the iterative update S130 of the indicator is performed for each image frame from i=1 to n. Furthermore, the number of feature vectors extracted for each frame j is m. i There are individual features, and therefore, for each image frame i, the respective extracted feature vector j=1→m i An update will be performed on it.

[0067] Starting with the first image frame i=1 S13505, following the first image frame in the first sequence of image frames, the indicator value decreases by a third amount S1355. This third amount is typically the same for all iterations across all image frames i=1→n. Therefore, for all iterations across all image frames i=1→n, the indicator value decreases by n·first. Starting with the first extracted feature vector j=1 S1360 for the first image frame i=1, it is determined whether the feature vector belongs to a cluster of feature vectors based on the similarity between the feature vector and the cluster of feature vectors S1365. If the feature vector belongs to a cluster of feature vectors, the indicator value increases by a fourth amount S1370. The feature vector index j is incremented by j=j+1 S1375, and the vector index j is incremented by j>m for image frame i=1. i For S1380, the number of feature vectors is m i As long as the following conditions are met, the process of determining whether feature vector j belongs to a feature vector cluster (S1365), incrementing it (S1370), and incrementing the feature vector index j (S1375) is repeated. Next, the index i of the image frame is incremented to i=i+1 (S1385). Then, the iterative update is performed for the remaining image frames i=2→n in the same manner as for the first image frame i=1.

[0068] Within the captured scene at least a predetermined portion of a given period DuringS150 determines that there is a movable object located and iteratively updates of completion Sometimes The condition is that the indicator value is greater than the second threshold, i.e., the cluster of feature vectors is alive.

[0069] The third quantity, the fourth quantity, and the second threshold are determined by the number of movable objects associated with the cluster of feature vectors over a given period of time. Our at least a certain portion During If located within the captured scene, the value of the indicator after iterative updates in Figure 3 should be greater than the second threshold, and the movable object associated with the cluster feature vector should be greater than the second threshold for a given period of time. Our at least a certain portion During If the object is not located within the captured scene, it is preferable that the value of the indicator after iterative updates in Figure 2 be set to be less than or equal to the second threshold. The value is such that the movable object associated with the first feature vector is Within the captured scene at least a predetermined portion of a given period between The value can be set by an iterative process that involves evaluating the ability to accurately determine whether or not something is located. The value depends on the type of movable object and the type of machine learning module used to extract the feature vectors of the movable object. The value also depends on the number of image frames in a first sequence of image frames, which depends on the length of a given period. The value also depends on a given portion.

[0070] Method 100 describes clusters of feature vectors identified in feature vectors extracted for a movable object in a second sequence of image frames. Method 100 can be performed on each cluster of feature vectors identified in feature vectors extracted for a movable object in a second sequence of image frames. Clusters of feature vectors can be identified in feature vectors extracted for a movable object at regular intervals, for example, intervals corresponding to the second sequence of image frames. Method 100 can then be performed on each cluster of feature vectors identified in feature vectors extracted for a movable object at regular intervals. Cluster identification should be such that each cluster of feature vectors is likely to correspond to one movable object. By periodically identifying clusters of feature vectors, this method can be performed on new movable objects entering the scene. After a given period associated with each cluster, the method terminates, and no further comparison with extracted feature vectors in subsequent image frames is performed on that cluster, and the indicator assigned to that cluster can be discarded.

[0071] Feature vectors can be received from a machine learning module trained to extract feature vectors that maximize the probability that the same movable object in various image frames is located in the same cluster according to a clustering algorithm S110. The feature vectors can then be determined to belong to a cluster of feature vectors according to a clustering algorithm S1370.

[0072] For example, density-based spatial clustering (DBSCAN) for noisy applications can be used as a clustering algorithm, as disclosed by Ester, M. et al., for instance. "A density-based algorithm for discovering clusters in large spatial databases with noise," Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (KDD-96). AAAI Press. pp.226-231. CiteSeerX 10.1.1.121.9220. ISBN 1-57735-004-9.

[0073] Alternatively, the feature vectors can be received from a machine learning module trained to minimize the distance between feature vectors extracted for the same moving object in various image frames S110. Instead of having a fixed fourth quantity, the fourth quantity may be determined based on the distance between the feature vector and the cluster. The fourth quantity is then determined to be larger the smaller the distance between the feature vector and the cluster, and can be added to the indicator without the condition that the feature vector belongs to the cluster.

[0074] Figure 4 shows Within the captured scene For a given period home at least a certain portion During A schematic diagram relating to an embodiment of the apparatus 200 of the present disclosure for determining whether there is a movable object located there. The apparatus 200 comprises a circuit 210, which is configured to perform the functions of the apparatus 200. The circuit 210 may include a processor 212, such as a central processing unit (CPU), a microcontroller, or a microprocessor. The processor 212 is configured to execute program code, which may be configured, for example, to perform the functions of the apparatus 200.

[0075] The device 200 may further include a memory 230. The memory 230 may be one or more of a buffer, flash memory, hard drive, removable media, volatile memory, non-volatile memory, random access memory (RAM), or another suitable device. In a typical configuration, the memory 230 may include non-volatile memory for long-term data storage and volatile memory that functions as device memory for the circuit 210. The memory 230 can exchange data with the circuit 210 over a data bus. There may also be accompanying control lines and an address bus between the memory 230 and the circuit 210.

[0076] The functions of device 200 may be embodied in the form of executable logic routines (e.g., lines of code, software programs, etc.) stored in the non-temporary computer-readable medium of device 200 (e.g., memory 230) and executed by circuit 210 (e.g., using processor 212). Furthermore, the functions of device 200 may be standalone software applications or may form part of a software application that performs additional tasks related to device 200. The described functions can be thought of as methods configured to be executed by a processing unit, e.g., the processor 212 of circuit 210. The described functions may also be implemented in software, but such functions may be implemented via dedicated hardware or firmware, or any combination of hardware, firmware and / or software.

[0077] The circuit 210 is configured to perform a receiving function 231, an assignment function 232, an update function 233, and a determination function 234.

[0078] The receiving function 231 is configured to receive multiple feature vectors for a movable object in a first sequence of image frames captured during a given period, and the feature vectors are received from a machine learning module trained to extract similar feature vectors for the same movable object in various image frames.

[0079] The assignment function 232 assigns an indicator having an initial value to a first feature vector extracted for a movable object in the first image frame of a first sequence of image frames, or preceding the first image frame. do It is configured to assign to clusters of feature vectors identified in the feature vectors extracted for movable objects in a second sequence of image frames captured prior to a given period.

[0080] The update function 233 is configured to iteratively update the indicator by updating the indicator value for each feature vector of a plurality of feature vectors for a movable object in an image frame, based on the similarity between the feature vector and the first feature vector, or the similarity between the feature vector and a cluster of feature vectors, respectively, for each image frame captured within a given period from the first image frame, following the first image frame in a first sequence of image frames.

[0081] The judgment function 234 is used when the indicator value is repeatedly updated by the update function 233. of completion Sometimes On the condition that the first feature vector or cluster of feature vectors is shown to be alive, Within the captured scene For a given period home at least a certain portion During It is configured to determine if there is a movable object in position.

[0082] The device 200 may be implemented in one or more physical locations such that one or more of its functions are performed at one or more physical locations, and one or more other functions are performed at other physical locations. Alternatively, all of the functions of the device 200 may be performed as a single device located at a single physical location. This single device may be, for example, a surveillance camera.

[0083] In the assignment function 232, the indicator may be assigned to a first feature vector, with an initial value greater than a first threshold. An indicator value greater than the first threshold indicates that the first feature vector is alive, while an indicator value less than or equal to the first threshold indicates that the first feature vector is not alive. The update function 233 can then iteratively update the indicator by decreasing the indicator value by a first amount, determining a second amount for each feature vector of multiple feature vectors for a movable object in the image frame based on the similarity between the feature vector and the first feature vector, and increasing the indicator value by the second amount. For further details and possible adaptations, see Figure 2 and the description referencing it.

[0084] Alternatively, the assignment function 232 allows an indicator to be assigned to a cluster of feature vectors, with an initial value of 0. An indicator value indicating that the number of feature vectors determined to belong to the cluster per unit time is greater than a second threshold indicates that the cluster of feature vectors is alive, while an indicator value indicating that the number of feature vectors determined to belong to the cluster per unit time is less than or equal to the second threshold indicates that the first feature vector is not alive. Then, the update function 233 can iteratively update the indicator by determining, for each feature vector of multiple feature vectors for a movable object in an image frame, whether the feature vector belongs to the cluster of feature vectors based on the similarity between the feature vector and the cluster of feature vectors, and if the feature vector belongs to the cluster of feature vectors, increasing the indicator value by a fourth amount. For further details and possible adaptations, see Figure 3 and the description referencing it.

[0085] The functions performed by circuit 210 can be further adapted as corresponding steps in embodiments of the method described in relation to Figures 1, 2, and 3.

[0086] Those skilled in the art will understand that the present invention is not limited to the embodiments described above. Rather, many modifications and variations are possible within the scope of the appended claims. Such modifications and variations can be understood and achieved by those skilled in the art practicing the claimed invention from the drawings, disclosures, and studies of the appended claims.

Claims

1. A computer method for determining whether there is a movable object located within a captured scene for at least a predetermined portion of a given period of time, Receiving a plurality of feature vectors for a movable object in a first sequence of image frames captured during a given period, wherein the plurality of feature vectors are received from a machine learning module trained to extract similar feature vectors for the same movable object in various image frames. Assigning an indicator having an initial value to a first feature vector extracted for a movable object in the first image frame of the first sequence of image frames, If the aforementioned initial value is greater than the first threshold, A value of the indicator greater than the first threshold indicates that the first feature vector is alive. Assigning an indicator such that the value of the indicator being below the first threshold indicates that the first feature vector is not alive, The iterative updating of the indicator is performed for each image frame following the first image frame, which is captured within a given period of time from the first image frame in the first sequence of image frames, wherein the iterative updating of the indicator is performed The value of the indicator is reduced by a first amount, For each of the feature vectors of the plurality of feature vectors for the movable object in the image frame, The second quantity is determined based on the similarity between the feature vector and the first feature vector, The value of the indicator is increased by the second amount, This involves performing the iterative updating of the indicator, The condition that the value of the indicator indicates that the first feature vector is alive when the iterative update is completed, is used to determine that there is a movable object located in the captured scene for at least the predetermined portion of the given period. A computer implementation method, including

2. The feature vector is received from a machine learning module trained to minimize the distance between feature vectors extracted for the same movable object in various image frames. The computer implementation method according to claim 1, wherein the second quantity is determined based on the distance between a further feature vector and the first feature vector.

3. The computer implementation method according to claim 2, wherein the second amount increases as the distance between the feature vector and the first feature vector decreases.

4. The computer implementation method according to claim 2, wherein the second quantity is determined to be a fixed non-zero quantity when the distance between the feature vector and the first feature vector is less than a threshold distance, and is determined to be zero when the distance between the feature vector and the first feature vector is greater than or equal to the threshold distance.

5. The feature vector is received from a machine learning module trained to extract feature vectors for the same movable object in various image frames, wherein the distance between them is shorter than the distance to feature vectors extracted for other movable objects. The computer implementation method according to claim 1, wherein the value of the indicator is updated based on the distance between the feature vector and the first feature vector, or the distance between the feature vector and the cluster of the feature vectors, respectively.

6. The computer implementation method according to claim 1, wherein the machine learning module comprises a neural network.

7. A non-temporary computer-readable storage medium storing instructions for carrying out the method described in any one of claims 1 to 6, when executed by a device having a processor and a receiver.

8. A device for determining whether there is a movable object located within a captured scene for at least a predetermined portion of a given period, wherein the device is A receiving function configured to receive a plurality of feature vectors for a movable object in a first sequence of image frames captured during a given period, wherein the plurality of feature vectors are received from a machine learning module trained to extract similar feature vectors for the same movable object in various image frames. An assignment function configured to assign an indicator having an initial value to a first feature vector extracted for a movable object in the first image frame of the first sequence of image frames, If the aforementioned initial value is greater than the first threshold, A value of the indicator greater than the first threshold indicates that the first feature vector is alive. An assignment function in which the value of the indicator being below the first threshold indicates that the first feature vector is not active, An update function configured to perform iterative updates of the indicator for each image frame following the first image frame, which is captured within a given period of time from the first image frame in the first sequence of image frames, wherein the iterative updates of the indicator are The value of the indicator is reduced by a first amount, For each of the feature vectors of the plurality of feature vectors for the movable object in the image frame, The second quantity is determined based on the similarity between the feature vector and the first feature vector, The value of the indicator is increased by the second amount, The update function is performed by, A determination function configured to determine that there is a movable object located within the captured scene for at least a predetermined portion of a given period, provided that the value of the indicator indicates that the first feature vector is alive when the iterative update is completed. A device comprising a circuit configured to perform a certain action.

9. A computer method for determining whether there is a movable object located within a captured scene for at least a predetermined portion of a given period of time, Receiving a plurality of feature vectors for a movable object in a first sequence of image frames captured during a given period, wherein the plurality of feature vectors are received from a machine learning module trained to extract similar feature vectors for the same movable object in various image frames. Assigning an indicator having an initial value to a cluster of feature vectors identified from feature vectors extracted for a movable object in a second sequence of image frames, wherein the second sequence of image frames is captured prior to the first image frame of the first sequence of image frames and before the given period. If the aforementioned initial value is greater than the second threshold, A value of the indicator greater than the second threshold indicates that the cluster of the feature vector is alive. Assigning an indicator such that the value of the indicator being below the second threshold indicates that the cluster of the feature vector is not alive, The iterative updating of the indicator is performed for each image frame following the first image frame, which is captured within a given period from the first image frame in the first sequence of image frames, wherein the iterative updating of the indicator is performed The value of the indicator is reduced by a third amount, For each of the feature vectors of the plurality of feature vectors for the movable object in the image frame, Based on the similarity between the feature vector and the cluster of the feature vector, it is determined whether or not the feature vector belongs to the cluster of the feature vector. If the feature vector belongs to the cluster of feature vectors, the value of the indicator is increased by a fourth amount, The repeated updating of the indicator is carried out by the following: The condition that the value of the indicator indicates that the cluster of feature vectors is alive at the completion of the iterative update, and that there is a movable object located in the captured scene for at least the predetermined portion of the given period. A computer implementation method, including

10. A non-temporary computer-readable storage medium storing instructions for carrying out the method described in claim 9, when executed by a device having a processor and a receiver.

11. The feature vectors are received from a machine learning module trained to extract feature vectors that maximize the probability that the same movable object in various image frames is located within the same cluster according to a clustering algorithm. The method according to claim 9, wherein the feature vector is determined to belong to a cluster of feature vectors according to the cluster algorithm.

12. A device for determining whether there is a movable object located within a captured scene for at least a predetermined portion of a given period, wherein the device is A receiving function configured to receive a plurality of feature vectors for a movable object in a first sequence of image frames captured during a given period, wherein the plurality of feature vectors are received from a machine learning module trained to extract similar feature vectors for the same movable object in various image frames. An assignment function configured to assign an indicator having an initial value to a cluster of feature vectors identified from feature vectors extracted for a movable object in a second sequence of image frames, wherein the second sequence of image frames is captured prior to the first image frame of the first sequence of image frames and before a given period of time. If the aforementioned initial value is greater than the second threshold, A value of the indicator greater than the second threshold indicates that the cluster of the feature vector is alive. The assignment function indicates that the value of the indicator, which is below the second threshold, indicates that the cluster of the feature vector is not alive. An update function configured to perform iterative updates of the indicator for each image frame following the first image frame, which is captured within a given period of time from the first image frame in the first sequence of image frames, wherein the iterative updates of the indicator are The value of the indicator is reduced by a third amount, For each of the feature vectors of the plurality of feature vectors for the movable object in the image frame, Based on the similarity between the feature vector and the cluster of the feature vector, it is determined whether or not the feature vector belongs to the cluster of the feature vector. When the feature vector belongs to the cluster of the feature vector, the value of the indicator is increased by a fourth amount, Update functions implemented by, A determination function configured to determine that there is a movable object located in the captured scene for at least a predetermined portion of a given period, provided that the value of the indicator indicates that the cluster of feature vectors is alive when the iterative update is completed. A device comprising a circuit configured to perform a certain action.

Citation Information

Patent Citations

  • Object detection device, method, and program

    JP2021051536A

  • Information processing system, information processing method, and program

    WO2014175356A1