Segmentation method

By identifying coherent areas in the image and adjusting the confidence score threshold, high-precision segmentation masks are generated, and the problem of insufficient segmentation mask generation efficiency and accuracy in the prior art is solved, and faster and more accurate object-class detection is achieved.

CN115908445BActive Publication Date: 2025-07-11AXIS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211054017.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-22
Filing Date
2022-08-31
Publication Date
2025-07-11
Estimated Expiration
2042-08-31

AI Technical Summary

Technical Problem

The prior art has shortcomings in processing efficiency and accuracy in segmentation mask generation, making it difficult to quickly and accurately detect separate instances of multiple object classes in an image.

Method used

By identifying coherent regions in the image, adjusting the confidence score threshold to generate a segmentation mask, performing multiple rounds of interpretation and estimation for coherent regions, reducing processing of incoherent regions, and optimizing the segmentation process using deep learning models and motion vector analysis.

Benefits of technology

Improve the processing speed and accuracy of segmentation masks, enable faster and more accurate detection of object-class instances in the image, and reduce waste of processing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115908445B_ABST
    Figure CN115908445B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a segmentation method. Specifically, a method for generating a segmentation result is disclosed, the segmentation result indicating separate instances of one or more object classes of an image in a sequence of images. The method includes: determining a coherent region of the image; processing the image to determine a tensor representing pixel-specific confidence scores; generating a series of temporary segmentation masks for the coherent region, where each temporary segmentation mask is generated by interpreting the tensor regarding a single object class using a different temporary confidence score threshold; estimating the series of temporary segmentation masks to determine whether an object mask condition is satisfied; based on the result of the estimation, setting the temporary confidence score threshold as the final confidence score threshold for the pixels of the temporary segmentation mask, or setting a default confidence score threshold as the final confidence score threshold for the coherent region; and generating a final segmentation result of the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of segmentation and, in particular, to a method of providing a segmentation mask that indicates individual instances of one or more object classes of an image in a sequence of images. Background Art

[0002] For security and other surveillance purposes, video surveillance of objects such as buildings, people, animals, roads, and vehicles has become increasingly common. However, even a trained and attentive viewer cannot extract meaningful information from the video. As a result, there is a continuing need for surveillance systems that can detect objects for surveillance and monitoring purposes using computer surveillance technology. Recently, deep learning has allowed for more sophisticated analysis of video from distributed cameras and cloud computing surveillance systems.

[0003] Deep learning is a type of machine learning that may involve training a model, often referred to as a deep learning model. A deep learning model can be based on a set of algorithms that are designed to model abstractions in data by using multiple processing layers. The processing layers can consist of non-linear transformations, and each processing layer can transform data before passing the transformed data to a subsequent processing layer. The transformation of the data can be performed by the weights and biases of the processing layer. The processing layers can be fully connected. By way of example and not limitation, deep learning models can include neural networks and convolutional neural networks. A convolutional neural network can consist of a hierarchy of trainable filters interleaved with non-linearities and pooling. Convolutional neural networks can be used for large-scale object recognition tasks.

[0004] Deep learning models can be trained in a supervised or unsupervised setting. In a supervised setting, a deep learning model is trained using a labeled data set to classify data or accurately predict a result. When input data is input into a deep learning model, the model adjusts its weights until the model is properly fitted, which occurs as part of a cross-validation process. In an unsupervised setting, an unlabeled data set is used to train a deep learning model. The deep learning model discovers patterns in the unlabeled data set that can be used to cluster the data in the data set into groups of data with common attributes. Common clustering algorithms are hierarchical models, k-means models, and Gaussian mixture models. Thus, deep learning models can be trained to learn a representation of the data.

[0005] With the development of deep learning, it has become possible to identify objects faster and more accurately from real-time video streams of camera networks. One technique for such identification is segmentation or image segmentation. The purpose of segmentation is to label image regions according to what is depicted. The image regions that are determined to depict an object (or a collection of objects) form a segmentation mask. The segmentation result or product of an image includes one or more segmentation masks.

[0006] Common types of segmentation include semantic segmentation and instance segmentation. In semantic segmentation, each pixel belonging to the same object class is segmented into one object and is thus part of a segmentation mask. For example, all pixels detected as a person are segmented into one object, and all pixels detected as a car are segmented into another object. On the other hand, instance segmentation aims to detect each distinct object instance in an image. For example, each person in an image is segmented into separate objects and thus forms a segmentation mask. Another known type of segmentation is panoptic segmentation, which can be described as a combination of semantic and instance segmentation. In panoptic segmentation, it aims to semantically distinguish different objects and detect individual instances of different object classes in an image.

[0007] The present application and the disclosed invention relate to segmentation techniques that aim to detect at least individual instances of one or more object classes in an image. Examples of such segmentation techniques include the aforementioned instance segmentation and panoptic segmentation.

[0008] The task of segmentation can be performed using a deep learning model that is configured and trained to perform instance segmentation. Instance segmentation can be part of a panoptic segmentation process or an instance segmentation process. At a general level, the input to such a deep learning model is the image data of one or more images (such as surveillance video images) to be segmented, and the output is a tensor representing the confidence scores of one or more object classes at the pixel level for the image data. In other words, the deep learning model determines for each pixel the probability that the pixel depicts an object of each of one or more object classes.

[0009] There are multiple different deep learning networks applicable to deep learning models that are configured and trained to perform instance segmentation, and the detailed form of the input and output (i.e., what format the input image data requires and what format the output tensor has) can vary between these networks. It should also be noted that while the term "tensor" is a suitable representation for the output of currently constructed deep learning models, other future terms for the output of deep learning models or for other methods or algorithmic constructs used for instance segmentation purposes should be considered equivalent. In other words, in future applications of the present invention, a data structure that represents the confidence scores of one of the multiple object classes of an image in some way can achieve the same purpose as the tensor discussed herein and should be considered equivalent to the tensor. The term tensor is thus interchangeable with future terms for data structures representing the confidence scores of the object classes of an image.

[0010] The output from a deep learning model is interpreted, possibly together with the results from an additional deep learning model, to form a segmentation mask of an image, i.e., a mask formed by image regions that are at least labeled with their object classes detected by one or more deep learning models. A standard technique for interpretation is thresholding, where the tensor output from a deep learning model performing instance segmentation is interpreted by setting a confidence score threshold for each object class. The threshold sets the minimum confidence score required for a pixel to be interpreted as depicting the corresponding object class. Thresholding is a non-trivial task and there are various concepts on how to optimally set the threshold. Generally, too low a threshold results in a noisy segmentation mask that includes various false object detections, while too high a threshold results in a poor segmentation mask that fails to include positive object detections.

[0011] A thresholding technique is disclosed in patent application US2009 / 0310822. The document discloses an object segmentation process in which so-called object prediction information is used during segmentation to adjust the confidence score threshold at the pixel level. Examples of object prediction information are object motion information, object class information, environmental information, object depth information, and interaction information. Based on the prediction information, a pixel is initially determined as a predicted foreground pixel or a predicted background pixel. If it is assumed that the pixel is a foreground pixel, the threshold for that pixel is lowered to increase the sensitivity of the segmentation process. Otherwise, the threshold is increased to decrease the sensitivity.

[0012] Even though there are various solutions for providing an accurate segmentation mask, there is still a need for methods that are improved in terms of, for example, processing efficiency and accuracy. Summary of the Invention

[0013] The object of the present invention is to provide an improved method that provides a segmentation mask that indicates individual instances of one or more object classes with respect to processing speed and segmentation accuracy (i.e., the level of accuracy provided by the scored segmentation mask). In other words, the object is to provide a fast segmentation method that outputs a high-precision segmentation mask.

[0014] According to a first aspect, the above and other objects are achieved in whole or at least in part by the method defined in claim 1.

[0015] The present invention is based on the recognition that thresholding, i.e., adjusting the interpretation of the tensor output from a segmentation deep learning model, advantageously focuses on image regions that are likely to depict an object. These regions can be located by identifying so-called coherent regions. Thresholding is configured to fine-tune the confidence score thresholds in the coherent regions to provide a more precise object segmentation mask. Less processing can be spent on thresholding the remaining regions that are less likely to depict an object. Thus, by spending on determining appropriate confidence score thresholds for the pixels in the regions that are likely to contain an object, potentially limited processing resources can be efficiently allocated. The fine-tuning of the confidence score thresholds in the coherent regions includes generating segmentation masks using different confidence score thresholds and estimating the generated segmentation masks with respect to one or more object mask conditions. In the case of a positive result, an acceptable segmentation mask has been determined for the object of the currently estimated object class, and the confidence score threshold of this segmentation mask is maintained in the final segmentation result for the entire image. Thus, contrary to the prior art, the present invention defines a method in which the confidence score thresholds are adjusted for all pixels in an image region (i.e., a group of adjacent pixels), rather than adjusting the thresholds for individual pixels independently of each other. In other words, multiple segmentation masks are generated for a group of pixels (forming an image sub-region), and the segmentation masks are estimated as entities, as opposed to independently adjusting the confidence score thresholds at the pixel level. Due to the adjustment and estimation of the image sub-regions that are selected as having a higher probability of depicting an object as indicated by the coherent regions, the proposed method provides a segmentation process with high precision and low processing cost.

[0016] As used herein, a coherent region represents a region of contiguous pixels in an image that forms an image sub-region. The identified coherent regions correspond to regions in the image that have moved approximately the same amount in the same direction. Coherent regions can, for example, correspond to moving objects described in a video. Coherent regions can also be referred to as motion-connected regions.

[0017] Coherent regions can be determined by comparing the motion vectors of image macroblocks to find pixel regions with similar motion. By assuming that a group of adjacent motion vectors with similar, i.e., approximately the same, direction and magnitude are likely to indicate that the corresponding pixels depict the same object, such a group of motion vectors is referred to herein as a coherent region, the presence of which is used as a guide for the image region in which an object is likely to be detected by the segmentation process. Other methods for finding coherent regions in an image can be, for example, by image analysis or by using an external sensor such as a radar to locate the object.

[0018] The process of determining coherent regions from motion vectors, which can be retrieved from the encoding process of an image, is described in more detail. Motion vectors are determined during image encoding, where the image is processed by an encoding algorithm in which motion vectors are determined for groups of pixels. Once the motion vectors of the image are determined, the motion vectors can be estimated to determine one or more coherent regions. Optionally, the motion vectors can be determined as an independent process rather than as part of the encoding process. Simply put, motion vectors can be determined by searching for macroblocks in a previously captured image, where a macroblock is a block of pixels, such as 4×4 or 8×8 pixels, of the current image. If found, the motion vector of the macroblock is determined, and the motion vector defines how much and in which direction the macroblock has moved since the previously captured image. Motion vectors are a well-known concept used for temporal encoding of video and are part of various video compression standards such as H.264 and H.265.

[0019] The condition of having approximately the same direction can be defined such that the direction of the motion vector or the main part of the motion vector lies within a predefined span. For example, the maximum angle between any two motion vectors in a coherent region is below a predefined threshold angle. The condition of having approximately the same magnitude can be defined such that the magnitude of the motion vector or the main part of the motion vector lies within a predefined span. For example, the maximum magnitude difference between any two motion vectors in a coherent region is below a predefined threshold. Which threshold is appropriate depends on parameters such as the configuration and accuracy of motion vector determination. Thus, the threshold may vary between different embodiments.

[0020] The task of finding coherent regions by estimating motion vectors should be regarded as trivial and can be solved by, for example, different known estimation algorithms.

[0021] The object mask condition as used herein is defined as a condition for a segmentation mask that indicates that the mask corresponds to the object being described. The dependent claims give some non-limiting examples of such conditions, including a non-fragmented mask, a merged segmentation mask, and a smooth mask edge. The object mask condition can be specific to an embodiment and thus can vary. For example, a segmentation method for an image describing a scene mainly with people can implement an object mask condition that is particularly effective for people but may be less effective for other object classes, while a segmentation method for an image describing a scene with different object classes can implement a mask condition suitable for multiple different object classes. The estimation of the segmentation mask for the object mask condition can consider only that segmentation mask or can also consider other segmentation masks. For example, the segmentation mask can be estimated by comparing how the corresponding segmentation masks in previous and / or subsequent images satisfy the object mask condition to see if it meets the object mask condition.

[0022] As used herein, the segmentation result refers to the result of interpreting a tensor representing pixel-specific confidence scores. According to the present invention, multiple rounds of interpretation of the tensor are performed on the pixels of the coherent region. Each round of interpretation is represented by a provisional segmentation mask, which is then estimated before generating the segmentation mask that forms part of the final segmentation result of the image. The provisional segmentation mask used herein can be regarded as a binary representation of an image region, which indicates whether a pixel is detected as an object describing a specific object class. Multiple provisional segmentation masks can be formed for the same image region or for overlapping image regions but for different object classes.

[0023] The step of processing an image to determine the tensor can include processing the image through a deep learning model configured and trained to determine pixel-level confidence scores for one or more object classes based on the image data of the image. Deep learning models configured and trained for, for example, segmentation or as part of a panoptic segmentation process are general examples of such deep learning models. Non-limiting examples of deep learning networks that can be configured and trained to form deep learning models suitable for instance segmentation tasks include YOLACT, MASK-R-CNN, fully convolutional instance-aware semantic segmentation techniques such as FCIS, and the Panoptic-DeepLab panoptic segmentation model architecture. Note that although embodiments of the present invention are disclosed in the context of deep learning, which is the preferred embodiment at the filing date of this application, the present invention is not limited to using deep learning models for instance segmentation. For example, if another type of machine learning model or algorithm performs instance segmentation based on image data and provides an output representing the confidence scores of the image, then that machine learning model or algorithm can be used instead of the deep learning model.

[0024] The generation of a series of provisional segmentation masks can be regarded as an iterative process, in which the tensor is interpreted using a provisional confidence score threshold that is adjusted between iterations. The adjustment of the provisional confidence score threshold can follow a predefined scheme. For example, the adjustment can gradually change the provisional confidence score threshold between a starting value and an ending value. The starting and ending values can form a threshold span. Thus, the iterative interpretation of the tensor can include always increasing or always decreasing the provisional confidence score threshold between iterations. The advantage of this method is that the estimation of a series of provisional segmentation masks can include estimating the trend in a series of provisional segmentation masks, such as the mask is expanding, which can enable the estimation to reach a conclusion more quickly as to whether the object mask condition is met or will be met.

[0025] According to different embodiments, a starting point for confidence score threshold adjustment can be selected, i.e., the initial confidence score threshold of the first generated mask in a series of provisional segmentation masks. As described above, the provisional confidence score threshold can be adjusted between a starting value and an ending value, thus adjusted between a maximum value and a minimum value, and vice versa. An alternative approach is to first use an initial confidence score threshold (e.g., a base confidence score threshold) to generate the first provisional segmentation mask in the series, and then estimate the spatial relationship between the first provisional segmentation mask and the coherent region. If the first provisional segmentation mask is smaller than the coherent region, it is more likely that the first provisional segmentation mask does not cover all object pixels. A second provisional confidence score threshold can be determined to include more pixels compared to the initial confidence score threshold, where the second provisional segmentation mask is generated using the second confidence score threshold. Conversely, if the first provisional segmentation mask is larger than the coherent region, it is possible that the first provisional segmentation mask covers too many pixels. A second provisional confidence score threshold can be determined to include fewer pixels compared to the initial confidence score threshold, where the second provisional segmentation mask is generated using the second confidence score threshold. Accordingly, the selection of generating a series of provisional segmentation masks with a consistently increasing or consistently decreasing provisional confidence score threshold can be based on the comparison between the first provisional segmentation mask in the series and the coherent region.

[0026] Preferably, a series of provisional segmentation masks are generated by interpreting a tensor with respect to a single selectable object class. In other words, the provisional segmentation masks represent objects of a single object class according to the tensor interpreted at different confidence score thresholds. The method can be executed in parallel for different object classes, where the final confidence score threshold for each object class can be determined. Before generating the final segmentation result, the segmentation process needs to select which objects to include in the final segmentation result and which to discard. Such filtering is a known part of the segmentation process.

[0027] If needed, when generating a series of provisional segmentation masks, which object class to select can be determined by analyzing the tensor. It may be necessary or desirable to make a selection to save time and / or processing resources. In one embodiment, a single object class is determined by identifying the object class with the highest sum of confidence scores of pixels in the coherent region. Alternatively, the highest number of pixels identified as a particular object class can set the single object class. When determining the single object class, the historical object classes determined for the coherent region can also be considered. For example, if the object class of a horse has been determined for the pixels in the coherent region in a previous image, the determination of the single object class can be adjusted to be more inclined to select the object class of a horse. For example, a higher weight can be given to the confidence score of the horse in the tensor compared to the weights of other object classes. Or, compared to the confidence score thresholds of other object classes, the confidence score threshold of the object class of a horse can be shifted to a lower value.

[0028] According to one embodiment, a series of temporary segmentation masks are generated for an image region consisting of a coherent region and a surrounding edge region. Thus, a series of temporary segmentation masks are generated for the coherent region and the relatively small surrounding edge region, yet not for the entire image. An advantage of this embodiment is that the generated temporary segmentation masks can extend beyond the coherent region.

[0029] In one embodiment, the final confidence score is reused in the subsequent segmentation process of the next image in the sequence of images. For example, the final confidence score of the first image can be used as an initial confidence score threshold for generating a first temporary segmentation mask in the coherent region of the second image determined to correspond to the same object as the coherent region of the previous image. To do this, it is necessary to analyze the coherent regions of the next, second image to determine whether it is caused by the same object as the coherent region of the first image. In addition, the condition that the coherent regions should have similar (i.e., approximately the same) spatial positions and sizes can also be implemented to ensure that the coherent regions are caused by the same object. A multi-object tracking algorithm can help determine this, as objects can be tracked and identified relative to each other. The results of the multi-object tracking algorithm can be used to verify whether the objects detected in the coherent regions of two consecutive images are the same.

[0030] Thus, according to a second aspect, the present invention is a method of generating a segmentation mask that indicates individual instances of one or more object classes of an image in a sequence of images, as defined in claim 9.

[0031] According to another embodiment, even in the case where no coherent region is identified in the next, second image, the final confidence score is reused in the subsequent generation of the segmentation mask of the next image in the sequence of images. The object can still be described in the second image, yet it can be stationary, i.e., not moving, and thus not cause any coherent region. Therefore, if there is a coherent region in a previous image, for example, in any one of the 10 previous images, the final confidence score threshold can be used for the same pixels in the second image as in the previous image. Using this method, an improved mask for a stationary object can be achieved with the help of an earlier generated segmentation mask of the same object when it was moving.

[0032] The method disclosed herein can advantageously be executed in the processing device of a camera. In this case, the final segmentation result can be transmitted by the camera together with the image in an encoded format.

[0033] According to a third aspect, the invention is an image capture device configured to generate a segmentation mask that indicates separate instances of one or more object classes of an image in a sequence of images, as defined in claim 12. The image capture device of the third aspect can generally be implemented in the same manner as the method of the first aspect and has the attendant advantages.

[0034] According to a fourth aspect, the invention is a computer-readable storage medium comprising computer code which, when loaded and executed by one or more processors or control circuitry, causes the one or more processors or control circuitry to perform the method according to the first aspect.

[0035] Further applications of the invention will become apparent from the detailed description given hereinafter. However, it should be understood that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the scope of the invention will become apparent to those skilled in the art from this detailed description.

[0036] Accordingly, it should be understood that the invention is not limited to the specific components of the described device or the steps of the described method, as such devices and methods may vary. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It must be noted that, as used in the specification and the appended claims, the articles "a", "the" and "said" are intended to mean that there is one or more elements unless the context clearly dictates otherwise. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The invention will now be described in more detail by way of example and with reference to the accompanying drawings, in which:

[0038] Figure 1 Modules of a camera according to an embodiment of the invention are shown.

[0039] Figure 2 is a flow chart providing an overview of the invention according to one embodiment.

[0040] Figure 3a A sequence of images is shown.

[0041] Figure 3b Coherent regions detected in a series of images are shown.

[0042] Figure 3c A series of temporary segmentation masks are shown.

[0043] Figure 4 The final segmentation result is shown.

[0044] Figure 5 is a flow chart of a method according to an embodiment of the invention. Detailed implementation manners

[0045] Figure 1 Shown is a camera 10 according to one embodiment, the camera 10 having a configuration adapted to execute a method for generating a segmentation mask for an acquired image. The camera 10 is a digital camera that can be adapted for surveillance purposes. The camera 10 can be, for example, a fixed surveillance camera configured to generate a video of a scene viewed by a remote user. The camera 10 includes an image sensor 101 for acquiring images according to known digital imaging techniques and a transmitter 106 for transmitting the video (i.e., one or more image sequences) to a receiver via wired or wireless communication. Before transmitting the video, the camera 10 adjusts and processes the acquired image data. An image processing pipeline (IPP) 103 performs known precoding enhancements on the acquired image data, such as gain and white balance adjustments. A coding processor 102 is provided for coding the original or precoded processed image data. The coding processor 102 is adapted to use a video compression algorithm based on predictive coding to code the image data. Non-limiting examples of suitable video compression algorithms include the H.264 and H.265 video compression standards. Predictive coding or temporal coding includes intra-frame (I-frame) coding and inter-frame (P-frame or B-frame) coding. During predictive coding, motion vectors are typically determined at the macroblock level in a known manner. A motion vector represents the position of a macroblock in one image relative to the position of the same or a similar macroblock in a previously acquired image. The macroblock size varies between different video compression standards. For example, a macroblock can be formed by a group of pixels of 8×8 or 16×16 adjacent pixels. The output of the coding processor 102 is the coded image data.

[0046] The camera 10 further includes a segmentation processor 104 configured to perform image segmentation including instance segmentation. In other words, the image segmentation processor 104 is adapted to process an image, more specifically, the original or precoded processed image data, to generate a segmentation mask that indicates separate instances of one or more object classes in the captured image. Segmentation can be performed on all images or selected images of the image sequence.

[0047] The image segmentation processor 104 and the coding processor 104 can be implemented as software, where circuitry forms the respective processors, such as a microprocessor, associated with computer code instructions stored on a storage 105, the storage 105 being a (non-transitory) computer-readable medium, such as a non-volatile memory, such that the camera 10 executes any (portion of) the methods disclosed herein. Examples of non-volatile storage include read-only storage, flash storage, ferroelectric RAM, magnetic computer storage devices, optical discs, etc.

[0048] The camera 10 can include other modules that serve other purposes not related to the present invention.

[0049] It should be noted that the segmentation processor 104 does not need to be part of the camera 10. In an alternative embodiment, the segmentation processor 104 is located at a remote location from the camera 10. Thus, image segmentation can be performed remotely, for example, on a server connected to the camera 10, where the camera 10 sends image data to the server and can receive the image segmentation result from the server.

[0050] A general overview of the embodiments will now be provided with further reference to Figure 2 FIG. Figure 2 FIG. shows a flowchart of image processing regarding image segmentation and encoding in the camera 10 according to an embodiment and their interaction therebetween. Through image capture 201, image data is acquired and optionally pre-coded. The image data is provided to the segmentation processor 104 for image segmentation and the encoding processor 102. In the encoding processor 102, video compression 102 is performed together with other encoding steps not shown herein. Video compression 102 includes determining motion vectors as part of predictive coding. For encoding purposes, the motion vectors are used to represent the image sequence in a bit-efficient manner. Additionally, according to an embodiment, the motion vectors are also used for segmentation. Thus, the motion vectors from video compression 202 are extracted into the segmentation, specifically the thresholding step 204 in the segmentation. The purpose of thresholding 204 is to determine a confidence score threshold for generating the final segmentation result of the image. The confidence score threshold sets a boundary value between the confidence score that should be interpreted as a positive detection and the confidence score that should be interpreted as a negative detection. The confidence score to be interpreted is given in the tensor output from the previous segmentation 203. As discussed, segmentation 203 includes processing the image data in a known manner to generate a determined confidence score. The confidence score is determined at the pixel level and can be, for example, a value between 1 and 100, which indicates the probability that the pixel depicts an object of a specific object class. Segmentation 203 can determine confidence scores for multiple object classes, which means that each pixel can be given multiple confidence scores. Depending on the algorithm used in segmentation 203, the output format, i.e., how the confidence scores are represented, can vary. In the present application, the term tensor represents any output of a segmentation algorithm that can be used for the purposes of the present invention. The tensor represents one or more confidence scores for each pixel in the image, where the confidence scores are given for one or more object classes. Non-limiting examples of object classes include vehicle, person, car, foreground, background, head, and bicycle. Thus, the segmentation algorithm can perform segmentation at a general level, for example, finding a single instance of an object without further determining the type of the object, or the segmentation algorithm can perform segmentation at a more specific level, for example, finding single instances of a bicycle and a car.

[0051] Returning to thresholding step 204, by using the motion vectors extracted from video compression 202, the embodiment provides a way to determine the confidence score threshold of an image in an efficient manner. First, the motion vectors are analyzed to determine one or more coherent regions of the image to be segmented. The analysis includes estimating the motion vectors to identify adjacent motion vectors, i.e., the motion vectors of adjacent macroblocks, which have similar directions and similar magnitudes, thus defining the coherent regions. The identified coherent regions indicate the parts of the scene that move in a coherent manner, such as the image region of a walking person, and the remaining part of the image may not describe any moving objects. This information is used in thresholding 204 to guide the process of finding the confidence score threshold that provides an accurate mask. Specifically, thresholding 204 applies an iterative search in the identified coherent regions to find a suitable confidence score threshold, because these regions are more likely to describe objects than the remaining regions. Less processing is spent on finding the confidence score threshold for the remaining image regions, which can be regarded as background regions and are assigned a default confidence score threshold or not processed at all in image segmentation. Thus, according to one embodiment, the method can assume that objects exist only in the coherent regions, and thus does not spend any resources trying to segment the image regions outside the coherent regions, i.e., determine instances of object classes. In one embodiment, a set of coherent regions identified for an image is preprocessed before being used for thresholding 204. The purpose of the preprocessing is to filter out relevant coherent regions and discard the coherent regions that may be caused by unrelated objects or motions in the scene. The preprocessing may include comparing the set of coherent regions with a set of segmentation masks, which are identified by using one or more basic confidence score threshold interpretation tensors. Different basic confidence score thresholds may be used in different regions of the image. The basic confidence score threshold may be a predefined value or may be dynamically adjusted during image capture. For example, the basic confidence score threshold may take the same value as the confidence score threshold used for the spatially corresponding image region in a previously captured image, which is preferably the just-captured image. The set of coherent regions is filtered such that the coherent regions that at least partially overlap with the segmentation masks in the set of segmentation masks are retained, and the degree of overlap is optionally above a threshold, while the coherent regions that do not overlap or have an overlap degree below the threshold are discarded and removed from the set of coherent regions. The remaining coherent regions in the set of coherent regions, which may be referred to as relevant coherent regions, are then used for thresholding 204 as disclosed herein. Even though the term relevant coherent region is not mentioned in the rest of this description, it should be understood that the optional preprocessing of coherent regions for filtering out relevant coherent regions as disclosed above can be used in any disclosed embodiment.

[0052] Reference will now be made further to Figures 3a to 3cMore specifically, an iterative search for a suitable confidence score threshold is described in one or more determined coherent regions. Figure 3a An image sequence 30 obtained in an image capture 201 is shown. The image 31 to be segmented depicts a moving car and three living beings. Two of the living beings (on the right) are stationary, and one living being (on the left) is moving. During the encoding of the image 31, motion vectors for other images in the reference image sequence 30 are determined. As described above, the motion vectors determined for the image 31 are analyzed to identify Figure 3b The coherent regions 32a, 32b shown. Since the coherent regions 32a, 32b are typically determined at the macroblock level at the pixel group level, they provide a rough indication of the image regions depicting coherent motion. Although the estimation of motion vectors is the preferred way to currently determine coherent regions, it should be noted that there are alternative ways to determine these regions. For example, the coherent regions 32a, 32b can be identified by performing image analysis of the image 31. Known image analysis algorithms for object detection or motion detection can be used for the purpose of finding coherent regions. Another alternative is to use sensors other than the image sensor. The sensor can be, for example, a radar or other distance sensor. The coherent regions can be determined by the sensor identifying moving objects and determining their corresponding spatial coordinates in the image. These two alternatives, namely determining coherent regions by image analysis or by using one or more additional sensors, are possible for a person skilled in the art without further details.

[0053] For each coherent region, thresholding 204 generates a series of temporary segmentation masks. Figure 3c A series of temporary segmentation masks 33 generated for the coherent region 32b are shown. Each mask in the series of temporary segmentation masks 33 is generated by interpreting the tensor output from the segmentation 203 using different temporary confidence score thresholds. According to the interpretation of the tensor, the black pixels of the segmentation mask indicate that the pixel depicts an object and forms part of the temporary segmentation mask, while the white pixels indicate that the pixel does not depict an object and is not part of the temporary segmentation mask. As shown, the temporary segmentation mask can be generated by interpreting the tensor in the coherent region and the surrounding region. In this case, all the black pixels in the estimated region, i.e., all the positive interpretations, are part of the temporary segmentation mask. It should be noted that the temporary segmentation mask can consist of separate mask fragments.

[0054] Generate a series of temporary segmentation masks 33 according to the object class, which means estimating the tensor for a single object class. In the illustrated embodiment, a series of temporary segmentation masks 33 are generated for the truck object class, which means generating the masks by interpreting the confidence scores of the truck object class in the tensor. Generate a first temporary segmentation mask 34 using a first initial temporary confidence score threshold. Generate a second temporary segmentation mask 35 using a second temporary confidence score threshold, which reduces the threshold for interpreting the confidence score as a positive detection. Generate an Nth temporary segmentation mask 36 using an Nth temporary confidence score that has been further reduced. Generate an (N + 1)th temporary segmentation mask 37 using an (N + 1)th temporary confidence score that has been further reduced compared to the Nth temporary confidence score. As shown, the series of temporary segmentation masks 33 includes the temporary segmentation masks between the second temporary segmentation mask 35 and the Nth temporary segmentation mask 36.

[0055] In this embodiment, the temporary confidence score thresholds used in the series of temporary segmentation masks 33 follow a decreasing scheme. Therefore, for each generated temporary segmentation mask, the confidence score threshold is adjusted to reduce the threshold for interpreting the confidence score as a positive detection.

[0056] In addition to generating a series of temporary segmentation masks 33, thresholding 204 also performs an estimation on these masks to determine whether any mask meets the object mask condition. The object mask condition is a predefined condition for a segmentation mask to be considered as representing an object. Thus, by estimating whether any temporary segmentation mask 33 meets the object mask condition, the assumed presence of the target indicated by the coherent region can be verified or discarded. The object mask condition defines one or more features of the temporary segmentation mask 33. Non-limiting examples of features include non-fragmented, meaning the mask is not composed of multiple isolated mask fragments, and a smooth mask edge. The smoothness of the mask edge can be given by the curvature of the mask edge. The object mask condition can be defined as the maximum allowable curvature of the mask edge. Alternatively, the object mask condition can be defined as the maximum allowable deviation or difference in the mask edge curvature. The object mask condition can be specific to the object class, which means the object mask condition for verifying a vehicle can be different from the object mask condition for verifying a living being.

[0057] The sequence of the temporary segmentation mask 33 can be estimated during sequence generation or when the complete sequence has been generated. Additionally, a series of temporary segmentation masks can be estimated at the individual level or the group level. For example, in the illustrated embodiment, each of the masks 34, 35, 36, 37 in a series of temporary segmentation masks 33 can be estimated individually to determine whether any of them has a smooth mask edge, which is defined by the object mask condition defined for the object class of the truck. Alternatively, the masks 34, 35, 36, 37 can be estimated to determine whether the masks 34, 35, 36, 37 are composed of individual segments that merge into a merged mask over a part or all of the entire series 33. In Figure 3c an example of merging segments is provided, where the first temporary segmentation mask 34 includes isolated segments, and the second temporary segmentation mask 35 generated later includes fewer segments compared to the previously generated mask 34. In this example, the object mask condition can include detecting merged segments in a series of temporary segmentation masks 33 between the masks, i.e., reducing the number of segments, where the temporary confidence threshold has been adjusted to lower the threshold for positive detection. In yet another alternative, the masks 34, 35, 36, 36 can be estimated to find the best fit to the object mask condition. For example, compared to the previous mask and the subsequent (N + 1)th mask 37, it can be found that the Nth mask 36 has the maximum best fit to the mask object condition defined by the smooth mask edge. Thus, this estimation can be described as finding the optimal value that satisfies the object mask condition.

[0058] When the coherent region 32b has been estimated relative to the object mask condition, the final confidence score threshold for the pixels of the coherent region 32b or a subset of the pixels or macroblocks therein is set. If the estimation of the temporary segmentation mask is successful and thus the mask is found to satisfy or conform to the object mask condition, the final confidence score threshold is set to the temporary confidence score threshold of the mask that satisfies the object mask condition. In this case, the final confidence score threshold is set for the pixels of the temporary segmentation mask that satisfies the object mask condition. If more than one temporary segmentation mask that satisfies the object mask condition is found, it is necessary to select which temporary segmentation mask and the corresponding temporary confidence score threshold. This selection can include determining and selecting the mask that most conforms to the object mask condition, or the mask generated using the temporary confidence score threshold representing the lowest threshold for positive detection.

[0059] However, if none of the temporary segmentation masks satisfy the object mask condition, the final confidence score threshold is set to the default confidence score threshold. The default confidence score threshold can be a predefined fixed threshold or the same threshold as determined in the segmentation of coherent regions in a previous image, which is preferably the image immediately preceding. The predefined fixed threshold can be the same as the threshold for the pixels or the surrounding region of the coherent region.

[0060] As illustrated, the final confidence score threshold can be temporarily stored for use in the segmentation of subsequent images. The final confidence score threshold of the first image can be used as the initial temporary confidence score threshold in the segmentation of a second image obtained later. In another embodiment, the final confidence score is applied to the image region in the second image corresponding to the coherent region of the first image, even if no coherent region is determined in that image region. Thus, even if the object does not move in the second image, the object that moves in the first image and is identified by the coherent region can be well segmented and thus no detection of the coherent region is caused.

[0061] In yet another embodiment, when segmenting a second subsequent image that detects a second coherent region, the final confidence score threshold of the first coherent region of the first image is used. In this embodiment, it is estimated whether the first and second coherent regions are caused by the same object. The estimation can include analyzing the similarity of the motion vectors of the coherent regions, analyzing the spatial relationship between the coherent regions, or by using a separate tracking algorithm, such as a multi-object tracking algorithm, which determines and tracks individual objects in the image. According to the tracking algorithm, by determining whether there are objects with the same identity in the coherent regions, it can be determined whether the coherent regions are caused by the same object. When it is determined that this is the case, the final confidence score threshold set for the coherent region of the first image or a subset of the pixels therein can be used as the initial confidence score threshold used when generating the first temporary segmentation mask for the coherent region of the second image. The advantage of this embodiment is that by starting to generate the temporary segmentation mask by generating a mask that has been found to be applicable to the described object, the temporary segmentation mask of the coherent region of the second image can more quickly satisfy the object mask condition.

[0062] For all the coherent regions 32a, 32b identified in image 31, a process of finding a suitable final confidence score threshold is performed. The process can also be performed multiple times for a single coherent region 32a, 32b for different object classes in order to determine a suitable final confidence score threshold for each object class. To generate the final segmentation mask of image 31, the final confidence score thresholds for different image regions and different object classes are provided to mask synthesis 205. Mask synthesis 205 performs mask combination not only for the coherent regions but also for the entire image. For the image regions outside the coherent regions, segmentation can be performed by using the confidence score threshold interpretation tensor, which is set as a standard or based on, for example, the thresholds used in the previous images. Mask synthesis 205 functions according to the known principles for generating the final segmentation result. Different known algorithms and conditions can be applied to select the object class most likely described by the image regions based on the received final thresholds. In addition, setting the spatial boundaries between the segmentation masks of different object classes can also be the task of mask synthesis 205. Thus, according to the known method and based on the final confidence score threshold and the tensor, the final segmentation result is determined and provided to output synthesis 206. Output synthesis 206 also receives the encoded image data from video compression 202 and combines the outputs from the image processing of camera 10. The output formats of the encoded image data and the final segmentation mask follow the traditional standards. For example, the encoded image data can be sent from camera 10 in the form of a video stream, and the final segmentation mask can be sent as metadata.

[0063] Figure 4 An example of a composite image 41 with the final segmentation result is shown, where the final segmentation result includes a first segmentation mask 44 representing the moving organisms in image 31 and a second segmentation mask 42 representing the moving vehicles in image 31. The first and second segmentation masks 44, 42 are determined by using the method disclosed herein, that is, by iteratively determining the temporary segmentation masks to find a suitable confidence score threshold to use. The final segmentation result also includes segmentation masks 45, 46, which are determined by segmentation by using a standard or basic confidence score threshold interpretation tensor.

[0064] Figure 5 An overview of method 5 for generating the final segmentation result of an image according to an embodiment is provided, where each step of method 5 has been discussed and illustrated above.

[0065] An image (i.e., the image data of the image) is processed 501 to determine coherent regions. For example, the coherent regions can be determined during an encoder process (in an encoder processor) or during a segmentation process (in a segmentation processor). The image is also processed 502 to determine a tensor representing pixel-specific confidence scores for one or more object classes. Step 501 and step 502 can be executed in parallel or serially. It does not matter which of step 501 and step 502 is executed before the other. When steps 501 and 502 have been executed, i.e., when the coherent regions and tensors required for thresholding are available for thresholding, thresholding in the image segmentation process is performed. If one of steps 501 and 502 is completed before the other, the result of the first completed step can be locally stored, e.g., in the memory of the camera, and retrieved by the segmentation processor when the result of the second completed step is available. Next, method 5 includes a step of generating 503 a series of temporary segmentation masks for each of one or more coherent regions. As previously mentioned, one or more coherent regions determined in step 501 may have been processed to filter out relevant coherent regions. In such a case, the step of generating 503 a series of temporary segmentation masks is performed for each of one or more relevant coherent regions.

[0066] A series of temporary segmentation masks is estimated 504 according to the method under discussion. Method 5 also includes setting 505 a final confidence score threshold for the pixels of the temporary segmentation masks or for the coherent regions based on the result of the estimation 504. In the case of several coherent regions, a final confidence score threshold is set for each region. Additionally, one or more final confidence score thresholds can be set for the pixels of the remaining image regions that do not belong to any coherent region. Thereafter, method 5 includes generating 506 a final segmentation result of the image based on one or more final confidence scores. As discussed, the generation of the final segmentation result can include known methods for estimating segmentation masks for different object classes to select which segmentation masks the final segmentation result should include.

Claims

1. A method for generating a segmentation result that indicates separate instances of one or more object classes in an image of a sequence of images, the method comprising: a. determining a coherent region of the image, wherein the coherent region is a region of contiguous pixels in the image that forms a sub-region of the image, b. processing the image to determine a tensor representing pixel-specific confidence scores for one or more object classes, c. generating a series of temporary segmentation masks for the coherent region, wherein each temporary segmentation mask is generated by interpreting the tensor with respect to a single object class using a different temporary confidence score threshold, d. estimating the series of temporary segmentation masks to determine whether one or more of the temporary segmentation masks satisfy an object mask condition, wherein the object mask condition is a condition indicating that the temporary segmentation mask corresponds to a segmentation mask of the object being described, e. in the case where one or more of the temporary segmentation masks satisfy the object mask condition, setting the temporary confidence score threshold used to generate one of the one or more temporary segmentation masks as the final confidence score threshold for the pixels of the temporary segmentation mask, f. in the case where the object mask condition is not satisfied, setting a default confidence score threshold as the final confidence score threshold for the coherent region, g. generating a final segmentation result for the image, wherein a portion of the final segmentation result covering the coherent region is generated by interpreting the tensor using the final confidence score threshold.

2. The method according to claim 1, wherein Step a includes determining an image region of adjacent pixels or groups of pixels having motion vectors with approximately the same direction and approximately the same magnitude.

3. The method according to claim 2, wherein, Step a includes processing the image by an encoding algorithm to determine motion vectors of groups of pixels.

4. The method according to claim 1, wherein, The object mask condition includes at least one of the following: a condition that the temporary segmentation mask defines an unfragmented object, and a condition for merging fragments of the temporary segmentation mask.

5. The method according to claim 1, wherein Step b includes processing the image by a deep learning model.

6. The method according to claim 1, wherein, The series of temporary segmentation masks is generated by iteratively interpreting the tensor using a temporary confidence score threshold that always increases or always decreases between iterations.

7. The method according to claim 6, further comprising, determining whether a first temporary segmentation mask generated using an initial confidence score threshold is larger or smaller than the coherent region, and selecting whether to always increase or always decrease based on whether the first temporary segmentation mask is larger or smaller than the coherent region.

8. The method according to claim 1, wherein The single object class is selected by identifying the object class in the coherent region having the highest sum of pixel confidence scores.

9. The method according to claim 1, wherein Step a includes generating the series of temporary segmentation masks for an image region composed of the coherent region and a surrounding edge region.

10. The method according to claim 1, wherein The method is executed in a processing device of a camera.

11. An image capture device configured to generate a segmentation result that indicates separate instances of one or more object classes in an image of a sequence of images, the image capture device comprising: One or more image sensors and an image processor configured to acquire a sequence of said images, an encoder, and a processor adapted to perform the method according to claim 1.

12. A non-transitory computer-readable storage medium comprising computer code which, when loaded and executed by one or more processors or control circuits, causes the one or more processors or control circuits to perform the method according to claim 1.

13. A method of generating a segmentation mask indicating separate instances of one or more object classes of an image in a sequence of images, the method comprising: performing the method according to claim 1 on a first image, determining a coherent region in a second image, estimating whether the coherent region of the second image is caused by the same object as the coherent region of the first image, performing steps b to g of claim 1 on the second image, wherein, if the coherent regions of the first image and the second image are caused by the same object, using the final confidence score threshold of the coherent region of the first image to generate a first temporary segmentation mask in the series of temporary segmentation masks.

14. The method according to claim 13, wherein, The step of estimating whether the coherent region of the second image is caused by the same object as the coherent region of the first image comprises: processing the first image and the second image by a multi-object tracking algorithm.

Citation Information

Patent Citations

  • Feedback object detection method and system

    US20090310822A1

  • Foreground object detection in a video surveillance system

    US20110052003A1

  • Systems And Methods For Performing Segmentation Based On Tensor Inputs

    US20210056363A1