Segmentation method

By adjusting the reliability score threshold specifically in coherent regions of images, the method enhances the precision and efficiency of segmentation mask generation, addressing the limitations of existing segmentation techniques.

JP7695222B2Active Publication Date: 2025-06-18AXIS
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022144992
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-22
Filing Date
2022-09-13
Publication Date
2025-06-18
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

Existing segmentation methods face challenges in achieving high processing speed and accuracy for generating segmentation masks that indicate individual instances of object classes in images.

Method used

The method focuses on thresholding in image areas likely to indicate objects, identified as coherent regions, where the reliability score threshold is finely adjusted to enhance object segmentation precision. This approach reduces processing resources by minimizing threshold adjustments in non-object areas.

Benefits of technology

This method achieves high-precision segmentation with reduced processing costs by efficiently allocating resources to determine suitable reliability score thresholds only in areas likely to contain objects, thereby improving both speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695222000001
    Figure 0007695222000001
  • Figure 0007695222000002
    Figure 0007695222000002
  • Figure 0007695222000003
    Figure 0007695222000003
Patent Text Reader

Abstract

To provide a method of generating a segmentation outcome which indicates individual instances of one or more object classes for an image in a sequence of images.SOLUTION: A method comprises: determining a coherent region of an image; processing the image to determine a tensor representing pixel-specific confidence scores; generating a series of temporary segmentation masks for the coherent region; evaluating the series of temporary segmentation masks to determine if an object mask condition is met; depending on the outcome of the evaluation, setting the temporary confidence score threshold as a final confidence score threshold for the pixels of the temporary segmentation mask, or setting a default confidence score threshold as a final confidence score threshold for the coherent region; and generating a final segmentation outcome for the image.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of segmentation, and more particularly to a method for providing a segmentation mask that indicates individual instances of one or more object classes for images in a sequence of images.

Background Art

[0002] Video monitoring of objects such as buildings, people, animals, roads, and vehicles for security and other surveillance purposes is becoming increasingly common. However, even trained and vigilant observers have limitations in what they can extract from video as meaningful information. As a result, there is a continuing increase in the demand for surveillance systems that can detect objects for monitoring and surveillance purposes using computer vision technology. Recently, deep learning has enabled more complex analysis of video feeds from distributed camera and cloud computing surveillance systems.

[0003] Deep learning is a type of machine learning that typically involves training a model, often referred to as a deep learning model. A deep learning model can be based on a set of algorithms designed to model abstractions in data by using several processing layers. The processing layers can consist of non-linear transformations, and each processing layer can transform the data and then pass the transformed data to a subsequent processing layer. The transformation of the data can be performed by the weights and biases of the processing layer. The processing layers can be fully connected. Deep learning models can include, by way of example and not limitation, neural networks and convolutional neural networks. A convolutional neural network can consist of a hierarchy of trainable filters interleaved with non-linearity and pooling. Convolutional neural networks can be used in large-scale object recognition tasks.

[0004] Deep learning models can be trained in a supervised or unsupervised setting. In a supervised setting, a deep learning model is trained using a labeled dataset to classify data or accurately predict an outcome. When input data is fed into a deep learning model, the model adjusts its weights until the model is properly fitted, which is done as part of an iterative process. In an unsupervised setting, a deep learning model is trained using an unlabeled dataset. From the unlabeled dataset, a deep learning model discovers patterns that can be used to cluster the data from the dataset into groups of data with common properties. Common clustering algorithms are hierarchical, k-means, and Gaussian mixture models. Thus, a deep learning model can be trained to learn the representation of data.

[0005] With the development of deep learning, faster and more accurate recognition of objects in live video streams from camera networks is becoming available. One technique for such recognition is segmentation or image segmentation. The purpose of segmentation is to label image regions according to what is shown. The image regions determined to indicate an object or a set of objects form a segmentation mask. The segmentation result or outcome for an image includes one or more segmentation masks.

[0006] Common types of segmentation are semantic segmentation and instance segmentation. In semantic segmentation, all pixels belonging to the same object class are segmented as one object and thus part of one segmentation mask. For example, all pixels detected as a person are segmented as one object, and all pixels detected as a car are segmented as another object. On the other hand, instance segmentation aims to detect every distinct object instance in an image. For example, each person in an image is segmented as an individual object and thus forms one segmentation mask. Another known type of segmentation is panoptic segmentation, which can be described as a combination of semantic segmentation and instance segmentation. In panoptic segmentation, the goal is to semantically distinguish different objects and to detect distinct instances of different object classes in an image.

[0007] The present application and the invention disclosed herein relate to segmentation techniques that aim, at least, to detect distinct instances of one or more object classes in an image. Examples of such segmentation techniques include the aforementioned instance segmentation and panoptic segmentation.

[0008] The task of segmentation can be performed using a deep learning model configured and trained to perform instance segmentation. Instance segmentation can be part of a panoptic segmentation process or an instance segmentation process. At a schematic level, the input to such a deep learning model is image data of one or more images, such as a surveillance video image to be segmented, and the output is a tensor representing reliability scores for one or more object classes for the image data at the pixel level. In other words, the deep learning model determines, for every pixel, the probability that the pixel represents an object of each of one or more object classes.

[0009] There are many different deep learning networks that are suitable for being configured and trained into a deep learning model for performing instance segmentation, and the detailed formats of the input and output, i.e., what format is required for the input image data and what format the output tensor has, may vary among these networks. Also, note that the term "tensor" is a suitable representation for the output of a deep learning model as they are constructed today, but other future terms for the output from a deep learning model or other method or algorithm structure used for the purpose of providing instance segmentation should be considered equivalent. In other words, in future applications of the present invention, a data structure that represents reliability scores for one or more object classes for an image in some way can fulfill the same purpose as the tensor described herein and should be considered equivalent to the tensor. Therefore, the term "tensor" is interchangeable with future terms for a data structure that represents reliability scores for object classes for an image.

[0010] Outputs from deep learning models are, in some cases, interpreted together with results from additional deep learning models to form a segmentation mask for an image, i.e., a mask formed by image regions labeled with at least those object classes as detected by one or more deep learning models. One standard technique for interpretation is thresholding, where a tensor output from a deep learning model performing instance segmentation is interpreted by setting a confidence score threshold for each object class. The threshold sets the minimum confidence score required for a pixel to be interpreted as indicating the corresponding object class. Thresholding is an important task and there are many ideas on how best to set the threshold. Generally, too low a threshold results in a noisy segmentation mask containing many false object detections, and too high a threshold results in an insufficient segmentation mask that cannot contain positive object detections.

[0011] One thresholding technique is disclosed in U.S. Patent Application Publication No. 2009 / 0310822. That document discloses an object segmentation process where so-called object prediction information at the pixel level is used to adjust the confidence score threshold during segmentation. Examples of object prediction information are object motion information, object category information, environmental information, object depth information, and interaction information. According to the prediction information, a pixel is pre-determined as a predicted foreground pixel or a predicted background pixel. If a pixel is considered to be a foreground pixel, the threshold for the pixel is decreased to increase the sensitivity of the segmentation procedure. In other cases, the threshold is increased to decrease the sensitivity.

[0012] There are various solutions for providing an accurate segmentation mask, but for example, improved methods regarding processing efficiency and accuracy are still needed.

Prior Art Documents

Patent Documents

[0013]

Patent Document 1

Summary of the Invention

[0014] An object of the present invention is to provide an improved method for providing a segmentation mask that indicates individual instances of one or more object classes with respect to processing speed and segmentation accuracy, i.e., the level of accuracy provided by the resulting segmentation mask. In other words, the object is to provide a fast segmentation method that outputs a high-precision segmentation mask.

[0015] According to a first aspect, these and other objectives are achieved, wholly or at least in part, by the method defined by claim 1.

[0016] The present invention is based on an implementation where thresholding, i.e., the adjustment of the interpretation of tensors output from a segmentation deep learning model, is advantageously focused on image areas that are likely to indicate objects. These areas can be located by identifying so-called coherent regions. The thresholding is configured to finely adjust the reliability score threshold in the coherent region in order to provide a more precise object segmentation mask. Less processing can be expended on thresholding in the remaining areas that are less likely to indicate objects. Thus, potentially limited processing resources can be efficiently allocated by expending them on determining suitable reliability score thresholds for the pixels in the areas that may contain objects. The fine adjustment of the reliability score threshold in the coherent region includes generating a segmentation mask using different reliability score thresholds and evaluating the generated segmentation mask with respect to one or more object mask conditions. When a positive result is obtained, an acceptable segmentation mask is determined for the objects of the currently evaluated object class, and the reliability score threshold of that segmentation mask is retained in the final segmentation result for the entire image. Thus, in contrast to the prior art, the present invention defines a method in which the reliability score threshold is adjusted for all pixels in an image region, i.e., in a group of adjacent pixels, instead of adjusting the thresholds for individual pixels that are independent of each other. That is, multiple segmentation masks are generated for groups of pixels that form image subregions, and these segmentation masks are evaluated as entities, as opposed to independent adjustment of the reliability score threshold at the pixel level. The proposed approach provides a segmentation process with high precision in order to reduce the processing cost for the adjustment and evaluation of image subregions selected to have a high likelihood of indicating objects, as indicated by the coherent regions.

[0017] As used herein, a coherent region means a region of contiguous pixels in an image that forms a sub-region of the image. The identified coherent region corresponds to an area in the image that has moved by approximately the same amount and in the same direction. A coherent region can correspond, for example, to a moving object shown in a video. A coherent region may also be referred to as a motion-connected area.

[0018] A coherent region can be determined by comparing motion vectors for image macroblocks to find pixel regions with similar motion. By assuming that a group of adjacent motion vectors that are similar, i.e., have approximately the same direction and magnitude, probably indicates that the corresponding pixels represent the same object, the presence of such a group of motion vectors, which is referred to herein as a coherent region, is used as a guide to an image region where an object should potentially be detected by a segmentation process. Another way to find coherent regions in an image can be to identify the position of an object by image analysis or by using an external sensor such as a radar.

[0019] Looking in more detail at the process of determining a coherent region from motion vectors, these vectors can be extracted from the encoding process for an image. Motion vectors are determined during the encoding of an image as the image is processed by an encoding algorithm, and the motion vectors are determined for groups of pixels. Once the motion vectors for an image are determined, the motion vectors can be evaluated to determine one or more coherent regions. Alternatively, the motion vectors can be determined as a separate process rather than as part of the encoding process. Briefly, the motion vectors can be determined by searching a previously captured image for macroblocks, which are pixel blocks of, for example, 4×4 or 8×8 pixels, of the current image. When found, the motion vectors are determined for the macroblocks, and the motion vectors define how much and in which direction the macroblocks have moved from the previously captured image. Motion vectors are thus a well-known concept used for the purpose of temporarily encoding video, and this is part of various video compression standards such as H.264 and H.265.

[0020] A state of having approximately the same direction can be defined as the direction of the motion vectors, or the dominant part of the motion vectors, being within a predetermined span, for example, the maximum angle between any two motion vectors in a coherent region being less than a predetermined threshold angle. A state of having approximately the same magnitude can be defined as the magnitude of the motion vectors, or the dominant part of the motion vectors, being within a predetermined span, for example, the maximum difference in magnitude between any two motion vectors in a coherent region being less than a predetermined threshold. Which threshold is suitable depends on parameters such as the configuration and accuracy of the motion vector determination. Thus, the threshold can vary between implementations.

[0021] The task of finding coherent regions by evaluating motion vectors should be considered straightforward and can be solved, for example, by different known evaluation algorithms.

[0022] As used herein, an object mask state is defined as a state for a segmentation mask that indicates that the mask corresponds to the object shown. Some non-limiting examples of such states are given by the dependent claims and include a non-fragmented mask, a merged segmentation mask, and a smooth mask edge. The object mask state can be implementation-specific and thus may vary. For example, a segmentation method for an image showing a scene mainly involving humans may implement an object mask state that works particularly well for humans but perhaps not as well for other object classes, while a segmentation method for an image showing a scene with objects of varying classes may implement a mask state suitable for many different object classes. The evaluation of a segmentation mask for an object mask state can consider only that segmentation mask or also other segmentation masks. For example, a segmentation mask can be evaluated to determine whether it satisfies the object mask state by comparing how the object mask state is satisfied by corresponding segmentation masks in previous and / or subsequent images.

[0023] The segmentation results used herein mean the result of the interpretation of a tensor representing pixel-specific reliability scores. According to the present invention, a plurality of interpretation rounds of the tensor are performed on the pixels of the coherent region. Each interpretation round is represented by a temporary segmentation mask, which is then evaluated before generating a segmentation mask that should be part of the final segmentation result for the image. The temporary segmentation mask used herein can be regarded as a binary representation of the image area indicating whether a pixel is detected as indicating an object of a specific object class. A plurality of temporary segmentation masks can be formed for the same image area or for overlapping image areas, but for different object classes.

[0024] The step of processing an image to determine a tensor may include processing the image by a deep learning model configured and trained to determine a pixel-level confidence score for one or more object classes based on the image data of the image. Deep learning models configured and trained for instance segmentation or for being part of a panoptic segmentation process are common examples of such deep learning models. Non-limiting examples of deep learning networks that may be configured and trained to form a deep learning model suitable for the task of instance segmentation include YOLACT, Mask-R-CNN, fully convolutional instance-aware semantic segmentation techniques such as FCIS, and the Panoptic-DeepLab panoptic segmentation model architecture. Embodiments of the present invention are disclosed in the context of deep learning, which is a preferred embodiment as of the filing date of this application, but it should be noted that the present invention is not limited to using a deep learning model for instance segmentation. For example, if another type of machine learning model or algorithm performs instance segmentation based on the image data and provides an output representing the confidence score for the image, that machine learning model or algorithm may be used instead of the deep learning model.

[0025] The generation of a series of temporary segmentation masks can be regarded as an iterative process in which tensors are interpreted using a temporarily adjusted confidence score threshold between iterations. The adjustment of the temporarily adjusted confidence score threshold can follow a predetermined manner. For example, the adjustment can vary the temporarily adjusted confidence score threshold step by step between a start value and an end value. The start value and the end value can form a threshold span. Thus, the iterative interpretation of the tensor can include always increasing or always decreasing the temporarily adjusted confidence score threshold between iterations. An advantage of this approach is that the evaluation of a series of temporary segmentation masks can include evaluating trends in the series of temporary segmentation masks, for example, that the mask is expanding, which can lead to a faster evaluation of whether the object mask state will be or is being satisfied.

[0026] The starting point for adjusting the reliability score threshold, i.e., the initial reliability score threshold of the first generated mask in a series of temporary segmentation masks, can be selected according to different embodiments. As described above, the temporary reliability score threshold can be adjusted between a starting value and an ending value, and thus between a maximum value and a minimum value, or vice versa. An alternative approach is to first generate the first temporary segmentation mask in the series using an initial reliability score threshold, e.g., a base reliability score threshold, and then evaluate the spatial relationship between the first temporary segmentation mask and the coherent region. If the first temporary segmentation mask is smaller than the coherent region, it is more likely that the first temporary segmentation mask does not cover all object pixels. A second temporary reliability score threshold can be determined to include more pixels compared to the initial reliability score threshold, and a second temporary segmentation mask is generated using the second reliability score threshold. Conversely, if the first temporary segmentation mask is larger than the coherent region, the first temporary segmentation mask may cover too many pixels. A second temporary reliability score threshold can be determined to include fewer pixels compared to the initial reliability score threshold, and a second temporary segmentation mask is generated using the second reliability score threshold. Thus, the choice of generating a series of temporary segmentation masks, which involves always increasing or always decreasing the temporary reliability score threshold, can be based on the comparison between the first temporary segmentation mask in the series of temporary segmentation masks and the coherent region.

[0027] A series of temporary segmentation masks is preferably generated by interpreting a tensor for a single, arbitrarily selected object class. That is, the temporary segmentation masks represent objects of a single object class according to tensors interpreted at different confidence score thresholds. The method can be performed in parallel for different object classes, and a final confidence score threshold for each object class can be determined. Before generating the final segmentation result, the segmentation process needs to select which objects should be included in the final segmentation result and which ones should be discarded. Such filtering is a known part of the segmentation process.

[0028] If necessary, when generating a series of temporary segmentation masks, which object class to select can be determined by analyzing the tensor. The selection may be required or desired to save time and / or processing resources. In one embodiment, a single object class is determined by identifying the object class with the highest total confidence score for pixels in a coherent region. Alternatively, the highest number of pixels identified as a particular object class can set the single object class. The historical object classes determined for the coherent region can also be considered when determining the single object class. For example, if in a previous image the object class of a horse was determined for pixels in a coherent region, the determination of the single object class can be adjusted to make the object class of the horse more likely to be selected. For example, a higher weight can be imposed on the confidence score of the horse in the tensor compared to the weights for other object classes. Alternatively, the confidence score threshold for the object class of the horse can be offset to a lower value compared to the confidence score thresholds for other object classes.

[0029] According to one embodiment, a series of temporary segmentation masks are generated for an image region composed of a coherent region and a surrounding margin area. Thus, a series of temporary segmentation masks are generated for the coherent region and, additionally, for a relatively small surrounding margin area, but not yet for the entire image. An advantage of this embodiment is that the generated temporary segmentation masks can extend outside the coherent region.

[0030] In one embodiment, the final reliability score is reused in a subsequent segmentation process for the next image in the sequence of images. For example, the final reliability score can be used as an initial reliability score threshold for generating a first temporary segmentation mask in the coherent region of a second image, determined to correspond to the same object as the coherent region of the previous image for the first image. For this to work, the coherent region of the next second image needs to be analyzed to determine whether that coherent region was caused by the same object as the coherent region of the first image. Moreover, a condition that the coherent regions should be similar, i.e., have approximately the same spatial position and size, can also be implemented to ensure that the coherent regions were caused by the same object. A multi-object tracking algorithm can assist in the determination as objects can be tracked and identified relative to each other. The results of the multi-object tracking algorithm can be used to verify whether the detected objects in the coherent regions of two subsequent images are the same.

[0031] Thus, according to a second aspect, the present invention is a method of generating a segmentation mask that indicates individual instances of one or more object classes for an image in a sequence of images, as defined by claim 9.

[0032] According to another embodiment, the final reliability score is reused in the subsequent generation of the segmentation mask for the next image in the image sequence even if the coherent region is not identified in the next second image. The object may still be shown in the second image, but the object is stationary, i.e., not moving, and thus may not give rise to a coherent region. Thus, if there is a coherent region in a previous image, for example, in any of the 10 previous images, the final reliability score threshold may be used for the same pixels in the second image as in the previous image. In this way, an improved mask for a stationary object can be achieved with the help of the segmentation mask generated previously for the same object when it was moving.

[0033] The method disclosed herein can advantageously be implemented in the processing device of a camera. The final segmentation result can then be transmitted by the camera, together with the image in an encoded format.

[0034] According to a third aspect, the invention is an image capture device configured to generate a segmentation mask indicating individual instances of one or more object classes for an image in an image sequence, as defined in claim 12. The image capture device of the third aspect can generally be implemented in the same way as the method of the first aspect with the attendant advantages.

[0035] According to a fourth aspect, the invention is a computer-readable storage medium comprising computer code that, when loaded and executed by one or more processors or control circuits, causes the one or more processors or control circuits to perform the method according to the first aspect.

[0036] The further scope of applicability of the present invention will become apparent from the following forms for carrying out the invention. However, it should be understood that various changes and modifications within the scope of the present invention will be apparent to those skilled in the art from the forms for carrying out the invention, so the forms for carrying out the invention and the specific examples are given as illustrative only and merely show preferred embodiments of the present invention.

[0037] Accordingly, it should be understood that the devices described and the methods described may vary, so the present invention is not limited to specific component parts of such devices or steps of such methods. It should also be understood that the terms used herein are for the purpose of describing specific embodiments only and are not limiting. It should be noted that, as used in this specification and the appended claims, the articles "a", "an", "the", and "said" are to be construed to mean one or more of the elements unless the context clearly dictates otherwise.

[0038] Next, by way of example and with reference to the accompanying schematic diagrams, the present invention will be described in more detail.

Brief Description of the Drawings

[0039]

Figure 1

Figure 2

Figure 3a

Figure 3b

Figure 3c

Figure 4

Figure 5

Best Mode for Carrying Out the Invention

[0040] FIG. 1 shows a camera 10 having a configuration suitable for implementing a method for generating a segmentation mask for a collected image according to an embodiment. The camera 10 is a digital camera that can be adapted for monitoring purposes. The camera 10 can be, for example, a fixed surveillance camera configured to generate a video of a scene viewed by a remotely located user. The camera 10 includes an image sensor 101 for collecting images according to known digital imaging techniques, and a transmitter 106 for transmitting a video, i.e., one or more image sequences, to a receiver via wired or wireless communication. Before transmitting the video, the camera 10 adjusts and processes its collected image data. An Image Processing Pipeline (IPP) 103 performs known pre-encoding enhancements of the collected image data, such as gain adjustment and white balance adjustment. An encoding processor 102 is provided for the purpose of encoding the raw image data or the pre-encoded processed image data. The encoding processor 102 is adapted to encode the image data using a video compression algorithm based on predictive coding. Non-limiting examples of suitable video compression algorithms include the H.264 and H.265 video compression standards. Predictive coding, or encoding temporarily, includes intra-frame (I-frame) coding and inter-frame (P-frame or B-frame) coding. During predictive coding, motion vectors are determined in a known manner, generally at the macroblock level. A motion vector represents the position of a macroblock in one image with respect to the position of the same or a similar macroblock in a previously collected image. The macroblock size varies between different video compression standards. For example, a macroblock can be formed by a group of pixels of adjacent pixels of 8×8 or 16×16. The output from the encoding processor 102 is the encoded image data.

[0041] Camera 10 further includes a segmentation processor 104 configured to perform image segmentation including instance segmentation. That is, the image segmentation processor 104 is adapted to process an image, more particularly raw image data or pre-encoded image data, in order to generate a segmentation mask that indicates individual instances of one or more object classes for the captured image. The segmentation can be performed for all of the images in the image sequence or for selected images in the image sequence.

[0042] The image segmentation processor 104 and the encoding processor 104 can be implemented as software, and circuitry forms respective processors, e.g., a microprocessor, which, in relation to computer code instructions stored in a memory 105 that is a (non-transitory) computer-readable medium such as a non-volatile memory, causes the camera 10 to perform any of the methods (partially) disclosed herein. Examples of non-volatile memory include read-only memory, flash memory, ferroelectric RAM, magnetic computer storage devices, optical disks, and the like.

[0043] Camera 10 may include additional modules that serve other purposes not related to the present invention.

[0044] Note that the segmentation processor 104 is not required to be an integral part of the camera 10. In an alternative embodiment, the segmentation processor 104 is located remotely with respect to the camera 10. Thus, the image segmentation can be performed remotely, e.g., on a server connected to the camera 10, and the camera 10 can transmit the image data to the server and receive the image segmentation results from the server.

[0045] Next, an overview of one embodiment is provided with further reference to FIG. 2. FIG. 2 shows a flowchart of image processing related to image segmentation and encoding in camera 10, and the interaction between them, according to one embodiment. Image data is collected by image capture 201 and optionally pre-encoded. The image data is provided to a segmentation processor 104 for image segmentation and an encoding processor 102. In the encoding processor 102, video compression 102 is performed along with other encoding steps not shown here. Video compression 102 includes determining motion vectors as part of predictive coding. For encoding purposes, motion vectors are used to represent the image sequence in a bit-efficient manner. Further, according to an embodiment, motion vectors are also used in segmentation. Thus, the motion vectors from video compression 202 are extracted for segmentation, particularly for the thresholding step 204 in segmentation. The purpose of thresholding 204 is to determine a reliability score threshold to be used to generate the final segmentation result for the image. The reliability score threshold sets a boundary value between the reliability score to be interpreted as a positive detection and the reliability score to be interpreted as a negative detection. The reliability score to be interpreted is given in the tensor output from the previous segmentation 203. As described, segmentation 203 includes processing the image data in a known manner to generate and determine the reliability score. The reliability score is determined at the pixel level and can be, for example, a value between 1 and 100 indicating the probability that a pixel represents an object of a certain object class. Segmentation 203 can determine reliability scores for multiple object classes, which means that each pixel can be given multiple reliability scores. Depending on the algorithm used in segmentation 203, the output format, i.e., how the reliability scores are represented, can vary.In this application, the term tensor represents any output from a segmentation algorithm that can be used for the purposes of the invention. A tensor represents one or more confidence scores per pixel in an image, where the confidence scores are given for one or more object classes. Non-limiting examples of object classes include vehicle, human, car, foreground, background, head, and bicycle. Thus, a segmentation algorithm may segment at a general level and find individual instances of an object, for example, without further determining the type of object, or the segmentation algorithm may segment at a more specific level and find individual instances of, for example, bicycles and cars.

[0046] Returning to the step of threshold processing 204, by using the motion vectors extracted from video compression 202, embodiments provide a way to determine the reliability score threshold for an image in an efficient manner. First, the motion vectors are analyzed to determine one or more coherent regions of the image underlying the segmentation. The analysis includes evaluating the motion vectors to identify adjacent motion vectors, i.e., the motion vectors of adjacent macroblocks having similar directions and similar magnitudes, and thus defining coherent regions. The identified coherent regions point to portions of the scene that are moving in a coherent manner, such as an image region showing a walking person, and the remaining portions of the image may not indicate moving objects. This information is used in threshold processing 204 to guide the process regarding where to exert effort in finding the reliability score threshold that provides a precise mask. Specifically, threshold processing 204 applies an iterative search for suitable reliability score thresholds in these identified coherent regions, since the identified coherent regions are more likely to indicate objects than the remaining areas. Less processing is expended in finding the reliability score threshold for the remaining image areas, which may be treated as background regions and assigned a default reliability score threshold or may not be processed at all in image segmentation. Thus, according to one embodiment, the method may assume that there are objects only in the coherent regions, and thus may not expend resources in attempting to segment, i.e., determine instances of object classes, in image areas outside the coherent regions. In one embodiment, the set of identified coherent regions for an image is preprocessed before being used in threshold processing 204. The purpose of the preprocessing is to filter out the relevant coherent regions and discard the coherent regions that may be caused by unrelated objects or movements in the scene.The preprocessing may include comparing a set of coherent regions with a set of segmentation masks identified by interpreting a tensor using one or more base confidence score thresholds. Different base confidence score thresholds may be used in different regions of the image. The base confidence score threshold may be a predetermined value or may be adjusted dynamically during image capture. For example, the base confidence score threshold may take the same value as the confidence score threshold used for a spatially corresponding image area in a previously captured image, preferably the most recently captured image. The set of coherent regions is filtered such that coherent regions that overlap at least partially, optionally to an extent above a threshold, with a segmentation mask in the set of segmentation masks are retained, and coherent regions that do not overlap or overlap to an extent below the threshold are discarded and removed from the set of coherent regions. The remaining coherent regions, sometimes referred to as relevant coherent regions in the set of coherent regions, are then used for the thresholding 204 disclosed herein. Although the term relevant coherent region is not referred to in the remainder of the description, it should be understood that the optional preprocessing of coherent regions disclosed above for filtering out relevant coherent regions may be used in any of the disclosed embodiments.

[0047] Next, an iterative search for suitable reliability score thresholds in one or more determined coherent regions is described in more detail with further reference to FIGS. 3a - 3c. FIG. 3a shows an image sequence 30 obtained in image capture 201. The image 31 to be segmented shows a moving car and three living beings. Two of the living beings (on the right side) are stationary and one living being (on the left side) is moving. During the encoding of the image 31, motion vectors pointing to other images in the image sequence 30 are determined. As described, the motion vectors determined for the image 31 are analyzed to identify the coherent regions 32a, 32b shown in FIG. 3b. Since the coherent regions 32a, 32b are determined at the pixel group level, generally at the macroblock level, those coherent regions provide a schematic indication of the image areas showing coherent movement. The evaluation of motion vectors is the currently preferred way to determine coherent regions, but it should be noted that there are alternative ways to determine these regions. For example, the coherent regions 32a, 32b can be identified by performing image analysis of the image 31. Known image analysis algorithms for object detection or motion detection can be used for the purpose of finding coherent regions. Another alternative is to use sensors in addition to the image sensor. The sensors can be, for example, radar or other distance sensors. The coherent regions can be determined by identifying the moving objects by the sensors and determining the corresponding spatial coordinates of those objects in the image. Both of these alternative forms, namely, determining the coherent regions by image analysis or by using one or more additional sensors, can be implemented by those skilled in the art without further detail.

[0048] For each coherent region, the thresholding process 204 generates a series of temporary segmentation masks. FIG. 3c shows a series of temporary segmentation masks 33 generated for the coherent region 32b. Each mask in the series of temporary segmentation masks 33 is generated by interpreting the tensor output from the segmentation 203 using different temporary reliability score thresholds. According to the interpretation of the tensor, the black pixels of the segmentation mask indicate that the pixel represents an object and forms part of the temporary segmentation mask, and the white pixels indicate that the pixel does not represent an object and is not part of the temporary segmentation mask. The temporary segmentation mask can be generated by interpreting tensors in the coherent region and also in the surrounding regions, as shown. All black pixels in the evaluated region, i.e., all positive interpretations, are part of the temporary segmentation mask in this context. Note that the temporary segmentation mask can be composed of separate mask fragments.

[0049] A series of temporary segmentation masks 33 are generated for each object class, which means that the tensor is evaluated for a single object class. In the illustrated embodiment, the series of temporary segmentation masks 33 are generated for the object class of the track, which means that those masks are generated by interpreting the reliability scores for the object class track in the tensor. A first temporary segmentation mask 34 is generated using a first initial temporary reliability score threshold. A second temporary segmentation mask 35 is generated using a second temporary reliability score threshold with the threshold lowered to interpret the reliability score as a positive detection. An Nth temporary segmentation mask 36 is generated using a further lowered Nth temporary reliability score. An (N + 1)th temporary segmentation mask 37 is generated using an (N + 1)th temporary reliability score that is further lowered compared to the Nth temporary reliability score. As indicated in the figure, the series of temporary segmentation masks 33 comprises a temporary segmentation mask between the second temporary segmentation mask 35 and the Nth temporary segmentation mask 36.

[0050] The temporary reliability score thresholds used in the series of temporary segmentation masks 33 follow a decreasing pattern in this embodiment. Thus, for each generated temporary segmentation mask, the reliability score threshold is adjusted to lower the threshold to interpret the reliability score as a positive detection.

[0051] In addition to generating a series of temporary segmentation masks 33, the thresholding process 204 performs an evaluation of those masks to determine whether the object mask state is satisfied by any of these masks. The object mask state is a predetermined state for a segmentation mask to be considered to represent an object. Thus, by evaluating whether any of the temporary segmentation masks 33 satisfy the object mask state, the hypothesized presence of an object indicated by the coherent region can be verified or discarded. The object mask state defines one or more characteristics of the temporary segmentation mask 33. Non-limiting examples of characteristics include not fragmented, which means that the mask does not consist of multiple separate mask fragments and smooth mask edges. The smoothness of the mask edge can be given by the curvature of the mask edge. The object mask state can be defined as the maximum allowable curvature of the mask edge. Alternatively, the object mask state can be defined as the maximum allowable deviation or tolerance in curvature for the mask edge. The object mask state can be specific to an object class, which means that the object mask state for verifying a vehicle can be different from the object mask state for verifying a living being.

[0052] A series of temporary segmentation masks 33 can be evaluated during the generation of the series or when the complete series of temporary segmentation masks 33 has been generated. Further, the series of temporary segmentation masks can be evaluated at an individual level or at a group level. For example, in the illustrated embodiment, each mask 34, 35, 36, 37 in the series of temporary segmentation masks 33 can be evaluated individually to determine whether any of them have smooth mask edges defined by an object mask state defined for the object class of the track. Alternatively, the masks 34, 35, 36, 37 can be evaluated to determine whether the masks 34, 35, 36, 37 are created from separate fragments that merge into a merged mask over part or the entire series 33. An example of merging fragments is provided in FIG. 3c, where the first temporary segmentation mask 34 comprises separate fragments and the second, later-generated temporary segmentation mask 35 comprises fewer fragments compared to the previously generated mask 34. The object mask state can, in this example, include the detection of merging fragments, i.e., reducing the number of fragments, in the series of temporary segmentation masks 33 where the temporary reliability score threshold is adjusted to lower the threshold for positive detection. In yet another alternative, the masks 34, 35, 36, 36 can be evaluated to find the best fit to the object mask state. For example, the Nth mask 36 can be found to have the greatest best fit to the mask object state defined by smooth mask edges compared to the previous mask and the subsequent (N + 1)th mask 37. Thus, the evaluation can be described as finding the optimal conditions for meeting the object mask state.

[0053] When the coherent region 32b is evaluated with respect to the object mask state, a final reliability score threshold is set for the pixels of the coherent region 32b, or for a subset of the pixels or macroblocks therein. If the evaluation of the temporary segmentation mask is successful and thus the mask is found to satisfy or meet the object mask state, the final reliability score threshold is set to the temporary reliability score threshold of the mask that satisfies the object mask state. The final reliability score threshold is then set for the pixels of the temporary segmentation mask that satisfies the object mask state. If two or more temporary segmentation masks that satisfy the object mask state are found, a selection must be made as to which temporary segmentation mask and corresponding temporary reliability score threshold. The selection may include determining and selecting the mask that best satisfies the object mask state or the mask generated using the temporary reliability score threshold that represents the lowest threshold for a positive detection.

[0054] However, if the object mask state is not satisfied by any one of the temporary segmentation masks, the final reliability score threshold is set to the default reliability score threshold. The default reliability score threshold can be a predetermined fixed threshold or the same threshold as determined in the segmentation of the coherent region in the previous, preferably immediately preceding, image. The predetermined fixed threshold can be the same as for the pixels or surrounding regions with respect to the coherent region.

[0055] As illustrated, the final reliability score threshold can be temporarily stored for use in subsequent image segmentation. The final reliability score threshold for the first image can be used as the initial temporary reliability score threshold in the segmentation of the second, later-collected image. In another embodiment, the final reliability score is applied in an image region of the second image corresponding to the coherent region of the first image, even if the coherent region is not determined in that image region. Thus, an object that has moved in the first image and is identified by the coherent region can be successfully segmented even if the object does not move in the second image and thus no detection of the coherent region occurs.

[0056] In yet another embodiment, the final reliability score threshold for the first coherent region of the first image is used when segmenting a second, later image in which a second coherent region is detected. In this embodiment, it is evaluated whether the first coherent region and the second coherent region were caused by the same object. The evaluation can include analyzing the similarity in the motion vectors of the coherent regions, analyzing the spatial relationship between the coherent regions, or utilizing a separate tracking algorithm for determining and tracking individual objects in the image, such as a multi-object tracking algorithm. According to the tracking algorithm, it can be determined that the coherent regions were caused by the same object by determining whether an object having the same identification information is present in the coherent regions. If it is determined that this has occurred by either the illustrated method or another evaluation method, the set of final reliability score thresholds for the coherent region or a subset of pixels therein of the first image can be used as the initial reliability score threshold when generating a first temporary segmentation mask for the coherent region of the second image. An advantage of this embodiment is that the temporary segmentation mask for the coherent region of the second image can more quickly satisfy the object mask state by generating a temporary segmentation mask that generates a mask already found suitable for the object shown, by starting the generation of the temporary segmentation mask.

[0057] A process for finding suitable final reliability score thresholds is performed for all the coherent regions 32a, 32b determined in image 31. The process can also be performed multiple times for a single coherent region 32a, 32b with respect to different object classes in order to determine suitable final reliability score thresholds for each object class. The final reliability score thresholds for different image areas and different object classes are provided for mask creation 205 for the purpose of generating the final segmentation mask for image 31. Mask creation 205 performs mask creation for the entire image, not just for the coherent regions. For the image areas outside the coherent regions, segmentation can be performed by interpreting the tensors using a reliability score threshold set as a standard or, for example, based on the threshold used in the previous image. Mask creation 205 functions according to known principles for generating the final segmentation result. Different known algorithms and conditions can be applied to select which object class the image area is most likely to represent based on the received final threshold. Moreover, setting the spatial boundaries between the segmentation masks of different object classes can also be a task for mask creation 205. Thus, according to known methods and based on the final reliability score threshold and the tensors, the final segmentation result is determined and provided for output creation 206. Output creation 206 also receives the encoded image data from video compression 202 and creates the output from the image processing of camera 10. The output format of the encoded image data and the final segmentation mask follow conventional standards. For example, the encoded image data can be sent from camera 10 in the form of a video stream, and the final segmentation mask can be sent as metadata.

[0058] FIG. 4 shows an example of a created image 41 with a final segmentation result that includes a first segmentation mask 44 representing a moving creature in image 31 and a second segmentation mask 42 representing a moving vehicle in image 31. The first segmentation mask 44 and the second segmentation mask 42 are determined by using the methods disclosed herein, i.e., by iteratively determining temporary segmentation masks to find a reliability score threshold process suitable for use. The final segmentation result also includes segmentation masks 45, 46 determined through segmentation by interpreting the tensor using a standard or base reliability score threshold.

[0059] FIG. 5 provides a schematic overview of a method 5 for generating a final segmentation result for an image according to one embodiment, and each step of method 5 has been described and illustrated above.

[0060] The image, i.e., the image data of the image, is processed at 501 to determine the coherent region. The coherent region can be determined, for example, in the encoder process (in the encoder processor) or in the segmentation process (in the segmentation processor). The image is also processed at 502 to determine a tensor representing the pixel-specific reliability score for one or more object classes. Step 501 and step 502 can be performed in parallel or sequentially. It is not important which of step 501 and step 502 is performed before the other. Thresholding in the image segmentation process is performed when both step 501 and step 502 have been performed, i.e., when both the coherent region and the tensor required for thresholding are available for thresholding. If one of step 501 and step 502 is completed before the other, the result of the first completed step is locally stored, for example, in the memory of the camera, and can be retrieved by the segmentation processor when the result of the second completed step is available. Next, method 5 includes step 503 of generating a series of temporary segmentation masks for each of the one or more coherent regions. As previously explained, the one or more coherent regions determined at step 501 may have been processed to remove related coherent regions with a filter. In that case, step 503 of generating a series of temporary segmentation masks is performed for each of the one or more related coherent regions.

[0061] A series of temporary segmentation masks are evaluated according to the method described at 504. Method 5 further includes setting, at 505, a final reliability score threshold for the pixels of the temporary segmentation mask or for the coherent regions, based on the results of the evaluation at 504. For some of the coherent regions, the final reliability score threshold is set for each area. Additionally, one or more final reliability score thresholds may be set for the pixels of the remaining image areas that are not part of the coherent regions. Method 5 then includes, at 506, generating a final segmentation result for the image based on one or more final reliability scores. As described, generating the final segmentation result may include known methods for evaluating segmentation masks of different object classes to select which segmentation masks the final segmentation result should include.

Claims

1. A method for generating a segmentation result that indicates individual instances of one or more object classes for an image in an image sequence, the method comprising: a. Determining a coherent region of the image, the coherent region being an area in the image that has moved in approximately the same amount and in approximately the same direction; b. Processing the image to determine a tensor representing pixel-specific reliability scores for one or more object classes; c. Generating a series of temporary segmentation masks for the coherent region, each temporary segmentation mask being generated by interpreting the tensor for a single object class using different temporary reliability score thresholds; d. Evaluating the series of temporary segmentation masks to determine whether an object mask state is satisfied by one or more of the temporary segmentation masks, the object mask state being a state that indicates that a temporary segmentation mask corresponds to an object; e. If the object mask state is satisfied by one or more of the temporary segmentation masks, setting the temporary reliability score threshold used to generate one of the one or more temporary segmentation masks as the final reliability score threshold for the pixels of the temporary segmentation mask; f. If the object mask state is not satisfied, setting a default reliability score threshold as the final reliability score threshold for the coherent region; g. Generating a final segmentation result for the image, wherein a portion of the final segmentation result covering the coherent region is generated by interpreting the tensor using the final reliability score threshold, generating a final segmentation result A method comprising.

2. The method according to claim 1, wherein step a includes determining a coherent region of the image of adjacent pixels or pixel groups having motion vectors in substantially the same direction and substantially the same magnitude.

3. The method according to claim 2, wherein step a includes processing the image by an encoding algorithm to determine a motion vector for a pixel group.

4. The object mask state is a state in which the temporary segmentation mask defines an unfragmented object, and a state in which temporary segmentation mask fragments merge The method according to claim 1, including at least one of.

5. The method according to claim 1, wherein step b includes processing the image to determine a tensor by a deep learning model.

6. The method according to claim 1, wherein the series of temporary segmentation masks is generated by repeatedly interpreting the tensor using a temporary reliability score threshold that always increases or always decreases between iterations.

7. Determining whether the first temporary segmentation mask generated using the initial reliability score threshold is larger or smaller than the coherent region, Selecting the temporary reliability score threshold to always increase according to the fact that the first temporary segmentation mask is larger than the coherent region, or selecting the temporary reliability score threshold to always decrease according to the fact that the first temporary segmentation mask is smaller than the coherent region The method according to claim 6, further comprising

8. The method according to claim 1, wherein the single object class is selected by identifying the object class having the highest total reliability score for the pixels in the coherent region

9. The method according to claim 1, wherein step a includes generating the series of temporary segmentation masks for an image region composed of the coherent region and a surrounding margin area

10. A method for generating a segmentation mask indicating individual instances of one or more object classes for an image in a sequence of images, the method comprising Performing the method according to claim 1 for a first image Determining a coherent region in a second image Evaluating whether the coherent region of the second image was caused by the same object as the coherent region of the first image Performing steps b to g according to claim 1 for the second image, wherein when the coherent region of the first image and the coherent region of the second image are caused by the same object, the final reliability score threshold for the coherent region of the first image is the first temporary segmentation mask in the series of temporary segmentation masks Performing steps b to g according to claim 1 for the second image used to generate Comprising a method.

11. The method according to claim 10, wherein the step of evaluating whether the coherent region of the second image was caused by the same object as the coherent region of the first image includes processing the first image and the second image by a multi-object tracking algorithm.

12. The method according to claim 1, wherein the method is implemented in a processing device of a camera.

13. An image capture device configured to generate a segmentation result indicating individual instances of one or more object classes for images in a sequence of images, the image capture device comprising: One or more image sensors and an image processor configured to collect the sequence of images; An encoder; A processor adapted to implement the method according to claim 1 and an image capture device.

14. A computer-readable storage medium comprising computer code that, when loaded and executed by one or more processors or control circuits, causes the one or more processors or control circuits to implement the method according to claim 1.

Citation Information

Patent Citations

  • Model training and instance segmentation method, device and system and storage medium

    CN108875732A

  • Video object detection and segmentation method based on space-time double-branch network

    CN110097568A

  • Feedback object detection method and system

    US20090310822A1

  • System and method for semantic segmentation of images

    US20190057507A1

  • System and method for training a neural network for visual localization based upon learning objects-of-interest dense match regression

    US20200364509A1