Method, data processing system, computer program product, and computer-readable medium for object segmentation

The method uses a neural network to parameterize object contours as a closed two-dimensional curve with a Fourier descriptor, addressing high computational issues in existing image segmentation methods, enabling efficient and accurate segmentation of complex shapes for real-time applications.

JP7704833B2Active Publication Date: 2025-07-08AIMOTIVE KFT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023502950
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-17
Filing Date
2020-12-16
Publication Date
2025-07-08
Estimated Expiration
2040-12-16

AI Technical Summary

Technical Problem

Existing image segmentation methods, particularly for autonomous driving and medical applications, face high computational requirements and inefficiencies in processing complex or concave-shaped object contours, which are critical for real-time decision-making.

Method used

A method using a machine learning system, preferably a neural network, to parameterize the object's contour as a closed two-dimensional curve with a compact representation, such as a Fourier descriptor, allowing for efficient reconstruction of complex shapes by estimating geometric transformations and reducing computational complexity.

Benefits of technology

Enables accurate and fast segmentation of objects with complex or concave shapes, reducing computational requirements and improving decision-making speed in safety-critical applications like autonomous vehicles and medical imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704833000001
    Figure 0007704833000001
  • Figure 0007704833000002
    Figure 0007704833000002
  • Figure 0007704833000003
    Figure 0007704833000003
Patent Text Reader

Abstract

The present invention relates to a method for object segmentation in an image, comprising inputting the image into a trained machine learning system and reconstructing a segmentation contour of the object. The method further comprises estimating, by the trained machine learning system, a representation of the segmentation contour of the object in the image, the segmentation contour being a closed two-dimensional parametric curve, each point defined by two coordinate components that are parameterized together, and the reconstruction of the object's segmentation contour is performed from the estimated representation of the segmentation contour. The present invention also relates to a data processing system, a computer program product, and a computer-readable medium for performing the above method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for object segmentation in an image. The present invention also relates to a data processing system, a computer program product, and a computer-readable medium for implementing the method.

Background Art

[0002] In modern computer vision, image understanding is generally considered through specific tasks such as object detection and semantic or instance-level segmentation, in other words, object segmentation. In object detection, in the case of autonomous driving applications, the position of an object or object instance (i.e., a specific sample / type of object within an object category) in an image, for example, an individual car, pedestrian, traffic sign, is usually predicted as the pixel coordinates of a box (rectangle) around the object, which is usually called a bounding box. On the other hand, the semantic or instance segmentation task aims to label the entire image at a high density at the pixel level by specifying object categories and / or specific instances of all pixels. In particular, the task of instance segmentation in an image is to label each pixel using an identification tag, number, or code of the instance to which the pixel belongs. As a result, a mask is provided for each object that marks the pixels related to the object in the image. This type of representation provides a more accurate description of the position, size, and shape of the visible objects in the scene than the commonly used bounding box (or bounding rectangle) representation.

[0003] The pixel-level segmentation method is disclosed in US10,067,509 B1 for detecting occluding objects. The method performs pixel-level instance segmentation by predicting, for each pixel, a) the semantic label of various target categories (e.g., cars and pedestrians), and b) a binary label indicating whether the pixel is a contour point. Individual instance masks can be recovered by separating the pixels of the category using the predicted contours.

[0004] The above technical solution is extended in US10,311,312 B2, where two separate classifiers are trained to handle static and dynamic cases separately. The dynamic classifier is used when the tracking of a particular vehicle is successful for a number of video frames, and otherwise, the static classifier is applied to individual frames. A pixel-level approach similar to the above documents is used for segmentation.

[0005] Also, US2018 / 0108137 A1 discloses an instance-level semantic segmentation system, where the approximate position of a target object in an image is determined by predicting a bounding box around each object. In a second step, a pixel-level instance mask is then predicted using the above bounding box of each object instance.

[0006] The main drawback of the pixel-level segmentation method is its high computational requirements and associated time consumption. In certain aspects of the segmentation task, such as in the case of autonomous vehicles, the speed of recognition is critical. Methods that require excessive computational power to produce results in real time or are simply too slow are not suitable for such applications.

[0007] The approach to accelerating the calculation speed leads to the following technical solutions, among which a smaller map (instance map) is generated, that is, at a lower resolution, and then the map is enlarged according to the size of the image.

[0008] One example is the publication "Mask R-CNN" (2017) by K. He et al., which discloses a two-stage approach for object instance segmentation. First, an object proposal step is applied to roughly localize all object categories or instances in the image. Then, the problem of instance segmentation in the second step is defined as a pixel labeling task, and the binary pixels of the instance segmentation mask are directly predicted on a grid of a fixed size (e.g., 14×14 pixels). Here, the binary 1 in the mask indicates the pixel position of the corresponding object. Then, the predicted mask is deformed / resized to the appropriate position and size of the object. The drawback of this solution is that even such a small grid requires a very complex neural network with at least an output dimension of 14×14 = 196. The amount of nodes and weighting coefficients slow down the segmentation, and moreover, the generated small map has to be enlarged and interpolated according to the size of the whole image, further reducing the speed and efficiency of this method.

[0009] A similar method is disclosed in US2009 / 0340462 A1, where a neural network is used to identify the pixels of prominent objects in the image. First, the resolution of the image is reduced, and the neural network is applied on this reduced image to identify the pixels of the main objects in the image, based on which the pixels belonging to the main objects in the original full-resolution image are identified.

[0010] The drawback of the above technical solutions is that an additional step is required to determine the contours or pixels of the objects in the image, which requires additional computing power and time.

[0011] Another approach to segmentation is to approximate the contour of an object with a polygon, where the polygon is preferably predicted by a trained neural network instead of the exact contour of the object. This approach significantly reduces the computation time and requirements compared to pixel-level segmentation techniques.

[0012] In the literature “Annotating Object Instances with a Polygon-RNN” by L. Castrejon et al. (The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 5230-5238), the authors propose a solution to represent an instance segmentation mask by a polygon that outlines the instance. The vertices of the polygon are continuously reconstructed one by one by a regression neural network. An extension of this approach by the same research group is “Polygon-RNN++” (2018). The drawback of this solution is that the computation is slower because the regression neural network has a complex structure.

[0013] A single-stage approach is introduced in the literature "FourierNet: Compact mask representation for instance segmentation using differentiable shape decoders" by N. Benbarka et al. (arXiv:2002.02709 [cs.CV], 2020). This literature discloses a single-stage segmentation method as opposed to a two-stage segmentation method. This approach represents the object's contour by a set of points that are the intersections of virtual rays emerging from near the centroid of the contour with the contour, which is a parameterization of a single component of the contour. If more intersections exist for a single ray, the intersections farther from the centroid are selected. A neural network is used to predict the Fourier coefficients (Fourier descriptors) of the set of points representing the contour, and the contour is reconstructed by inverse Fourier transform. However, the process used in this method, on the one hand, limits the complexity of the shape being modeled and, on the other hand, reduces the information present in the ignored contour coordinates. The biggest drawback of this method is that the contour of an object with a concave shape can only be approximated by the envelope of the object's contour without being accurately predicted or reconstructed. However, for certain applications, an accurate shape or contour reconstruction is required.

[0014] As far as known approaches are concerned, there is a need for a method capable of performing segmentation of an object in an image for any object having a contour that includes a concave-shaped contour. SUMMARY OF THE INVENTION

[0015] The main object of the present invention is to provide a method for object segmentation in an image, the method having none of the drawbacks of prior art approaches to the extent possible.

[0016] The present invention aims to provide a method capable of segmenting an object in an image in a more efficient way than prior art approaches, in order to enable the segmentation of an object having any shape or contour. Accordingly, the present invention aims to provide a reliable segmentation method capable of reconstructing the contour of an object having any shape in an image.

[0017] A further object of the present invention is to provide a data processing system including means for performing the steps of the method according to the invention.

[0018] Furthermore, the present invention aims to provide a non - transitory computer program product for performing the steps of the method according to the invention on one or more computers, and a non - transitory computer - readable medium including instructions for performing the steps of the method according to the invention on one or more computers.

[0019] The object of the present invention can be achieved by the method according to claim 1. The object of the present invention can further be achieved by the data processing system according to claim 13 and the non - transitory computer program product according to claim 14 and the non - transitory computer - readable medium according to claim 15 The preferred embodiments of the present invention are defined in the dependent claims.

[0020] The main advantage of the method according to the present invention compared to prior art approaches stems from the fact that it can reconstruct the contour (segmentation contour) of an object having any shape, including complex shapes and even concave shapes. Since the position of the object can be determined with higher precision by this method, more accurate object segmentation can be achieved than by any method known in the prior art.

[0021] By using the parameterization of two coordinates of the contour, it has been recognized that an accurate representation of any two-dimensional closed curve, that is, the complex contour of an object in an image, can be achieved without ambiguity. The segmentation method is frequently used in the decision-making process, for example, in autonomous driving applications, where the speed of decision-making can be critical. A common choice to facilitate the decision-making process is to use a predetermined simple shape that can be easily and quickly recognized even from a small number of feature points. In contrast to this approach, the method according to the present invention is suitable for recognizing any complex shape. Determining any complex shape can increase the computational requirements of the method, but it has also been recognized that it increases the accuracy of the decision-making process based on the detected contour, which is desirable in various safety-critical applications such as those related to autonomous vehicles or medical applications. Furthermore, to balance the accuracy and computational efficiency of the method, the parameterization of the segmentation contour according to the present invention provides flexibility and control.

[0022] Also, in order to reduce the computational requirements for estimating the contour representation, instead of the simple two-coordinate representation of the contour, a transformed (e.g., Fourier-transformed) representation can be used by a machine learning system implementing any known machine learning algorithm or method, including neural networks such as convolutional neural networks (CNNs), thereby resulting in an efficient estimation of the contour representation. By using a fixed-length, transformed representation that provides a compact representation of the contour, the complexity of the trained machine learning system can be reduced compared to the current state of the art for pixel-level instance descriptions, resulting in a faster processing speed and a smaller memory footprint. Also, it is advantageous that the contour can be easily reconstructed from the compact representation.

[0023] Another advantage is that, by reducing the amount of computation required, the method according to the present invention can reconstruct the contour of an object with higher accuracy compared to the prior art solutions using the same computing power.

[0024] The method according to the present invention can segment a number of objects in an image including occluded or partially hidden objects. An occluded or partially hidden object is, for example, an object that is not entirely visible in the image because at least a part of it is hidden behind another object, in which case the visible part of the object can be segmented and, depending on a particular embodiment of the method, the occluded part of the object can be ignored or assigned to the visible part of the same object.

[0025] The method according to the present invention isBy estimating the typical appearance (basic representation or reference contour) of the shape of an object and also by estimating geometric parameters by at least one of, or a combination of, geometric transformations such as scaling, rotation, mirroring, or translation of the object, where the geometric parameter(s) correspond to the size, position, and orientation of the object in the image, the contour of the object can be reconstructed. Separating the basic shape of the object and the above-described geometric transformations provides a representation of the object contour that can be estimated in a more efficient way, where the basic shape or reference contour is invariant to the above geometric transformations. Certain machine learning algorithms / methods, for example, convolutional neural networks, are invariant to translation and fit well with such a decomposed representation of the object contour. By applying this decomposed representation, the same reference contour can be used to estimate the same object located in different parts of the image regardless of their size, position, and orientation. Information regarding the exact size, position, and orientation can be encoded in a small number of geometric parameters. Further, in practical applications, the geometric transformation well approximates a rigid body transformation in 3D space, i.e., the movement of an object projected onto the image. Thus, when multiple images, for example, images of a camera stream, are processed in sequence, the consecutive images are similar to each other and the overall shape of the objects in the images is mostly the same, but the size, position, or orientation may vary slightly. By the approach of determining the shape and the corresponding geometric parameters, the computational requirements of the method are further reduced and fast segmentation of the objects in the image becomes possible. Such a representation is more easily learned by machine learning methods including, but not limited to, convolutional neural networks.

[0026] Therefore, the method according to the present invention can be used in any vision-based scene understanding system, including medical applications (medical image processing) or improving the vision of autonomous vehicles.

Brief Description of the Drawings

[0027] Preferred embodiments of the present invention are described below as examples with reference to the following drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0028] The present invention relates to a method for segmenting objects or object instances in an image, collectively referred to as object segmentation. Object instances are preferably restricted to a specific set of applications of interest, such as, for example, vehicles, pedestrians, etc. in the case of autonomous driving applications, or various organs in the case of medical applications. Throughout the description, the word "object" can refer to different object instances of the same category or objects of different categories. Further, the term "object segmentation" is used for the task of instance segmentation, i.e., labeling the pixels of an image with the identification tags of the corresponding object instances to which the pixels belong. In applications where only one object is present in the image, object segmentation simplifies to semantic segmentation, i.e., labeling each pixel with its category.

[0029] In the case of object segmentation, the normal task is to predict a label (identification tag, e.g., a number, code, or tag) for each pixel corresponding to a specific object in the image, resulting in an object mask at the pixel level. In the method according to the present invention, the objects to be segmented are represented by their contours (segmentation contours) in the image, based on which an object mask can be created, i.e., by including the pixels within the segmentation contour, with or without the segmentation contour itself.

[0030] According to the present invention, instead of directly determining the coordinates of the segmentation contour points in the real space, a representation, preferably a compact representation, is generated from the segmentation contour points. This representation of the segmentation contour (usually called a contour descriptor or descriptor) can be learned by a machine learning system. The machine learning system preferably implements any known machine learning algorithm or method. For example, the machine learning system includes a neural network, preferably a convolutional neural network. The trained machine learning system can preferably determine the descriptor by inverse transformation, and the segmentation contour can be reconstructed from the descriptor. The embodiment of the method according to the present invention shown in the figure is implemented by using a neural network as a machine learning algorithm due to its high efficiency in the segmentation task compared to other machine learning algorithms / methods known in the art. However, other machine learning algorithms / methods, such as filtering or feature extraction methods (e.g., Scale-Invariant Feature Transform (SIFT), Histogram of Oriented Gradients (HOG), Haar filter, or Gabor filter), regression methods (e.g., Support Vector Regression (SVR) or decision tree), ensemble methods (e.g., random forest, boosting), feature selection (e.g., Minimum Redundancy Maximum Relevance (MRMR)), dimensionality reduction (e.g., Principal Component Analysis (PCA)), or any suitable combination thereof can also be used. The machine learning algorithm / method must be trained so that the image matches the representation (descriptor) of the contour of the object from which the segmentation contour can be reconstructed.

[0031] The method according to the present invention for object segmentation in an image is inputting the image into a trained machine learning system, estimating, by the trained machine learning system, a representation of the segmentation contour of the object in the image, wherein the segmentation contour is a closed two-dimensional parametric curve, each point of the segmentation contour is defined by two coordinate components, and both coordinate components are parameterized, and A step of reconstructing an object's segmentation contour from an estimated representation of the segmentation contour is included.

[0032] According to the present invention, the segmentation contour of an object is a closed two-dimensional parametric curve, and its points (contour points) are defined by two coordinate components, both of which are parameterized. The use of a discrete number of contour points can limit the complexity of the method and reduce the required amount of calculation.

[0033] Preferably, the two coordinate components of the segmentation contour are parameterized independently, for example, by a time-like parameter, preferably by a single time-like parameter. The parameterized coordinate components in the 2D plane can be represented in any coordinate system and reference coordinate system using, for example, rectangular coordinates, polar coordinates, or complex (or any alternative) coordinate representations. The advantage of parameterizing the coordinate components of a two-dimensional curve together is that it can represent curves having any shape (including concave shapes). In a preferred embodiment of the method according to the present invention, the segmentation contour is represented by rectangular coordinates, and more preferably, the segmentation contour is represented by rectangular coordinates parameterized by a time-like parameter t that encodes the trajectory r of the curve, that is, r(t) = (x(t), y(t)), where x and y are functions that define the respective rectangular coordinates of the contour points of the segmentation contour. In another preferred embodiment, the parameterization of the segmentation contour is encoded via its tangent vector, i.e., the velocity along the trajectory, and the tangent vector can be extracted as the displacement vector of the contour points. In a further preferred embodiment, the segmentation contour is parameterized as a series of standardized line segments connecting the points of the segmentation contour.

[0034] Instead of directly estimating the contour points of the segmentation contour, the method according to the present invention estimates a representation, preferably a transformed and compact representation of the contour, by a trained machine learning system. The accuracy of the method, i.e., the approximation of the segmentation contour to the exact contour of the object, can be controlled by the dimension of the transformed representation, taking into account, for example, the available computing resources. Also, the transformed representation enables a decomposed representation of the segmentation contour including the general shape of the object (e.g., the reference contour) and the geometric transformations imposed on the shape. In a preferred embodiment of the present invention, the compact representation can be generated by a Fourier transform, more preferably by a discrete Fourier transform.

[0035] Thus, in a preferred embodiment of the present invention, the above displacement vector sequence is preferably transformed from the spatial domain to the frequency domain by a Fourier transform, more preferably by a discrete Fourier transform. As a result, the segmentation contour is represented by the amplitudes of the Fourier harmonics. In the literature (F.P. Kuhl and C.R. Giardina, “Elliptic Fourier features of a closed contour”, Computer Graphics and Image Processing, 1982), this representation is generally referred to as the elliptic Fourier descriptor (EFD) of the curve. The advantage of the discrete Fourier transform is that it can be performed for any two-component parameterization of the curve. To obtain a compact representation of the segmentation contour, the number of coefficients of the descriptor is limited to a fixed value. When estimating the representation (descriptor) of the segmentation contour, this value can be an input parameter for the machine learning algorithm and controls the accuracy of the reconstructed segmentation contour. By representing the segmentation contour of the object by a single vector of coefficients, a compact representation of fixed length is provided. The length of this vector is proportional to the number of harmonics used, e.g., in the case of a Fourier transform, the number of Fourier harmonics indicating the order of the transform. Hereinafter, this fixed-length vector is referred to as the Fourier descriptor.

[0036] For a single frequency, the two real-valued Fourier coefficients account for the amplitude and phase of a given harmonic, respectively. Generally, four real-valued coefficients are required to represent the single-frequency component of the trajectories of two components along a two-dimensional real-space contour. As a result, when the segmentation contour is represented by the Fourier descriptor of an ellipse, the length of the descriptor is 4×O, where O indicates the number of harmonics (also referred to as the degree in the literature) of the transform. Thus, the method according to the present invention simplifies the object segmentation task to the regression of a fixed-length vector containing the descriptor of the segmentation contour. This task can be learned from an existing set of training data that includes image and segmentation contour (or object mask) pairs, from which the above vector representation can be derived. The regression can be performed in any form including machine learning methods / algorithms, for example, by a convolutional neural network. The segmentation contour can be reconstructed from the descriptor by applying the inverse of the transform, i.e., in the case of the Fourier descriptor of an ellipse, the inverse discrete Fourier transform can be used.

[0037] It is emphasized that any suitable representation of the coefficients, such as Cartesian coordinates, polar coordinates, or complex vectors, is equivalent to the proposed method.

[0038] FIG. 1 and FIG. 2 illustrate a preferred embodiment of the method according to the present invention, and the trained machine learning system includes a neural network (20). The neural network (20) is trained to estimate the representation of the segmentation contour (40) of the object in the image (10) in step (S100) (FIG. 2), and the representation of the segmentation contour (40) is a Fourier descriptor (30), preferably a Fourier descriptor of an ellipse. In step (S110) (FIG. 2), the segmentation contour (40) can be reconstructed therefrom by inverse Fourier transform. An example of the Fourier descriptor (30) is shown in FIG. 5. In this embodiment, the neural network (20) directly determines the Fourier descriptor (30) and can directly reconstruct the segmentation contour (40) therefrom, that is, the reconstruction does not require a transformation of the Fourier descriptor (30). The deviation of the reconstructed segmentation contour (40) from the exact contour (boundary) of the object to be segmented depends on the number of Fourier coefficients used in the Fourier descriptor (30). By increasing the number of Fourier coefficients in the Fourier descriptor (30), the reconstructed segmentation contour (40) approximates the exact contour (boundary) of the object, but even with a limited number of Fourier coefficients, for example, 32 Fourier coefficients corresponding to a Fourier transform of degree 8, result in a reconstructed segmentation contour (40) that approximates the exact contour quite well (see FIG. 7 and its description).

[0039] Figures 3 and 4 illustrate further preferred embodiments of the method according to the present invention. Also, in this embodiment, the machine learning system comprises a neural network (20) trained to estimate a representation of a reference contour of an object in step (S100’) (Figure 4), the reference contour belonging to the typical appearance of the object. The neural network (20) is further trained to estimate at least one geometric parameter (34) of a geometric transformation in step (S120) (Figure 4). Thus, the estimated representation of the segmentation contour includes the representation of the reference contour belonging to the typical appearance of the object and at least one geometric parameter (34) of the geometric transformation. The neural network (20) is preferably a convolutional neural network, and the geometric transformation is preferably any kind of geometric transformation such as scaling, translation, rotation, mirroring, or any suitable combination thereof. The geometric parameter (34) may represent the actual size, position, and orientation of the object within the image (10). By exploiting these characteristics, a disentangled / decomposed representation can be created such that these geometric factors are separated from the shape descriptor (reference contour). By using this compact and disentangled representation, the regression problem becomes easier to learn by the machine learning system as the representations of the reference contour and the geometric transformation parameters are processed independently. This disentangled representation allows for the application of a less complex neural network (20), resulting in a faster inference time and a smaller memory footprint. Furthermore, the neural network (20) typically less overfits to the learning of simpler representations, thereby increasing the generalization characteristics of the learned model.

[0040] In the embodiments illustrated in FIGS. 3 and 4, the representation of the segmentation contour includes Fourier descriptors that are the Fourier transform of the reference contour. The output of the neural network (20) is at least one geometric parameter (34) of the Fourier descriptor (30') of the reference contour of the object to be segmented and a geometric transformation. In step (S130) (FIG. 4), the Fourier descriptor (30') of the reference contour and the geometric parameter (34) are integrally combined to form an adjusted descriptor (36), and the adjusted descriptor (36) is an estimated representation of the segmentation contour (40'). The segmentation contour (40') is reconstructed from the adjusted descriptor (36) by applying an inverse Fourier transform, preferably an inverse discrete Fourier transform (IDFT), in step (S110') (FIG. 4). An example of the steps of the above embodiment of the method can be seen in FIG. 6.

[0041] In a further preferred embodiment of the method according to the invention (not shown, the reference numerals refer to those in FIGS. 3 and 4), the estimated representation of the segmentation contour preferably includes a representation of the reference contour belonging to the typical appearance of the object and at least one geometric parameter (34) of a geometric transformation. The geometric transformation is preferably any kind of geometric transformation such as scaling, translation, rotation, mirroring, or any suitable combination thereof, and the geometric parameter (34) can represent the actual size, position, and orientation of the object. The representation of the segmentation contour is preferably a Fourier descriptor, preferably a Fourier descriptor of an ellipse, and preferably includes a Fourier descriptor that is the Fourier transform of the reference contour. For the reconstruction of the segmentation contour (40'), first, the reference contour is preferably reconstructed from the representation of the reference contour by applying an inverse Fourier transform, more preferably an inverse discrete Fourier transform to the Fourier descriptor of the reference contour. Then, in a second step, the reconstructed reference contour is transformed into the segmentation contour (40') by applying a geometric transformation to the reconstructed reference contour.

[0042] Figure 5 shows typical values of Fourier descriptors (30), in this case Fourier descriptors of an ellipse, estimated by a neural network (20) included in a machine learning system according to the methods of FIGS. 1 and 2. In the illustrated case, the Fourier transform up to the 8th order is used to represent the segmentation contour (40) of the object, and thus 8×4 Fourier coefficients were estimated by the neural network (20). By applying an inverse Fourier transform to these estimated coefficients that make up the Fourier descriptor (30), the segmentation contour (40) of the object can be reconstructed.

[0043] The implementation of the method according to FIGS. 3 and 4 is illustrated in FIG. 6. The input to the machine learning system comprising the neural network (20) is provided with the image (10) to be segmented, and the neural network (20) is preferably a convolutional neural network. The neural network (20) is trained to estimate a Fourier descriptor (30') corresponding to at least one geometric parameter (34) of the reference contour (shape) of the object and a geometric transformation, and the geometric parameter (34) corresponds to the size, position, and / or orientation of the object. Similar to FIG. 5, the Fourier descriptor (30') is exemplified by the estimated Fourier coefficients. In this case, the geometric parameters (34) include the horizontal and vertical displacements of the object in the image (10) represented by Δx and Δy, respectively, and a scale factor. The Fourier descriptor (30') and the geometric parameter (34) combine to form an adjusted descriptor (36), from which the segmentation contour (40') of the object can be reconstructed by an inverse Fourier transform.

[0044] Also, FIG. 6 includes a manually annotated contour, i.e., the ground truth contour (12) of the image (10). From a qualitative comparison between the ground truth contour (12) and the reconstructed segmentation contour (40'), it can be confirmed that the latter gives a good approximation of the accurate contour, i.e., the position, size, and overall shape of the object coincide with those of the ground truth contour (12).

[0045] A detailed comparison of the reconstructed segmentation contours determined by manual annotation, the method according to FIG. 2, and the method according to FIG. 4 is illustrated in FIG. 7. The first column of FIG. 7 consists of the images (10a), (10b), (10c) to be segmented. Since the images (10a), (10b), (10c) are grayscale or color images showing the same object (car) in different fields of view, the size and position of the object also differ. The second column of FIG. 7 shows the ground truth contours (12a), (12b), (12c) of the object determined by manual annotation.

[0046] The third column of FIG. 7 shows the reconstructed segmentation contours (40a), (40b), (40c) of the images (10a), (10b), (10c) respectively, according to a preferred embodiment of the method according to FIG. 2. The centroid of each reconstructed segmentation contour (40a), (40b), (40c) is represented by a cross symbol. The reconstructed segmentation contours (40a), (40b), (40c) match the objects seen in the images (10a), (10b), (10c) and the ground truth contours (12a), (12b), (12c). The reconstructed segmentation contours (40a), (40b), (40c) are reconstructed from the Fourier descriptors (30) determined by the trained machine learning system by the neural network (20) of the trained machine learning system according to FIGS. 1 and 2. The Fourier descriptor (30) in this particular example has 32 coefficients corresponding to a Fourier transform with 8 harmonics (the degree of the Fourier transform is 8).

[0047] The fourth column of FIG. 7 shows the segmented contours (40’a), (40’b), (40’c) reconstructed according to the preferred embodiment of the method according to FIG. 4 for each of the images (10a), (10b), (10c). The centroid of each reconstructed segmented contour (40a), (40b), (40c) is represented by a plus sign.

[0048] As can be seen in FIG. 7, different embodiments of the method according to the present invention, for example, the method according to FIG. 2 and the method according to FIG. 4, result in similar reconstructed segmented contours (40a), (40b), (40c) and reconstructed segmented contours (40’a), (40’b), (40’c). The reconstructed segmented contours (40a), (40b), (40c) and the reconstructed segmented contours (40’a), (40’b), (40’c) are all similar to their respective ground truth contours (12a), (12b), (12c).

[0049] FIG. 8 represents a comparison chart of the values of the Fourier descriptor coefficients (Fourier coefficients) according to FIG. 7. The Fourier coefficients are grouped according to the representation of the two coordinates of the segmented contour, i.e., the horizontal coordinate component and the vertical coordinate component of the segmented contour in the Cartesian basis. The chart in FIG. 8 compares the respective values of the Fourier coefficients. The white bars represent the values of the ground truth contours (12a), (12b), (12c) according to FIG. 7 (second column), the black bars represent the values of the Fourier coefficients according to the method of FIG. 2 (third column of FIG. 7), and the striped bars represent the values of the Fourier coefficients according to the method of FIG. 4 (fourth column of FIG. 7). As can be seen from the chart in FIG. 8, since the reconstructed segmented contours (40a), (40b), (40c), (40’a), (40’b), (40’c) provide a good approximation of the ground truth contours (12a), (12b), (12c), the embodiments of the method according to the present invention can be used for fast and reliable segmentation of objects in an image.

[0050] Figure 9 illustrates an example of the use of the method according to the present invention for reconstructing the segmentation contour of an object whose field of view in the image (10) is blocked / obstructed, for example, a partially hidden object. In this example, a part of the object in the image (10) is artificially covered, and in other cases, the object may be covered by a different object (the shielding object). In a particular application of the method according to the present invention, the shielded part of the object may be ignored, or in other applications, the shielded part will be assigned to the visible part of the same object.

[0051] In the case of shielding, it is desirable to represent parts of the same object with the same identification tag during segmentation. According to a preferred embodiment of the method according to the present invention, for example, an ordering parameter representing depth or layer can be determined for the shielding object. For example, based on the ordering parameters having the same or similar values, the segmented contours belonging to the same shielding object can be identified, and the same identification tag can be assigned to the segmentation contours belonging to the same object.

[0052] In a further preferred embodiment, in order to process the shielding, a visibility score value is preferably generated by a machine learning algorithm for the estimated representation of each segmentation contour. The visibility score value preferably indicates the visibility or non-visibility of each object part resulting from the division of the object into a plurality of parts by the shielding. Based on the visibility score value, the non-visible object parts can be ignored or omitted, for example, excluded from the segmented image, or the non-visible object parts can be assigned to the visible part of the same object, that is, made possible by assigning the same identification tag. The same identification tag is preferably assigned based on the ordering parameters as described above.

[0053] According to the embodiment shown in FIG. 9, the trained machine learning system comprises a neural network (20), and the neural network (20) is trained to detect a single object that constitutes a predetermined number of objects and / or a predetermined number of parts. In the example according to FIG. 9, the maximum number of parts constituting the object is 3, or 3 individual objects are segmented. The neural network (20) according to this embodiment of the method thus estimates three Fourier descriptors (30) (3 sets of Fourier coefficients), preferably Fourier descriptors of an ellipse, and the values of each Fourier descriptor (30) are shown in the graph as in FIG. 5. Further, the neural network (20) determines a visibility score value indicating the visibility of each object or object part. If the object or object part is not visible (obscured), its visibility score value is 0. In this example, only two visible objects (i.e., two parts of the same object) are in the image (10), and thus only these two visibility score values are non-zero.

[0054] The visibility score value of the visible object part in this example is 1, but other non-zero values can be used to indicate further parameters or features of the visible object or object part. In a particular embodiment of the method according to the present invention, the visibility score value can include a value corresponding to an ordering parameter value, for example, the distance from the camera taking the image (10). Based on the visibility score value and / or the ordering parameter, a relationship, preferably a spatial relationship of the segmentation contours, can be determined, and the segmentation contours belonging to the same object can be identified.

[0055] In the example according to FIG. 9, the visibility score value of a visible object or object part in the image (10) is 1, and the visibility score value of an invisible object or object part (hidden or shielded object or object part) in the image (10) is 0. According to FIG. 9, the reconstruction of the segmentation contour is performed only for visible objects or object parts, that is, only for objects / object parts having a visibility score value indicating visibility, in this case objects / object parts whose visibility score value is not 0, via the inverse discrete Fourier transform (IDFT). The reconstructed segmentation contour (40) of each object / object part is shown in the same reconstructed segmentation contour image.

[0056] The present invention further relates to a data processing system including means for performing the steps of the method according to the present invention. The data processing system is preferably implemented on one or more computers and is trained to provide object segmentation, for example, an estimation of the representation of the segmentation contour of an object. The input to the data processing system is an image to be segmented, and the image includes one or more objects or object parts. The segmentation contour of an object is represented as a closed two-dimensional parametric curve, each point is defined by two coordinate components, and both coordinate components are parameterized. The characteristics of the representation of the segmentation contour are described in more detail in relation to FIGS. 1 and 2. The data processing system preferably comprises a machine learning system trained by any training method known in the art, and the machine learning system is preferably trained on a segmented image having a manually annotated contour (ground truth contour) and a representation of the segmentation contour that is a closed two-dimensional parametric curve, each point being defined by two coordinate components, and both coordinate components being parameterized. Preferably, the representation of the segmentation contour is a Fourier descriptor, more preferably a Fourier descriptor of an ellipse.

[0057] Preferably, the machine learning system of the data processing system is further trained to provide an estimate of the parameters of at least one geometric transformation and / or the identification tag of each object, the geometric transformation including scaling, translation, rotation, and / or mirroring, and the identification tag preferably being a unique identifier of each object.

[0058] In a preferred embodiment, the same identification tag is assigned to the same object part. In a further preferred embodiment, the machine learning system of the data processing system is trained to segment a plurality of objects in the image and / or objects divided into parts by occlusion. A preferred data processing system comprises a machine learning system trained to determine a visibility score value for each object or object part related to the visibility of each object or object part. To process occlusion, the visibility score value may include a value of an ordering parameter representing the relative position of the occluding object, based on which the same identification tag can be assigned to object parts belonging to the same object.

[0059] The machine learning system of the data processing system preferably includes a neural network trained for object segmentation, more preferably a convolutional neural network.

[0060] Furthermore, the present invention relates to a computer program product including instructions that, when executed by a computer, cause the computer to execute an embodiment of the method according to the present invention.

[0061] The computer program product may be executable by one or more computers.

[0062] Also, the present invention relates to a computer-readable medium including instructions that, when executed by a computer, cause the computer to execute an embodiment of the method according to the present invention.

[0063] The computer-readable medium may be a single one or may include more separate parts.

[0064] Of course, the present invention is not limited to the preferred embodiments described in detail above, but further variations, modifications, and developments are possible within the scope of protection defined by the claims. Furthermore, all embodiments that can be defined by any combination of the dependent claims belong to the present invention.

[0065] List of reference signs 10 Image 10a, 10b, 10c Images 12 Ground truth contour 12a, 12b, 12c Ground truth contours 20 Neural network 30, 30’ Fourier descriptor 34 Geometric parameters 36 Adjusted descriptor 40’, 40 Segmentation contour 40a, 40b, 40c Segmentation contours 40’a, 40’b, 40’c Segmentation contours S100, S100’ (Fourier descriptor estimation) step S110, S110’ (Contour reconstruction) step S120 (Geometric parameter estimation) step S130 (Adjusted descriptor generation) step

Claims

1. A method for object segmentation in an image (10), the method comprising: inputting the image (10) into a trained machine learning system, and estimating, by the trained machine learning system, a representation of a segmentation contour (40, 40') of an object in the image (10), wherein the segmentation contour (40, 40') is a closed two-dimensional parametric curve, the estimating step including reconstructing the segmentation contour (40, 40') of the object having any contour including a concave-shaped contour from the estimated representation of the segmentation contour (40, 40'), each point of the segmentation contour (40, 40') being defined by two coordinate components, both coordinate components of the segmentation contour (40, 40') being parameterized independently, and the estimated representation including at least one parameter of a geometric transformation estimated by the trained machine learning system, and a representation of a reference contour belonging to a typical appearance of the object estimated by the trained machine learning system, the method being characterized in that.

2. The reconstruction of the segmentation contour (40, 40') is generating an adjusted representation by combining at least one parameter of the geometric transformation with the reference contour, and reconstructing the segmentation contour (40, 40') from the adjusted representation, or reconstructing the reference contour from the representation of the reference contour, and transforming the reconstructed reference contour into the segmentation contour (40, 40') by the geometric transformation The method according to claim 1, characterized in that it is performed by.

3. The method according to claim 1 or 2, characterized in that the geometric transformation includes scaling, translation, rotation, and / or mirroring.

4. The method according to any one of claims 1 to 3, characterized in that the representation of the segmentation contour (40, 40') is obtained by Fourier transform, the estimated representation includes Fourier descriptors estimated by the trained machine learning system, and the reconstruction of the segmentation contour (40, 40') includes applying an inverse Fourier transform to the Fourier descriptors.

5. The method according to claim 4, wherein the Fourier descriptor is an elliptical Fourier descriptor. **Claim 6** The method according to any one of claims 1 to 5, further comprising the step of generating an identification tag for each segmentation contour (40, 40') by the trained machine learning system. **Claim 7** To process occlusion, a visibility score value is generated by the trained machine learning system for the representation of each segmentation contour (40, 40'), and the visibility score value indicates whether an object or object part is visible, hidden, or occluded. The method according to claim 6, wherein the segmentation contour (40, 40') is reconstructed only for the representation having a visibility score value indicating the visibility of the object. **Claim 8** The method according to claim 7, wherein in the case of occlusion, the same identification tag is assigned to the segmentation contours (40, 40') belonging to the same object. **Claim 9** The method according to any one of claims 1 to 8, wherein the trained machine learning system includes a neural network. **Claim 10** The method according to claim 9, wherein the neural network is a convolutional neural network. **Claim 11** A computer system, into which a computer program including instructions for causing the computer to perform the method according to any one of claims 1 to 10 is loaded when the program is executed by the computer. **Claim 12** A non-transitory computer-readable medium having recorded thereon a program for causing a computer to perform the method according to any one of claims 1 to 10 when the program is executed by the computer.