Method, system and computer program for image segmentation

By using a controlled environment and neural networks to determine optimal conditions, the method addresses the challenges of image segmentation for objects that are out of focus, in shadow, or poorly lit, achieving precise object separation and masking.

WO2025153628A1PCT designated stage expired Publication Date: 2025-07-24PROFOTO BV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/051056
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-19
Filing Date
2025-01-16
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing image segmentation methods struggle with objects that are out of focus, blend with the background, are in shadow, or poorly lit, making it difficult to accurately separate them from the image.

Method used

A method and system that uses a controlled environment to capture images under different lighting and framing conditions, employing neural networks to determine optimal lighting configurations and camera settings for high-quality image segmentation, enabling accurate separation of objects from backgrounds.

Benefits of technology

Enables high-quality image segmentation by identifying optimal lighting and camera settings, allowing for precise object masking and extraction from challenging images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025051056_24072025_PF_FP_ABST
    Figure EP2025051056_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Method (200) for image segmentation for use in separation of an object in an image from a background. The object is arranged inside a volume configured to shut out ambient light. A first image is captured with a desired framing, lighting and camera settings. A second image is captured using the same framing but with a lighting setup for optimal image segmentation. The segmentation for the second image is then used to separate the object in the first image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD, SYSTEM AND COMPUTER PROGRAM FOR IMAGE SEGMENTATION

[0002] TECHNICAL FIELD

[0003] The present disclosure relates to methods, systems and computer programs for image segmentation.

[0004] BACKGROUND ART

[0005] Generating masks for objects in images relies on discerning which pixels belong to the object and which pixels do not, i.e. it involves image segmentation. Image segmentation can be difficult or impossible when features of an image are out of focus, blend in with the background, are in shadow and / or are poorly lit. There is a need in the art for methods that addresses these challenges.

[0006] SUMMARY OF THE DISCLOSURE

[0007] The present disclosure relates to a method for configuring a system for image segmentation. The system comprises a volume configured to prevent light from outside the volume to enter into the volume. The method comprises, for each object of a set of objects, arranging the object by itself inside the volume of the system. The method further comprises, for each object when arranged in the volume, determining a set of framing configurations, each framing configuration relating to a position and attitude with respect to the object and a current set of camera settings. The method also comprises, for each object when arranged in the volume and for each framing configuration and for each light configuration of a set of predetermined light configurations, capturing at least one image with the camera. The method additionally comprises, for each object when arranged in the volume and for each framing configuration, determining a ground truth image segmentation based on an image of the captured images. The method further comprises generating, using an automated process, an image segmentation for each of the captured images. The method also comprises, for each object when arranged in the volume and for each framing configuration, determining a score for each light configuration based on a comparison between the image segmentations generated by the automated process and the respective ground truth image, the score relating to a quality of the image segmentation generated by the automated process. The method additionally comprises configuring the system to automatically choose a light configuration for a framing configuration based on the determined scores for the light configurations.

[0008] The method thereby enables configuration of a system that is to be configured to capture images of an object that may be hard to cut out of an image, e.g. due to blurry edges and / or the edges being difficult or impossible to discriminate from the background due to the lighting of the object, and provide a mask for cutting out the object and / or an image with the object cut out.

[0009] According to some aspects, using an automated process comprises using a neural network. According to some aspects, the step of generating an image segmentation comprises generating a semantic image segmentation. According to some aspects, the step of generating an image segmentation comprises generating an instance segmentation.

[0010] Neural networks can capture complex relationships in images that enable segmentation, i.e. classification of individual pixels to belong to a set of predetermined categories. In the simplest case, the set of categories only comprises two categories; this can be used as to generate a mask that classifies each pixel in an image if it shows the object or not. The object arranged inside the volume may comprise a set of items to be photographed, e.g. a plurality of shirts or pants and a shirt. In such case, it may be desirable to also mask the individual items of the object. Semantic image segmentation provides a solution to this problem by enabling classification of each pixel to belong to a predetermined category, not just whether it shows part of the object or not. Instance segmentation further enables isolation of individual elements, such as providing respective masks for the individual shirts in the above example. Instance segmentation can be combined with semantic segmentation.

[0011] The present disclosure further relates to a method for image segmentation for use in separation of an object in an image from the background, said object being arranged inside a volume configured to prevent light from outside the volume to enter into the volume, wherein a camera is configured to capture an image of the object when the object is arranged within the volume. The camera is arranged at a first distance and first attitude with respect to the object, and with a first set of camera settings. The first distance, attitude and the first set of camera settings define a first framing of the object. The method comprises configuring a set of light devices, the set of light devices being arranged to provide light inside the volume, to provide light according to a first light configuration. The method further comprises capturing a first image of the object using the camera at said first distance and attitude, with said first set of camera settings, and using said first light configuration. The method also comprises configuring the set of light devices to provide light according to a second light configuration, wherein the second light configuration is configured to facilitate generating an image segmentation configured to separate the object from the background. The method additionally comprises configuring the camera settings to a second set of camera settings, wherein the second set of camera settings is configured to facilitate generating an image segmentation configured to separate the object from the background. The method also comprises capturing a second image of the object using the camera at said first distance and attitude, with the second set of camera settings, and using the second light configuration, wherein the first distance and attitude and second set of camera settings define a second framing of the object that is the same as the first framing of the object. The method further comprises generating an image segmentation configured to separate the object from the background based on the second image.

[0012] By generating a mask (the image segmentation) from an image that allows the silhouette of the object to be clearly identified, i.e. the second image, the method thereby enables generating a mask for an object in an image, i.e. the first image, that is difficult or impossible to accurately mask from that image alone. With the image segmentation from the second image, the corresponding part can be extracted from the first image, which has the specific lighting and camera parameter settings desired by the user.

[0013] According to some aspects, the method is performed using a system configured according to the method for configuring a system for image segmentation, as disclosed above and below.

[0014] According to some aspects, generating the image segmentation is performed using a neural network. According to some aspects, the neural network is trained based on training data generated when configuring a system for image segmentation according to the method for configuring a system for image segmentation, as disclosed above and below.

[0015] Neural networks enable learning relevant features of a wide variety of objects and thereby enables generating masks for a diverse set of objects as long as the lighting conditions and camera focus are set up to facilitate the image segmentation. Neural networks further enable continuous improvement of the method, specifically the image segmentation accuracy and ability to mask an increasingly diverse set of objects, based on data generated during usage of the method.

[0016] According to some aspects, generating the image segmentation comprises generating a semantic image segmentation. According to some aspects, the step of generating an image segmentation comprises generating an instance segmentation.

[0017] According to some aspects, the method further comprises generating a composite image based on the first image and the image segmentation, the composite image comprising a cutout of the object from the first image, the cutout being based on the image segmentation.

[0018] The present disclosure also relates to a computer program comprising computer program code which, when executed by the control circuitry of a system for image segmentation, as described above and below, causes the system to carry out the method for image segmentation, as described above and below. The computer program has all the technical effects and advantages of the disclosed method and system, as described above and below.

[0019] BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 illustrates aspects of the method for configuring a system for image segmentation;

[0021] Figure 2 illustrates aspects of the method for image segmentation for use in separation of an object in an image from a background; and

[0022] Figures 3a-d illustrate aspects of the system for image segmentation for use in separation of an object in an image from a background.

[0023] DETAILED DESCRIPTION

[0024] Figure 1 illustrates aspects of a method 100 for configuring a system for image segmentation. The system comprises a volume configured to prevent light from outside the volume to enter into the volume.

[0025] The basic idea to configure the system is to collect a set of images of objects under different lighting conditions, framings of the objects and using different camera settings, and then determine which light configurations provide the optimal conditions for masking an object given a framing and a set of camera settings. Camera, as used herein, i.e. throughout the present disclosure, comprises a camera body and a lens, e.g. a zoom lens or a prime lens. A zoom lens is a lens that is configured to enable the focal length be varied, while a prime lens has a fixed focal length. Camera settings, as used herein, comprise settings of both the camera body and of lenses used with the camera, such as aperture, focal length and focus.

[0026] Knowing the optimal conditions, the system can be configured to choose corresponding optimal conditions, such as light configurations and camera settings, in future use given a current framing and current set of camera settings. In particular, the system is configured to be able to automatically determine the optimal conditions, possibly using a neural network. Thus, according to some aspects, the method for configuring the system comprises collecting a set of images and corresponding parameters that can function as a training and validation set, e.g. for training a neural network.

[0027] Thus, the method 100 comprises, for each object of a set of objects, arranging S110 the object by itself inside the volume of the system.

[0028] By arranging the object inside a volume configured to prevent light from outside the volume to enter into the volume, it is ensured that a given lighting configuration will always have the same effect; there is no risk of ambient light or external light sources interfering with the image capture.

[0029] The method further comprises, for each object when arranged in the volume, determining S120 a set of framing configurations, each framing configuration relating to a position and attitude with respect to the object and a current set of camera settings.

[0030] The relative position and attitude between the camera, the camera comprising a camera body and a lens, and the object will determine where and how much of the object will be in frame when an image is captured. If other objects are present, they too may obscure at least part of the object. Framing, as used herein, refers to the set of pixels representing the object in an image captured at said relative specific position and attitude.

[0031] The method also comprises, for each object when arranged in the volume and for each framing configuration and for each light configuration of a set of predetermined light configurations, capturing S130 at least one image with the camera. In order to create a mask for the image, i.e. performing image segmentation wherein the object constitutes a segment, it is typically preferable if the object stands out with respect to its surroundings. By capturing images at a set of predetermined light configurations, features facilitating the image segmentation, such as the visible contour of the object, can be enhanced. Upon examining the effect of different light configurations, some light configurations will be better at facilitating image segmentation than others. By repeating the process for different objects at different framings, the effect of different light configurations for different objects and different framings on facilitating image segmentation can be deduced.

[0032] In order to evaluate how effective a particular light configuration is at facilitating image segmentation for an object at a particular framing, it is preferable to compare the resulting image segmentation with a ground truth image segmentation.

[0033] Thus, the method additionally comprises, for each object when arranged in the volume and for each framing configuration, determining S140 a ground truth image segmentation based on an image of the captured images. According to some aspects, the ground truth image segmentation is determined manually.

[0034] The method further comprises generating S150, using an automated process, an image segmentation for each of the captured images. According to some aspects, using an automated process comprises using a neural network.

[0035] In some scenarios it is desirable to insert multiple objects simultaneously in the volume and / or generating masks for components of an object in the volume. It might therefore be of interest to be able to separate different types of objects from each other and / or different instances of objects. Thus, according to some aspects, the step of generating S150 an image segmentation comprises generating S152 a semantic image segmentation. According to some aspects, the step of generating an image segmentation comprises generating an instance segmentation.

[0036] The method also comprises, for each object when arranged in the volume and for each framing configuration, determining S160 a score for each light configuration based on a comparison between the image segmentations generated by the automated process and the respective ground truth image, the score relating to a quality of the image segmentation generated by the automated process. By assigning a score, the quality of generated image segmentations can be compared and be used as a basis for choosing which light configuration to use given a specific set of camera settings and position and attitude relative to the object being photographed.

[0037] If a neural network is employed for generating S150 the image segmentation, the score can be used as a basis for updating the parameters of the neural network, e.g. via backpropagation.

[0038] The method additionally comprises configuring S170 the system to automatically choose a light configuration for a framing configuration based on the determined scores for the light configurations. The method can thereby choose the lighting configuration that generates the image segmentation of the highest quality.

[0039] Figure 2 illustrates aspects of the method 200 for image segmentation for use in separation of an object in an image from a background, said object being arranged inside a volume configured to prevent light from outside the volume to enter into the volume. A camera is configured to capture an image of the object when the object is arranged within the volume, the camera being arranged at a first distance and first attitude with respect to the object, and with a first set of camera settings, wherein the first distance, attitude and the first set of camera settings define a first framing of the object. According to some aspects, the method 200 is performed using a system for image segmentation as, described above and below. Thus, the method 200 may be used by a system configured according to the method 100 described in relation to Figure 1, above.

[0040] The method 200 comprises configuring S210 a set of light devices, the set of light devices being arranged to provide light inside the volume, to provide light according to a first light configuration. The method 200 further comprises configuring capturing S220 a first image of the object using the camera at said first distance and attitude, with said first set of camera settings, and using said first light configuration.

[0041] The first image will typically be related to a desired final end product, such as a product photograph. However, separating the object from the background may be challenging. For example, focus settings, shallow depth of field and / or a lighting setup that doesn't enable the contours of the object to be distinguished from the background may compromise the quality of an image segmentation or make it practically impossible. Thus, a second image is captured under different conditions that enable a high-quality image segmentation to be performed, but with the same framing as for the first image, so that there is a pixel-to-pixel correspondence between the first and second images; an image segmentation, i.e. a mask, that works well for the second image, i.e. is highly accurate, will apply to the first image as well, thereby eliminating the above mentioned potential problems.

[0042] Thus, the method 200 also comprises configuring S230 the set of light devices to provide light according to a second light configuration, wherein the second light configuration is configured to facilitate generating an image segmentation configured to separate the object from the background.

[0043] The method 200 additionally comprises configuring S240 the camera settings to a second set of camera settings, wherein the second set of camera settings is configured to facilitate generating an image segmentation configured to separate the object from the background.

[0044] The method 200 further comprises capturing S250 a second image of the object using the camera at said first distance and attitude, with the second set of camera settings, and using the second light configuration, wherein the first distance and attitude and second set of camera settings define a second framing of the object that is the same as the first framing of the object.

[0045] According to some aspects, the second image may be generated from a plurality of captured images. For instance, it may be desirable to perform so-called focus stacking, wherein images of the object are captured with a shallow depth of field, focusing on the object at different distances from the camera. The images are then blended, i.e. stacked, such that the portions in focus from each image are merged to a single second image with a wider depth of field that has the entire object in focus. In other words, the steps of configuring S240 the camera settings and capturing S250 the second image may be performed repeatedly. The second image that is subsequently used for the image segmentation may by generated from the set of second images, as in the illustrated example using focus stacking.

[0046] The method 200 also comprises generating S260 an image segmentation configured to separate the object from the background based on the second image. According to some aspects, generating S260 the image segmentation is performed using a neural network. According to some aspects, the neural network is trained based on training data generated when configuring a system for image segmentation according to the method described in relation to Figure 1, above. It may be desirable to able make the image segmentation based on more than just the presence of an object inside the volume, e.g. object category and / or having more than one object present inside the volume. Thus, according to some aspects, generating S260 the image segmentation comprises generating S262 a semantic image segmentation. According to some aspects, generating S260 the image segmentation comprises generating an instance image segmentation. According to some further aspects, generating S260 the image segmentation comprises simultaneously generating a semantic image segmentation and an instance segmentation.

[0047] According to some aspects, comprising generating S270 a composite image based on the first image and the image segmentation, the composite image comprising a cutout of the object from the first image, the cutout being based on the image segmentation. If there are multiple objects present, the method may provide image segmentations for each object based on instance and / or semantic segmentation.

[0048] The method may also be performed for video by repeatedly applying the disclosed method. For instance, the object may be placed on a rotating platform, which would mean that the camera will have different attitudes with respect to the object as the object is rotated. For every first image of the resulting first image sequence, the object is then rotated in the same manner, but under optimal lighting conditions that enable capturing corresponding second images from which image segmentations can be generated. The resulting set of image segmentations can then be used to generate a composite image sequence comprising a sequence of cutouts of the object from the first image sequence, thereby resulting in a video segmentation.

[0049] Figure 3a illustrates aspects of a system 300 for image segmentation for use in separation of an object 302 in an image from a background. The system comprises a volume 310 configured to prevent light from outside the volume to enter into the volume. The system further comprises a camera 320 configured to capture images of an object, when the object is arranged inside the volume. The system also comprises a set of light devices 330a-d, the set of light devices being arranged to provide light inside the volume. The system additionally comprises control circuitry 340 configured to adjust a set of camera settings of the camera, wherein the control system is configured to adjust a distance, d, and attitude, a, of the camera with respect to the object, wherein the distance, attitude and the set of camera settings define a framing of the object. The control system is configured to determine a light configuration of the set of light devices, and is further configured to cause the system to carry out the method for image segmentation as disclosed above and below. The system has all the technical effects and advantages of the disclosed method for image segmentation.

[0050] Figures 3b-3d illustrate an example of the system 300 generating an image segmentation for a shoe.

[0051] Figure 3b illustrates capturing a first image of the object and illustrates some potential challenges with generating an image segmentation of the object. The shoe 302 is arranged inside the volume 310. According to some aspects, the volume 310 can be accessed via a door. According to some aspects, the volume 310 is formed by connecting a set of pieces defining the volume. With the shoe 302 arranged inside the volume and all light outside the volume 310 being shut out, the user defines the settings that will result in the desired image. Thus, the camera 320 is arranged at a first distance and first attitude with respect to the object, and with a first set of camera settings, wherein the first distance, attitude and the first set of camera settings define a first framing of the object. A set of light devices 330e is then configured to provide light according to a first light configuration. With the lighting and camera set, a first image of the object is captured using the camera at said first distance and attitude, with said first set of camera settings, and using said first light configuration. In the illustrated example, the first camera settings might use a depth of field that keeps a first portion of the shoe 302a in focus and a second portion 302b out of focus, wherein the second portion 302b makes generating an accurate image segmentation. Likewise, the first light configuration might only light the first portion 302a of the shoe, with the second portion 302b being so dark that it blends in with the background, thereby making accurate image segmentation generation difficult or impossible.

[0052] Figure 3c illustrates how the system 300 overcomes the challenges of Figure 3b. After the first image has been captured, the system adjusts camera settings and the lighting configuration such that the shoe, in particular the silhouette of the shoe, is as clear as possible, which greatly facilitates image segmentation. Thus, the set of light devices is configured to provide light according to a second light configuration, wherein the second light configuration is configured to facilitate generating an image segmentation configured to separate the object from the background. In the illustrated example, the shoe is lit from different angles, e.g. backlit, in a way that reveals a portion 302c of the silhouette. Likewise, the camera settings are configured to a second set of camera settings, wherein the second set of camera settings is configured to facilitate generating an image segmentation configured to separate the object from the background. In the present example, the second set of camera settings comprises changing the depth of field to ensure that the entire shoe is in focus, will keeping the framing of the shoe unchanged with respect to the first image. Then a second image of the object is captured using the camera at said first distance and attitude, with the second set of camera settings, and using the second light configuration, wherein the first distance and attitude and second set of camera settings define a second framing of the object that is the same as the first framing of the object.

[0053] According to some aspects, a plurality of second images of the object is captured, each second image being captured at a respective focus point on the object. The second image used for image segmentation is generated based on the plurality of second images via focus stacking.

[0054] The settings in Figure 3c, resulting in the second image, enables a highly accurate image segmentation to be generated from the second image. The image segmentation can then optionally be used as a mask to extract the shoe from the first image, as illustrated in Figure 302d.

[0055] The present disclosure also relates to a computer program comprising computer program code which, when executed by the control circuitry of a system for image segmentation, as described in relation to Figure 3, causes the system to carry out the method for image segmentation as described in relation to Figure 2.

Claims

CLAIMS1. A method (100) for configuring a system for image segmentation, the system comprising a volume configured to prevent light from outside the volume to enter into the volume, the method (100) comprising for each object of a set of objects, arranging (S110) the object by itself inside the volume of the system, for each object when arranged in the volume, determining (S120) a set of framing configurations, each framing configuration relating to a position and attitude with respect to the object and a current set of camera settings, for each object when arranged in the volume and for each framing configuration and for each light configuration of a set of predetermined light configurations, capturing (S130) at least one image with the camera, for each object when arranged in the volume and for each framing configuration, determining (S140) a ground truth image segmentation based on an image of the captured images, generating (S150), using an automated process, an image segmentation for each of the captured images, for each object when arranged in the volume and for each framing configuration, determining (S160) a score for each light configuration based on a comparison between the image segmentations generated by the automated process and the respective ground truth image, the score relating to a quality of the image segmentation generated by the automated process, configuring (S170) the system to automatically choose a light configuration for a framing configuration based on the determined scores for the light configurations.

2. The method according to claim 1, wherein using an automated process comprises using a neural network.

3. The method according to claim 2 or 3, the step of generating (S150) an image segmentation comprises generating (S152) a semantic image segmentation.

4. Method (200) for image segmentation for use in separation of an object in an image from a background, said object being arranged inside a volume configured to prevent light from outside the volume to enter into the volume, wherein a camera is configured to capture an image of the object when the object is arranged within the volume, the camera being arranged at a first distance and first attitude with respect to the object, and with a first set of camera settings, wherein the first distance, attitude and the first set of camera settings define a first framing of the object, the method (200) comprising configuring (S210) a set of light devices, the set of light devices being arranged to provide light inside the volume, to provide light according to a first light configuration, capturing (S220) a first image of the object using the camera at said first distance and attitude, with said first set of camera settings, and using said first light configuration, configuring (S230) the set of light devices to provide light according to a second light configuration, wherein the second light configuration is configured to facilitate generating an image segmentation configured to separate the object from the background, configuring (S240) the camera settings to a second set of camera settings, wherein the second set of camera settings is configured to facilitate generating an image segmentation configured to separate the object from the background, capturing (S250) a second image of the object using the camera at said first distance and attitude, with the second set of camera settings, and using the second light configuration, wherein the first distance and attitude and second set of camera settings define a second framing of the object that is the same as the first framing of the object, generating (S260) an image segmentation configured to separate the object from the background based on the second image.

5. The method (200) according to claim 4, wherein the method (200) is performed using a system configured according to any of claims 1-3.

6. The method (200) according to claim 4 or 5, wherein generating (S260) the image segmentation is performed using a neural network.

7. The method (200) according to claim 6, wherein the neural network is trained based on training data generated when configuring a system for image segmentation according to the method according to claims 1-3.

8. The method (200) according to any of claims 4-7, wherein generating (260) the image segmentation comprises generating (S262) a semantic image segmentation.

9. The method (200) according to any of claims 4-8, further comprising generating (S270) a composite image based on the first image and the image segmentation, the composite image comprising a cutout of the object from the first image, the cutout being based on the image segmentation.

10. A system (300) for image segmentation for use in separation of an object (302) in an image from a background, the system comprising a volume (310) configured to prevent light from outside the volume to enter into the volume, a camera (320) configured to capture images of an object, when the object is arranged inside the volume, a set of light devices (330a-d), the set of light devices being arranged to provide light inside the volume, control circuitry (340) configured to adjust a set of camera settings of the camera, wherein the control system is configured to adjust a distance (d) and attitude (a) of the camera with respect to the object, wherein the distance, attitude and the set of camera settings define a framing of the object, wherein the control system is configured to determine a light configuration of the set of light devices, andwherein the control system is configured to cause the system to carry out the method for image segmentation according to any of claims 4-9.

11. A computer program comprising computer program code which, when executed by the control circuitry of a system for image segmentation according to claim 10, causes the system to carry out the method for image segmentation according to any of claims 4-9.

Citation Information

Patent Citations

  • Method, device, system and computer-program product for setting lighting condition and storage medium

    US11631230B2

  • Product onboarding machine

    US20200089997A1

  • Imaging apparatus for providing background separated images

    US7931380B2

  • Portable studio for item photography

    US9625794B2