System and method for three-dimensional scene perception
The integration of macropixel histograms and image segmentation with a neural network in time-of-flight systems enhances the precision of object boundary and size detection in 3D scene perception, addressing accuracy limitations in existing techniques.
Patent Information
- Application Number
- PCT/EP2025/072321
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-06
- Filing Date
- 2025-08-04
- Publication Date
- 2026-02-12
AI Technical Summary
Existing 3D scene perception techniques using time-of-flight systems face limitations in accuracy due to the distance between neighboring light spots and low light intensity of individual spots, which affect the precision of depth information at object edges and the detection of small objects.
A system and method that incorporates a time-of-flight device with macropixel histograms and image segmentation, using a neural network to match depth information with image segments, enhancing the determination of object boundaries and sizes by fusing image and depth data.
Improves the accuracy of boundary and size determination of objects in a scene, enabling more precise three-dimensional scene perception for applications like augmented reality and autonomous systems.
Smart Images

Figure EP2025072321_12022026_PF_FP_ABST
Abstract
Description
[0001] Sony Semiconductor Solutions Corporation et al.
[0002] SYSTEM AND METHOD FOR THREE-DIMENSIONAL SCENE
[0003] PERCEPTION
[0004] TECHNICAL FIELD
[0005] The present disclosure generally pertains to a system including a time-of-flight device, and a method for three-dimensional scene perception.
[0006] TECHNICAL BACKGROUND
[0007] In recent years, there has been growing interest in the field of three-dimensional (“3D”) scene perception, which may be applied across a broad range of fields, such as augmented or mixed or virtual reality, robotics, advanced driver assistance systems, or security and surveillance.
[0008] Some known approaches to 3D perception rely on stereo vision techniques, which typically involve using two or more cameras to capture multiple views of a scene and then triangulating the corresponding points to estimate depth. While effective, these methods can be computationally expensive as they may require significant processing power.
[0009] More recently, some known approaches leverage active sensing techniques, such as time-of- flight (“ToF”) systems, which typically use an illuminator and a ToF sensor (ToF receiver) to measure the ToF of light signals. Direct time-of-flight (“dToF”) systems typically emit a series of light pulses and measure the time it takes for the reflected light to return to the receiver. Indirect time-of-flight (“iToF”) systems typically measure the phase shift of the returning light waves relative to the emitted light. By analyzing ToF data, distances of objects in a scene to the ToF system may be estimated.
[0010] Moreover, spot ToF devices are known in which the illuminator emits a plurality of light spots to the scene, for example, a light pattern of separated high-intensity light areas and low-intensity light areas.
[0011] Typically, detailed depth information about objects in the scene may be obtained due to a large number of small light spots.
[0012] However, the distance between neighboring light spots and a low light intensity of a single spot may limit, in some cases, the accuracy of the depth information at edges of the objects.
[0013] Although there exist techniques for 3D scene perception, it is generally desirable to improve such techniques. Sony Semiconductor Solutions Corporation et al.
[0014] SUMMARY
[0015] According to a first aspect, the disclosure provides a system, comprising: a time-of-flight device configured to perform a time-of-flight measurement to obtain time-of-flight data representing distance information about a scene, wherein the time-of-flight data include one or more macropixel histograms; and circuitry configured to: segment image information received by the circuitry, the image information representing a captured image of the scene, to obtain image segments; and match the obtained distance information with the image segments.
[0016] According to a second aspect, the disclosure provides a method, comprising: segmenting received image information, the image information representing a captured image of a scene, to obtain image segments; obtaining time-of-flight data representing distance information about the scene, wherein the time-of-flight data include one or more macropixel histograms; and matching the obtained distance information with the image segments.
[0017] Further aspects are set forth in the dependent claims, the drawings and the following description.
[0018] BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Embodiments are explained by way of example with respect to the accompanying drawings, in which:
[0020] Fig. 1 schematically illustrates an embodiment and an operating principle of a time-of-flight device;
[0021] Fig. 2 schematically illustrates in a block diagram an embodiment of a system;
[0022] Fig. 3 A schematically illustrates in a block diagram an embodiment of a scene illuminated by a light spot, and, in a graph, a macropixel histogram representing the time-of-flight data associated with the light spot illuminating the scene;
[0023] Fig. 3B schematically illustrates in a block diagram an embodiment of a scene illuminated by four light spots, and, in a graph, four macropixel histograms representing the time-of-flight data associated with the light spots illuminating the scene;
[0024] Fig. 3C schematically illustrates in a block diagram an embodiment of a scene with full-field illumination, and, in a graph, an embodiment of a macropixel histogram representing the time- of-flight data associated with the full-field illumination of the scene; Sony Semiconductor Solutions Corporation et al.
[0025] Fig. 4A schematically illustrates in a block diagram an embodiment of a dynamic illumination in switching mode;
[0026] Fig. 4B schematically illustrates in a block diagram an embodiment of a dynamic illumination in adaptive optics mode;
[0027] Fig. 4C schematically illustrates in a block diagram an embodiment of a dynamic illumination by using an addressable illuminator;
[0028] Fig. 5 schematically illustrates in a block diagram an embodiment of training a neural network; and
[0029] Fig. 6 schematically illustrates in a flow diagram an embodiment of a method.
[0030] DETAILED DESCRIPTION OF EMBODIMENTS
[0031] Before a detailed description of some embodiments under reference of Fig. 2 is given, general explanations are made.
[0032] As mentioned in the outset, three-dimensional (“3D”) scene perception and depth estimation may be based on time-of-flight (“ToF”) techniques to detect objects in a scene, determine distances to objects in a scene, and track motions, rendering such techniques suitable for application in various fields, such as augmented or mixed or virtual reality, robotics and autonomous systems, medical imaging, etc.
[0033] For enhancing the general understanding of the present disclosure, an embodiment and an operating principle of a ToF device 1 is discussed in the following under reference of Fig. 1, which schematically illustrates the operating principle of the ToF device 1, and which may also apply to other embodiments of the present disclosure.
[0034] The ToF device 1 includes an illuminator 2 and a ToF camera module 3 which are controlled by a controller 4.
[0035] The illuminator 2 may include an addressable illuminator or emitter array, for example, an addressable VCSEL (“Vertical-Cavity Surface-Emitting Laser”) array. In other embodiments, the illuminator 2 may include at least two illuminators or emitters, which are individually addressable and placed next to each other. In some embodiments, the illuminator 2 may include an array of n x m illuminators where n and m are different or the same integer values.
[0036] The illuminator 2 illuminates a scene 9 with spotted light. The scene 9 includes a foreground object 10 which at least partially throws back the spotted light. Sony Semiconductor Solutions Corporation et al.
[0037] The spotted light has a spatial light pattern including high-intensity areas 11 and low-intensity areas 12 and, thus, a plurality of light spots corresponding to the high-intensity areas 11 is projected onto the scene 9. Generally, the spatial light pattern is not limited to a pattern of light spots, for example, the spatial light pattern may include or may correspond to a plurality of light patterns, wherein a light pattern may be a spot, a cross, a rectangle or may have any other geometrical shape.
[0038] The ToF camera module 3 includes, for example, a lens, an aperture and a ToF sensor (not shown) to detect light returning from the object 10 in the scene 9.
[0039] The controller 4 controls the overall operation of the ToF device 1, in particular, the ToF data acquisition by controlling that a ToF measurement is performed.
[0040] The ToF device 1 is based on the dToF technique and uses a Single-Photon Avalanche Diode (“SPAD”) array as the ToF sensor for its operation to acquire depth data, wherein the SPAD array includes a plurality of light detection pixels, wherein each light detection pixel includes one or more SPADs.
[0041] As depicted in Fig. 1, the dToF technique is based on a synchronized process of illuminating a scene 9 by the illuminator 2 and of acquiring the reflected illumination light returning from the scene 9 by the ToF camera module 3.
[0042] The process starts, for example, with the emission of short light pulses toward the scene 9. When these light pulses interact with objects in the scene 9, a portion of the photons is thrown back towards the ToF camera module 3. The reflected illumination light is detected by the SPADs that have the ability of creating an avalanche current for each received photon.
[0043] Then, the ToF sensor records the arrival time for each light detection event with respect to the time of emission of a light pulse and groups them in discrete time intervals (bins) to create a histogram, thereby generating ToF data (histogram data). Typically, the ToF sensor generates a histogram for each light detection pixel in which the light detection events generated by the respective light detection pixel are recorded.
[0044] In order to improve the signal-to-noise (“SNR”) ratio, this process may be repeated several times, and the final histogram may thus be the sum of the histograms for each emitted light pulse. Sony Semiconductor Solutions Corporation et al.
[0045] Afterwards, the ToF data (histogram data) is processed to detect peaks in the histogram indicating the arrival time of the reflected pulse and, thus, a distance to an object in the scene (depth data) is obtained.
[0046] As mentioned above, each light detection pixel in the ToF sensor may include one or more SPADs, and typically one histogram is built for each pixel (the light detection events of each SPAD of the light detection pixel may be summed) such that a 3D point cloud for the target field-of-view of the ToF device 1 may be obtained therefrom.
[0047] Returning to the general explanations, typically, as mentioned in the outset, detailed depth or distance information about objects in the scene may be obtained due to the large number of small light spots, in a non-limiting example, 576 spots.
[0048] However, as further mentioned in the outset, the distance between neighboring light spots and a low light intensity of a single spot may limit, in some cases, the accuracy of the depth information at edges of the objects. Furthermore, depth information of small objects may not be obtained.
[0049] This may limit the accuracy of the depth information at edges of the objects in some cases.
[0050] The histograms associated with edges of objects in the scene may have (at least) two detectable peaks, which may result from the object and a background (object). These peaks may have an intrinsic spatial ambiguity, since it may be difficult to determine which distance belongs to which object or background present in the scene.
[0051] Generally, it has been recognized that an accurate scale or size of the objects may be useful in various applications. It has further been recognized that in such cases a detailed knowledge of the depth profile of the objects may be less important compared to the boundaries and size of the objects.
[0052] For example, for augmented or mixed or virtual reality applications, the knowledge of accurate boundaries and sizes of the objects may improve the realism of constructed images or videos, which include real world objects and overlay ed or composited virtual objects. The virtual objects may be scaled and placed according to the detected environment and, thus, an accurate knowledge of scales and boundaries may improve the realism. Therefore, the virtual objects may be scaled and / or placed accurately, for instance, between two objects or on top of one object or the like. In some embodiments the virtual objects may be placed at a distance which is interpreted by the human eye as intermediate (for example equidistant or non-equi distant) between two objects. Sony Semiconductor Solutions Corporation et al.
[0053] Another example may be related to autonomous systems such as autonomous mobile platforms (e.g., robots or self-driving cars). The autonomous mobile platforms may navigate more accurately around objects if the boundaries and the size of the objects are known more precisely.
[0054] In view of the above-mentioned circumstances, it has been recognized that it may be beneficial to improve the determination of the boundaries and sizes of the objects present in a scene.
[0055] It has been recognized that a fusion of image information about a scene with depth or distance information about the scene may improve the determination of the boundaries and sizes of the objects in the scene.
[0056] It has been recognized that light detection pixels generated by neighboring light detection pixels may be recorded together in a single macropixel histogram in order to improve the SNR and / or capture light from a larger part of the scene potentially including object edges.
[0057] In this way, the spatial resolution is traded for time resolution, which may be beneficial, for instance, for detecting (at least) two peaks in the histograms resulting, for example, from light spots that impinge on edges of objects in the scene.
[0058] Using only depth data from histograms with (at least) two peaks is beneficial because boundaries and sizes of objects can be determined with a system with a small number of light spots or fullfield illumination, and, assuming on chip processing, less data needs to be transferred from the time-of-flight sensor.
[0059] Moreover, it has been recognized that the image information may provide image segments, since, for example, image segmentation based on edge detection is known.
[0060] It has been recognized, however, that these image segments lack depth information and that the fusion of the image information and the depth information may result in ambiguities at the edges of the objects in the scene.
[0061] It has thus been further recognized that the depth information from macropixel histograms may be matched with the image segments, in particular, by using an artificial neural network for improving the determination of the boundaries and sizes of the objects in the scene, since the correlations between the image segments of a scene, obtained based on the image information, and the depth information may be learned such that the ambiguities may be disentangled.
[0062] Hence, the determination of the boundaries and sizes of objects in a scene may be improved.
[0063] Hence, some embodiments pertain to a system, wherein the system includes: Sony Semiconductor Solutions Corporation et al. a time-of-flight device configured to perform a time-of-flight measurement to obtain time-of-flight data representing distance information about a scene, wherein the time-of-flight data include one or more macropixel histograms; and circuitry configured to: segment image information received by the circuitry, the image information representing a captured image of the scene, to obtain image segments; and match the obtained distance information with the image segments, in particular, wherein at least one micropixel histogram represents distance information for at least portions of two image segments and is matched to the image segments.
[0064] The system may be configured as or may be a mobile electronic device, such as a smartphone, a tablet, a laptop, a head mounted display etc. and may be used in applications in the fields of, for example, augmented or mixed or virtual reality, robotics, and advanced driver assistance systems, without limiting the disclosure in this regard.
[0065] The apparatus may be a robot, a vehicle, a medical device or the like.
[0066] The system may be, for example, configured as a measurement system suitable for integration in a robot, a vehicle, a medical device or the like.
[0067] The circuitry may include one or more processors. A processor may be or may include an application processor, a central processing unit (“CPU”), a graphical processing unit (“GPU”), a digital signal processor (“DSP”), a field-programmable gate array (“FPGA”), an application specific integrated circuit (“ASIC”) etc.
[0068] The circuitry may include one or more memory components, one or more input / output interfaces, one or more communication interfaces, one or more data bus interfaces and respective data busses to exchange data, one or more point-to-point connections to exchange data, etc.
[0069] The functionality of the circuitry may be implemented by typical electronic components configured to achieve the functionality as described herein. The functionality of the circuitry may be implemented in parts by typical electronic components and in parts by software configured to achieve the functionality as described herein. The functionality of the circuitry may be implemented by software configured to achieve the functionality as described herein.
[0070] In some embodiments, the system includes an image sensor configured to capture the image of the scene. Sony Semiconductor Solutions Corporation et al.
[0071] The image sensor may be or may include, for example, a charge-coupled device sensor (“CCD") or a complementary metal-oxide-semiconductor sensor (“CMOS”) or a SPAD array to generate image data. The CCD or CMOS sensor or the SPAD array may include a plurality of image pixels arranged in rows and columns, wherein each image pixel generates a pixel signal which is processed to obtain a pixel value.
[0072] The image data may be, for instance, RGB (“red-green-blue”) data or grey scale data representing an RGB image or grey scale image, respectively, and include the pixel values of the plurality of image pixels.
[0073] As mentioned above, the circuitry is configured to segment received image information, wherein the image information represent a captured image of the scene to obtain image segments.
[0074] Generally, image segmentation techniques are known, for example, the image segments may be obtained by edge detection and edge linking, thresholding methods, clustering methods or the like.
[0075] The image segments may represent groups of neighboring or adjacent pixels of the image such that each image segment may be characterized by a set of image pixel indexes identifying the image pixels of the image sensor associated with the group of neighboring pixels of the image.
[0076] As mentioned above, the system includes a time-of-flight (“ToF”) device which is configured to perform a ToF measurement to obtain ToF data representing distance information about the scene, wherein the ToF data include one or more macropixel histograms. The ToF device is a direct ToF device.
[0077] Distance information about the scene obtained by the ToF device may include the distances of one or more objects in the scene to the ToF device, wherein the object may represent anything included in the scene.
[0078] The ToF device includes an illuminator configured to emit one or more light patterns to the scene. The illuminator may emit one large light pattern which may also be referred to as flooded or full-field illumination. The light pattern may be a spot, a cross, a rectangle or may have any other geometrical shape.
[0079] The illuminator may include one or more (light) emitters, wherein each emitter may be a Light Emitting Diode (“LED”), an edge emitting laser, a Vertical-Cavity Surface-Emitting Laser (“VCSEL”) or the like. Each emitter may be individually controlled to emit a light pulse and / or groups of emitters can be individually controlled to emit a light pulse. The illuminator may Sony Semiconductor Solutions Corporation et al. include a plurality of drivers to drive each emitter or each group of emitters individually according to a respective control signal.
[0080] The illuminator may include optical parts such as lenses such as glass or plastic or liquid lenses or metalenses, mirrors, or optical filters. The illuminator may include mechanical parts to move the optical parts, for example, piezo actuators.
[0081] The ToF device includes a ToF sensor including a plurality of light detection pixels. Each light detection pixel may be configured to generate light detection events in response to incident light.
[0082] In some embodiments, each light detection pixel includes one or more single-photon avalanche diodes, wherein each single-photon avalanche diode is configured to perform photoelectric conversion on incident light to generate light detection events.
[0083] The ToF sensor is configured to measure the time-of-flight of each light detection event.
[0084] The ToF sensor is configured to count, for each light detection pixel, the light detection events generated by the respective light detection pixel.
[0085] The ToF sensor may be configured to record, for each light detection pixel, the time-of-flight of light detection events of the respective light detection pixel in a histogram.
[0086] Hence, the ToF sensor is configured to use one or more macropixels. A macropixel corresponds to a group of typically adjacent light detection pixels. The ToF sensor thus records, for each configured macropixel, the time-of-flight of light detection events of the respective group of light detection pixels in a histogram, which is referred to as macropixel histogram.
[0087] In other words, the ToF sensor is configured to record, for each macropixel, light detection events generated by one or more light detection pixels associated with the respective macropixel in a respective macropixel histogram.
[0088] This may increase the probability of detecting more than one peak for light spots that impinge, for example, on edges of objects. Thus, such regions of the scene may be detected based on the macropixel histograms that show more than one peak. These regions may then be matched with the image segments to disentangle the depth ambiguity and increase the spatial resolution at the edges of objects such that accurate boundaries and sizes of the objects are obtained.
[0089] Even though the spatial resolution of the ToF measurement is reduced due to the use of macropixel histograms or less light spots, the spatial resolution at the edges of objects may be increased by matching the depth information with the image segments. Sony Semiconductor Solutions Corporation et al.
[0090] It has further been recognized that light spots emitted to a scene may be small compared to the size of the objects in the scene, which may limit the probability that a light spot impinges on more than one object.
[0091] It has thus been recognized that increasing the size of light spots or light patterns illuminating a scene increases the probability of a light spot impinging on more than one object.
[0092] This may further increase the probability of detecting more than one peak in macropixel histograms associated with such regions of a scene.
[0093] Hence, in some embodiments, the ToF sensor is configured to use one or more macropixels which are configured according to the one or more light patterns.
[0094] In some embodiments, for example, the ToF sensor uses a number of light detection pixels associated with a macropixel according to a size of a light spot or light pattern emitted by the illuminator that is associated with the macropixel. For instance, a larger size of the light spot may result in a configuration with a larger number of light detection pixels associated with the macropixel and, thus, light detection events of more light detection pixels are recorded in the respective macropixel histogram.
[0095] In some embodiments, for example, the ToF sensor uses a number of macropixels according to a number of light spots emitted by the illuminator.
[0096] As mentioned above, the ToF device is configured to perform a ToF measurement. A ToF measurement may include at least a measurement time period during which a light pulse is emitted, and light detection events are recorded in macropixel histograms. A ToF measurement may include an output time period during which the macopixel histograms are output.
[0097] The circuitry may detect peaks in the macropixel histograms in some embodiments. Typically, in some embodiments, the time-of-flight device is configured to detect peaks in the micropixel histograms, for example, the time-of-flight sensor processes the micropixel histograms on-chip and is thus configured to detect peaks in the micropixel histograms.
[0098] The peaks in the macropixel histograms may be identified based on a predetermined threshold or signal-to-noise ratio or a shape, profile or the like, as it is generally known.
[0099] The circuitry may calculate, for each macropixel histogram (that has at least two peaks), the distances for the (at least two) peaks, in some embodiments, to obtain the distance information. Typically, in some embodiments, the time-of-flight device is configured to calculate, for each macropixel histogram, the distances for the peaks, for example, the time-of-flight sensor Sony Semiconductor Solutions Corporation et al. processes the micropixel histograms on-chip and is thus configured to calculate the distances for the peaks.
[0100] The distances for the macropixel histogram peaks may be inferred from the speed of light and the photon travel times.
[0101] As mentioned above, the circuitry matches the obtained distance information with the image segments.
[0102] The matching may be performed according to one or more predetermined rules relating to the properties of the micropixel histogram data and the image segments.
[0103] In some embodiments, a machine learning algorithm is used to match the distance information with the image segments.
[0104] The machine learning algorithm may be stored by the circuitry.
[0105] The machine learning algorithm may be or may include a support vector machine, a k-clustering algorithm, an artificial neural network or the like. The artificial neural network may include an input layer, one or more hidden layers and an output layer, as generally known.
[0106] The obtained distance information and the image segments are input to the machine learning algorithm which may then output the image segments, wherein to each image segment a distance is assigned. For example, the micropixel histograms may be used as input representing the distance information or the calculated distances may be used as input representing the distance information.
[0107] In some embodiments, the machine learning algorithm is or includes a trained neural network. The neural network may be trained with labelled data (e.g., segmented RGB images and macropixel histograms or the calculated distances) showing a large variety of scenes in which segments and their distances from the camera are included.
[0108] In some embodiments, the circuitry is further configured to output a segmented view of the scene including a size of each image segment and a distance for each image segment.
[0109] The circuitry may calculate (e.g., triangulate), based on the distances assigned to the image segments, a size of each image segment in one or more directions. For example, the circuitry may calculate at first a center point of the image segment and then calculates the size of the image segment in a horizontal and vertical direction.
[0110] In some embodiments, the circuitry is further configured to perform, based on the segmented view of the scene, three-dimensional scene perception. Sony Semiconductor Solutions Corporation et al.
[0111] The three-dimensional scene perception may include various techniques for processing the visual and depth or distance information. For example, the three-dimensional scene perception may include determining visual features and shapes of objects for obtaining semantic information about the scene. For example, the three-dimensional scene perception may include determining an arrangement of objects and a spatial relationship between objects for obtaining semantic information about the scene. Three-dimensional scene perception may include, for instance, determining a movement of objects between subsequent segmented views of the scene. Three- dimensional scene perception may include, for instance, identifying a correlation between the movement of two objects or determining a correlation between the two objects from image frame to image frame to determine a motion.
[0112] In some embodiments, the circuitry is further configured to control, based on a scene complexity, the number of light patterns emitted by the illuminator.
[0113] The scene complexity may correlate with the number of objects in the scene. The scene complexity may thus be indicated by the number of image segments. The scene complexity may be indicated by a number of peaks in one or more macropixel histograms.
[0114] A high scene complexity may indicate a larger number of light patterns to be emitted, while a low scene complexity may indicate a lower number of light patterns to be emitted.
[0115] In some embodiments, the circuitry is further configured to control, based on a scene complexity, a size of the light patterns emitted by the illuminator.
[0116] A high scene complexity may indicate use of a lower size of the light patterns, while a low scene complexity may indicate use of a larger size of the light patterns.
[0117] The illuminator may include adaptive optics for varying the size of the light spots.
[0118] In some embodiments, the circuitry is further configured to detect, based on a segmented view of the scene, regions-of-interest in the scene and to control the illuminator to emit light spots to the regions-of-interest. Regions-of-interest may include a particular set of one or more objects in the scene or predetermined visual features or the like.
[0119] The illuminator may be an addressable illuminator such that each emitter may be individually controlled by a respective control signal.
[0120] Some embodiments pertain to an illuminator, including: a plurality of electronically addressable emitters configured to generate light patterns varying in size; and Sony Semiconductor Solutions Corporation et al. circuitry configured to: receive a control signal representing data indicative of a complexity of a scene for adapting an illumination; and select, based on the control signal, emitters to provide light patterns sized to overlap edges of objects in the scene.
[0121] Some embodiments pertain to an illuminator, including: a plurality of electronically addressable emitters configured to generate light patterns; and circuitry configured to: receive a control signal representing data indicative of a complexity of a scene for adapting an illumination; and adapt, based on the control signal, a number of emitters to provide light patterns.
[0122] Some embodiments pertain to an illuminator, including: a plurality of emitters configured to generate light patterns; an optical element configured to image the light patterns onto a scene; and circuitry configured to: receive a control signal representing data indicative of a complexity of the scene for adapting an illumination; and adapt, based on the control signal, an imaging property of the optical element to focus or defocus the light patterns.
[0123] The optical element may be a lens or the like, for example, a glass or plastic lens or a liquid lens.
[0124] The imaging property may be adapted, for example, by adjusting a position of the optical element relative to the plurality of emitters or the imaging property may be adapted, for instance, by adjusting the focal length of the liquid lens by varying an electric signal applied to the liquid lens.
[0125] The illuminator may include a plurality of optical elements, for example, one for each emitter such that the imaging property may be individually adapted for the different emitters.
[0126] Some embodiments pertain to a system, including circuitry configured to: derive distance information about a scene from one or more received micropixel histograms; and match the distance information with received image segments of the scene.
[0127] The system may be the time-of-flight device as described herein or an information processing device or a combination thereof or any other system as described herein. Sony Semiconductor Solutions Corporation et al.
[0128] Some embodiments pertain to a method, including: deriving distance information about a scene from one or more received micropixel histograms; and match the distance information with received image segments of the scene.
[0129] Some embodiments pertain to a method, wherein the method includes: segmenting received image information, the image information representing a captured image of a scene, to obtain image segments; obtaining time-of-flight data representing distance information about the scene, wherein the time-of-flight data include one or more macropixel histograms; and match the obtained distance information with the image segments.
[0130] The methods may be performed by the circuitry described herein or the system as described herein or by an information processing device such as computer, a server, or the like.
[0131] The methods may be used in applications in the fields of, for example, augmented reality, robotics, and advanced driver assistance systems, without limiting the disclosure in this regard.
[0132] In some embodiments, the methods further includes outputting a segmented view of the scene including a size of each image segment and a distance for each image segment.
[0133] In some embodiments, the methods further includes performing, based on the segmented view of the scene, three-dimensional scene perception.
[0134] In some embodiments, as mentioned above, the matching may be performed using a machine learning algorithm which is or includes a trained neural network. The neural network may be trained with labelled data (segmented RGB images and macropixel histograms) showing a large variety of scenes in which segments and their distances from the camera may be annotated.
[0135] In some embodiments, the methods further include: obtaining the captured image from an image sensor of a system; and obtaining the time-of-flight data from a time-of-flight device of the system, wherein the time-of-flight device is configured to emit one or more light patterns to the scene and is configured to use one or more macropixels which are according to the one or more light patterns to obtain the one or more macropixel histograms.
[0136] In some embodiments, the methods further includes controlling, based on a scene complexity, the number of light patterns emitted by the time-of-flight device. The scene complexity may correlate with the number of objects in the scene. Sony Semiconductor Solutions Corporation et al.
[0137] In some embodiments, the methods further includes controlling, based on a scene complexity, a size of the light patterns emitted by the time-of-flight device.
[0138] In some embodiments, the methods further includes detecting, based on a segmented view of the scene, regions-of-interest in the scene and controlling the time-of-flight device to emit light patterns to the regions-of-interest. The segmented view may include image segments. Regions- of-interest may comprise a particular set of one or more objects in the scene, as mentioned above.
[0139] In some embodiments, the system is a mobile electronic device, such as a smartphone, a tablet, a laptop, a head mounted display etc.
[0140] Detailed descriptions of some embodiments are provided in Figs. 2 to 5.
[0141] Fig. 2 schematically illustrates in a block diagram an embodiment of a system 7, which is discussed in the following.
[0142] The system 7 is a mobile electronic device, for example a smartphone, and includes a ToF device 1, an image sensor 5, and circuitry 6.
[0143] As discussed under reference of Fig. 1 above, the ToF device 1 includes an illuminator 2 configured to emit light spots 11 to the scene 9 which includes an object 10, a ToF camera module 3 configured to obtain ToF data of the scene 9, and a controller 4 configured to control the ToF device 1 to perform a ToF measurement to obtain the ToF data.
[0144] The basic functionality of the ToF device 1 is discussed under reference of Fig. 1.
[0145] Here, however, the ToF device 1 is configured to use one or more macropixels which are configured according to the light spots 11.
[0146] The controller 4 provides the ToF data to the circuitry 6.
[0147] The image sensor 5 is configured to capture a color image of the scene 9 to obtain RGB or grey scale image data representing the image, wherein the image data is provided to the circuitry 6.
[0148] The position and orientation of the image sensor 5 and the ToF camera module 3 against each other is fixed and calibrated.
[0149] The circuitry 6 segments the image to obtain image segments.
[0150] Moreover, the ToF device 1 detects peaks in the macropixel histograms and calculates, for each macropixel histogram (that has at least two peaks), the distances for the (at least two) peaks. In other embodiments, the ToF device 1 outputs the ToF data including one or more macropixel Sony Semiconductor Solutions Corporation et al. histograms corresponding to the configured one or more macropixels to the circuitry 6 which then processes the one or more micropixel histograms.
[0151] Then, the circuitry 6 matches the calculated distances with the image segments, as discussed herein.
[0152] Each of Figs. 3A, 3B, and 3C schematically illustrate a scene 20 including an object 21 and a background wall 22, which are discussed in the following under reference of Fig. 2 and Figs. 3 A, 3B and 3C, respectively.
[0153] It is assumed that the system 7, i.e. the ToF device 1 and the image sensor 5, is oriented towards the scene 20. The object 21 is closer to the apparatus 8 than the background wall 22.
[0154] Figs. 3A to 3C differ in the number and / or size of the light spots illuminating the scene 20 and, thus, in the macropixel histograms obtained by a ToF measurement.
[0155] Fig. 3A schematically illustrates the scene 20 that is illuminated by a single light spot 23, and, in a graph, a macropixel histogram representing the ToF data associated with the light spot 23. Data in the histogram to the left indicate a distance closer to the system 7.
[0156] Since the portion of the light spot 23 impinging on the object 21 is smaller than the portion of the light spot 23 impinging on the background wall 22, less photons are returned from the object 21 and, thus, the histogram peak representing a shorter distance is smaller than the peak representing a longer distance.
[0157] Fig. 3B schematically illustrates the scene 20 illuminated by four light spots 24 to 27 and, in a graph, four macropixel histograms representing the time-of-flight data associated with the light spots 24 to 27 illuminating the scene 20.
[0158] The macropixel histogram representing the ToF data of the light spot 24 has a small peak, representing the shorter distance to the comer of object 21 illuminated by light spot 24, and a larger peak, representing the longer distance to the background wall 22.
[0159] The macropixel histogram representing the ToF data of the light spot 25 has a small peak which is somewhat larger than the small peak corresponding to light spot 24 and represents the shorter distance to the edge of object 21 illuminated by light spot 25. The larger peak represents the longer distance to the background wall 22.
[0160] The macropixel histogram representing the ToF data of the light spot 26 has a large peak. Since the light spot 26 only illuminates the object 21 and not any other objects, the macropixel histogram displays one peak. Sony Semiconductor Solutions Corporation et al.
[0161] The macropixel histogram representing the ToF data of the light spot 27 has two equally large peaks since a first portion of the light spot 27 impinges on the object 21 and a second portion impinges on the background wall 22, wherein the first and second portion of the light spot 27 may lead to the detection of approximately the same number of photons in the case of similar reflectivity of object 21 and wall 22.
[0162] Fig. 3C schematically illustrates in a block diagram an embodiment of a scene with full-field illumination 28 and, in a graph, an embodiment of a macropixel histogram representing the time- of-flight data associated with the full-field illumination of the scene 20.
[0163] The macropixel histogram includes two peaks representing the object 21 closer to the observer and the background wall 22 further away from the observer. Due to the smaller size of the object 21 compared to the background wall 22, less photons impinge on the object 21 and, thus, less photons are returned and detected. Therewith, the peak at the shorter distance, that is the peak corresponding to the object 21, is smaller than the peak at the longer distance corresponding to the background wall 22.
[0164] Hence, as illustrated by Figs. 3A to 3C, the utilization of a few larger light spots and a few corresponding larger macropixels may allow to detect regions of a scene that include edges of objects which are indicated by at least two peaks in the macropixel histograms.
[0165] However, a depth ambiguity remains for these regions, since the spatial resolution is decreased.
[0166] Thus, the image segments are matched with the calculated distances, for example, using a neural network in order to obtain a segmented view of the scene in which the boundaries and sizes of the objects are accurately determined.
[0167] A training of a corresponding neural network will be discussed under reference of Fig. 5 below.
[0168] In the following, some control procedures are discussed under reference of Fig. 4, which may be used to adapt the illumination according to the scene.
[0169] Fig. 4A schematically illustrates in a block diagram an embodiment of a dynamic illumination in switching mode.
[0170] Based on the complexity of the scene, the size of the light spot emitted by the illuminator may be varied as indicated by the arrow. Fig. 4A depicts a switching mode between spotted illumination (left) and full-field illumination (right).
[0171] Fig. 4B schematically illustrates in a block diagram an embodiment of a dynamic illumination in adaptive optics mode. Sony Semiconductor Solutions Corporation et al.
[0172] Depending on the complexity of the scene, the size of the light spots may be adapted with large spots (left) and small spots (right), for example, adaptive optics may include focusing or defocusing the light spots with a lens or controlling a liquid lens with an electric signal to change its focal length. Moreover, some (adjacent) emitters of an addressable emitter array may be switched on or off to control a size of the light spots, for example, when a plurality of emitters are used together to provide one light spot.
[0173] Fig. 4C schematically illustrates in a block diagram an embodiment of a dynamic illumination by an addressable illuminator which may emit light spots to regions-of-interest in the scene.
[0174] Fig. 5 schematically illustrates in a block diagram an embodiment of training a neural network 40. The neural network 40 is used to match the measured distances with the segments of the RGB images.
[0175] Training the neural network 40 includes updating weights (and biases) 42 of the neural network 40, wherein the weights (and biases) 42 are optimized such that the difference between the prediction 43 of the neural network 40 and the evaluation data 44 - that is the desired output - is minimized.
[0176] Data for training the neural network 40 includes segmented RGB images and macropixel histograms. The RGB images include a large variety of scenes, which may be a mix of real and synthetic data, totaling, for example, a magnitude of training data on the order of 100.000 frames.
[0177] The data is labelled such that each segment of an RGB image is assigned a peak of a macropixel histogram and thus a distance of the segment to the ToF sensor.
[0178] In addition to the segmented RGB images and the macropixel histograms, the number of photon counts in the histogram peak, the spatial overlap between segment and light spot, the peak height, and an estimate of the reflectivity of the material from RGB image may be used as input data 41 for the neural network 40 to match image segments and histogram peaks.
[0179] The data is split into a training dataset to train the neural network 40 and a testing dataset for an independent performance evaluation of the neural network 40. Both training and testing datasets include input data 41 that is provided to the neural network 40 making a prediction 43, and evaluation data 44 representing the desired output that the prediction 43 is evaluated against.
[0180] The neural network 40 includes a plurality of weights and biases 42 which are to be updated in the training process to improve the prediction. In the first training iteration, the weights and Sony Semiconductor Solutions Corporation et al. biases 42 of the neural network 40 are initialized randomly and the training data is fed into and propagated forward through the neural network 40, resulting in the prediction 43. The difference between the prediction 43 and the evaluation data 44 is quantified by a loss 45.
[0181] In a backward propagation step, the contributions of each weight and bias 42 to the loss 45 is determined. The weights and biases 42 are updated to minimize the loss 45. The step size of the updates 46 of the weights and biases 42 is determined by a learning rate.
[0182] In subsequent training iterations, the training process is repeated using updated weights and biases 42 until either a convergence criterium or a predetermined number of iterations is met.
[0183] Since the testing dataset is not used in the training procedure, the testing dataset serves to monitor the performance of the neural network 40 and to avoid overfitting of the weights and biases 42 to the training dataset.
[0184] Fig. 6 schematically illustrates in a flow diagram an embodiment of a method 50, which is discussed in the following.
[0185] The method 50 may be performed by the circuitry as described herein or the system as described herein or by an information processing device such as a computer, a server or the like.
[0186] At 51, a captured image of a scene is obtained, as discussed herein.
[0187] At 52, the captured image is segmented to obtain image segments, as discussed herein.
[0188] At 53, time-of-flight data representing distance information about the scene is obtained, wherein the time-of-flight data include one or more macropixel histograms, as discussed herein.
[0189] At 54, peaks in the macropixel histograms are detected, as discussed herein.
[0190] At 55, for each macropixel histogram that has at least two peaks, the distances for the at least two peaks are calculated, as discussed herein.
[0191] At 56, the calculated distances are matched with the image segments using a machine learning algorithm, as discussed herein.
[0192] Returning to the general explanations, summarizing some aspects of some embodiments:
[0193] A system for 3D perception is proposed which uses a few bigger light spots.
[0194] With this system the spatial resolution may be lower, however, this may be compensated for by looking at multiple peaks in the histograms because the multiple peaks may come from different objects in the scene. With the multiple peaks and one macropixel histogram, the distance to Sony Semiconductor Solutions Corporation et al. multiple objects in the scene may be measured that are additionally imaged using an image sensor.
[0195] The lower spatial resolution may be compensated by getting more information of distances of objects in the scene (from macro-pixel histograms).
[0196] Note that the present technology can also be configured as described below.
[0197] (1) An apparatus, including: a time-of-flight device configured to perform a time-of-flight measurement to obtain time-of-flight data representing distance information about a scene, wherein the time-of-flight data include one or more macropixel histograms; and circuitry configured to: segment image information received by the circuitry, the image information representing a captured image of the scene, to obtain image segments; and match the obtained distance information with the image segments.
[0198] (2) The apparatus according (1), wherein the circuitry is further configured to output a segmented view of the scene including a size of each image segment and a distance for each image segment.
[0199] (3). The apparatus according to (1) or (2), wherein the circuitry is further configured perform, based on the segmented view of the scene, three-dimensional scene perception.
[0200] (4) The apparatus of any one of (1) to (3), wherein a machine learning algorithm is used to match the distance information with the image segments, in particular, wherein the machine learning algorithm includes a trained neural network.
[0201] (5) The apparatus according to any one of (1) to (4), wherein the time-of-flight device includes: an illuminator configured to emit one or more light patterns to the scene; and a time-of-flight sensor including a plurality of light detection pixels, each light detection pixel being configured to generate light detection events in response to incident light, wherein the time-of-flight sensor is configured to use one or more macropixels which are configured according to the one or more light patterns, and wherein the time-of-flight sensor is further configured to record, for each macropixel, light detection events generated by one or more light detection pixels associated with the respective macropixel in a respective macropixel histogram. Sony Semiconductor Solutions Corporation et al.
[0202] (6) The apparatus according to any one of (1) to (5), wherein the circuitry is further configured to control, based on a scene complexity, the number of light patterns emitted by the illuminator.
[0203] (7) The apparatus according to any one of (1) to (6), wherein the circuitry is further configured to control, based on a scene complexity, a size of the light patterns emitted by the illuminator.
[0204] (8) The apparatus according to any one of (1) to (7), wherein the circuitry is further configured to detect, based on a segmented view of the scene, regions-of-interest in the scene and to control the illuminator to emit light patterns to the regions-of-interest.
[0205] (9) The apparatus according to any one of (1) to (8), wherein the apparatus further includes an image sensor configured to capture the image of the scene.
[0206] (10) The apparatus according to any one of (1) to (9), wherein the apparatus is a mobile electronic device, in particular, a smartphone.
[0207] (11) A method, including: segmenting received image information, the image information representing a captured image of a scene, to obtain image segments; obtaining time-of-flight data representing distance information about the scene, wherein the time-of-flight data include one or more macropixel histograms; and matching the obtained distance information with the image segments.
[0208] (12) The method according to (11), further including outputting a segmented view of the scene including a size of each image segment and a distance for each image segment.
[0209] (13) The method according to (11) or (12), further including performing, based on a segmented view of the scene, three-dimensional scene perception.
[0210] (14) The method according to any one of (11) to (13), wherein a machine learning algorithm is used to match the obtained distance information with the image segments, in particular, wherein the machine learning algorithm includes a trained neural network.
[0211] (15) The method according to any one of (11) to (14), further including: receiving the captured image from an image sensor of an apparatus; and receiving the time-of-flight data from a time-of-flight device of the apparatus, wherein the time-of-flight device is configured to emit one or more light patterns to the scene and is Sony Semiconductor Solutions Corporation et al. configured to use one or more macropixels which are configured according to the one or more light patterns to obtain the one or more macropixel histograms.
[0212] (16) The method according to any one of (11) to (15), further including controlling, based on a scene complexity, the number of light patterns emitted by the time-of-flight device.
[0213] (17) The method according to any one of (11) to (16), further including controlling, based on a scene complexity, a size of the light patterns emitted by the time-of-flight device.
[0214] (18) The method according to any one of (11) to (17), further including detecting, based on a segmented view of the scene, regions-of-interest in the scene and controlling the time-of-flight device to emit light patterns to the regions-of-interest.
[0215] (19) The method according to any one of (11) to (18), wherein the apparatus is a mobile electronic device.
[0216] (20) The method according to (19), wherein the mobile electronic device is a smartphone.
[0217] (21) A computer program comprising program code causing a computer to perform the method according to anyone of (11) to (20), when being carried out on a computer.
[0218] (22) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (11) to (20) to be performed.
[0219] (23) A system, including circuitry configured to: derive distance information about a scene from one or more received micropixel histograms; and match the distance information with received image segments of the scene.
[0220] (24) A method, including: deriving distance information about a scene from one or more received micropixel histograms; and match the distance information with received image segments of the scene.
[0221] (25) A computer program comprising program code causing a computer to perform the method according to (24), when being carried out on a computer.
[0222] (26) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to (24) to be performed. Sony Semiconductor Solutions Corporation et al.
[0223] (27) An illuminator, including: a plurality of electronically addressable emitters configured to generate light patterns varying in size; and circuitry configured to: receive a control signal representing data indicative of a complexity of a scene for adapting an illumination; and select, based on the control signal, emitters to provide light patterns sized to overlap edges of objects in the scene.
[0224] (28) An illuminator, including: a plurality of electronically addressable emitters configured to generate light patterns; and circuitry configured to: receive a control signal representing data indicative of a complexity of a scene for adapting an illumination; and adapt, based on the control signal, a number of emitters to provide light patterns.
[0225] (29) An illuminator, including: a plurality of emitters configured to generate light patterns; an optical element configured to image the light patterns onto a scene; and circuitry configured to: receive a control signal representing data indicative of a complexity of the scene for adapting an illumination; and adapt, based on the control signal, an imaging property of the optical element to focus or defocus the light patterns.
Claims
Sony Semiconductor Solutions Corporation et al.CLAIMS1. A system, comprising: a time-of-flight device configured to perform a time-of-flight measurement to obtain time-of-flight data representing distance information about a scene, wherein the time-of-flight data include one or more macropixel histograms; and circuitry configured to: segment image information received by the circuitry, the image information representing a captured image of the scene, to obtain image segments; and match, the obtained distance information with the image segments.
2. The system of claim 1, wherein the circuitry is further configured to output a segmented view of the scene including a size of each image segment and a distance for each image segment.
3. The system of claim 2, wherein the circuitry is further configured perform, based on the segmented view of the scene, three-dimensional scene perception.
4. The system of claim 1, wherein a machine learning algorithm is used to match the distance information with the image segments, in particular, wherein the machine learning algorithm includes a trained neural network.
5. The system of claim 1, wherein the time-of-flight device includes: an illuminator configured to emit one or more light patterns to the scene; and a time-of-flight sensor including a plurality of light detection pixels, each light detection pixel being configured to generate light detection events in response to incident light, wherein the time-of-flight sensor is configured to use one or more macropixels which are configured according to the one or more light patterns, and wherein the time-of-flight sensor is further configured to record, for each macropixel, light detection events generated by one or more light detection pixels associated with the respective macropixel in a respective macropixel histogram.
6. The system of claim 5, wherein the circuitry is further configured to control, based on a scene complexity, the number of light patterns emitted by the illuminator.
7. The system of claim 5, wherein the circuitry is further configured to control, based on a scene complexity, a size of the light patterns emitted by the illuminator.
8. The system of claim 5, wherein the circuitry is further configured to detect, based on a segmented view of the scene, regions-of-interest in the scene and to control the illuminator to emit light patterns to the regions-of-interest.Sony Semiconductor Solutions Corporation et al.
9. The system of claim 5, wherein the system further includes an image sensor configured to capture the image of the scene.
10. The system of claim 1, wherein the system is a mobile electronic device, in particular, a smartphone.
11. A method, comprising: segmenting received image information, the image information representing a captured image of a scene, to obtain image segments; obtaining time-of-flight data representing distance information about the scene, wherein the time-of-flight data include one or more macropixel histograms; and matching the obtained distance information with the image segments.
12. The method of claim 11, further comprising outputting a segmented view of the scene including a size of each image segment and a distance for each image segment.
13. The method of claim 11, further comprising performing, based on a segmented view of the scene, three-dimensional scene perception.
14. The method of claim 11, wherein a machine learning algorithm is used to match the obtained distance information with the image segments, in particular, wherein the machine learning algorithm includes a trained neural network.
15. The method of claim 11, further comprising: receiving the captured image from an image sensor of a system; and receiving the time-of-flight data from a time-of-flight device of the system, wherein the time-of-flight device is configured to emit one or more light patterns to the scene and is configured to use one or more macropixels which are configured according to the one or more light patterns to obtain the one or more macropixel histograms.
16. The method of claim 15, further comprising controlling, based on a scene complexity, the number of light patterns emitted by the time-of-flight device.
17. The method of claim 15, further comprising controlling, based on a scene complexity, a size of the light patterns emitted by the time-of-flight device.
18. The method of claim 15, further comprising detecting, based on a segmented view of the scene, regions-of-interest in the scene and controlling the time-of-flight device to emit light patterns to the regions-of-interest.
19. The method of claim 15, wherein the system is a mobile electronic device.Sony Semiconductor Solutions Corporation et al.
20. The method of claim 19, wherein the mobile electronic device is a smartphone.
Citation Information
Patent Citations
Dynamic structured light for depth sensing systems
US20190355138A1
Multimode detector for different time-of-flight based depth sensing modalities
US20220294998A1
Electronic device, method and computer program
US20230393278A1