Electronic device and method
Patent Information
- Application Number
- PCT/EP2026/058502
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026058502_01102026_PF_FP_ABST
Abstract
Description
[0001] Sony Semiconductor Solutions Corporation et al.
[0002] ELECTRONIC DEVICE AND METHOD
[0003] TECHNICAL FIELD
[0004] The present disclosure generally pertains to an electronic device and a method.
[0005] TECHNICAL BACKGROUND
[0006] Generally, it is known to generate a virtual representation of a real object in a virtual environment based on information obtained about the real object.
[0007] However, in some cases, the realistic or natural appearance of the virtual representation of the real object in the virtual environment may be improved.
[0008] Although there exist techniques for generating a virtual representation of a real object in a virtual environment, it is generally desirable to improve the existing techniques.
[0009] SUMMARY
[0010] According to a first aspect, the disclosure provides an electronic device, comprising circuitry configured to:
[0011] acquire depth information and color information of a real object to generate a colored virtual object in a virtual environment;
[0012] generate material information of the real object to identify texture information; and display the colored virtual object in the virtual environment in accordance with the texture information.
[0013] According to a second aspect, the disclosure provides a method, comprising:
[0014] acquiring depth information and color information of a real object to generate a colored virtual object in a virtual environment;
[0015] generating material information of the real object to identify texture information; and displaying the colored virtual object in the virtual environment in accordance with the texture information.
[0016] Further aspects are set forth in the dependent claims, the drawings and the following description.
[0017] BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Embodiments are explained by way of example with respect to the accompanying drawings, in which:
[0019] Fig. 1 schematically illustrates in a block diagram an embodiment of an electronic device;Sony Semiconductor Solutions Corporation et al.
[0020] Fig. 2 schematically illustrates in a flow diagram an embodiment of a method;
[0021] Fig. 3 A schematically illustrates a situation in which it may be difficult to separate a target peak from other peaks, wherein two target peaks overlap in a histogram;
[0022] Fig. 3B schematically illustrates a situation in which it may be difficult to separate a target peak from other peaks, wherein two separated peaks are present in a histogram;
[0023] Fig. 4 schematically illustrates in a flow diagram an embodiment of a method;
[0024] Fig. 5 schematically illustrates an embodiment of a peak contrast;
[0025] Fig. 6 schematically illustrates in a flow diagram an embodiment of a method;
[0026] Fig. 7 schematically illustrates in a block diagram an application example of the method of Fig. 4 or Fig. 6;
[0027] Fig. 8 schematically illustrates in a block diagram an application example of the method of Fig. 4 or Fig. 6;
[0028] Fig. 9 schematically illustrates in a flow diagram an embodiment of a method;
[0029] Fig. 10 schematically illustrates in a flow diagram an embodiment of a method; and
[0030] Fig. 11 schematically illustrates in a block diagram an embodiment of a multi-purpose computer.
[0031] DETAILED DESCRIPTION OF EMBODIMENTS
[0032] Before a detailed description of the embodiments under reference of Fig. 1 is given, general explanations are made.
[0033] As mentioned in the outset, generally, it is known to generate a virtual representation of a real object in a virtual environment based on information obtained about the real object.
[0034] However, in some cases, the realistic or natural appearance of the virtual representation of the real object in the virtual environment may be improved.
[0035] Some embodiments pertain to an electronic device, wherein the electronic device includes circuitry configured to:
[0036] receive depth information and color information of a real object to generate a colored virtual object in a virtual environment;
[0037] generate material information of the real object to identify texture information; and display the colored virtual object in the virtual environment in accordance with the texture information.Sony Semiconductor Solutions Corporation et al.
[0038] The electronic device may be an information processing apparatus such as a computer with an electronic visual display (e.g., a liquid crystal display (“LCD”), a light-emitting diode (“LED”) display, etc.).
[0039] The electronic device may be used in a system for virtual or digital content generation, for example, a system for virtual cinema productions.
[0040] Thus, some embodiments pertain to a system for virtual content generation and / or editing, wherein the system includes:
[0041] a color camera configured to capture a color image to acquire color information of a real object;
[0042] a time-of-flight sensor which includes a plurality of time-of-flight pixels, and wherein the time-of-flight sensor is configured to acquire histogram data with each of the plurality of time-of-flight pixels to acquire depth information of the real object;
[0043] an electronic device comprising circuitry configured to:
[0044] receive the depth information and the color information of the real object to generate a colored virtual object in a virtual environment;
[0045] generate material information of the real object to identify texture information; and
[0046] display the colored virtual object in the virtual environment in accordance with the texture information.
[0047] The color camera may be, for instance, RGB (“Red-Green-Blue”) camera or a Pan-Tilt-Zoom camera.
[0048] Some embodiments pertain to an electronic device, wherein the electronic device includes circuitry configured to:
[0049] acquire depth information and color information of a real object to generate a colored virtual object in a virtual environment;
[0050] generate material information of the real object to identify texture information; and display the colored virtual object in the virtual environment in accordance with the texture information.
[0051] The electronic device may be an information processing apparatus (e.g., a camera device, a computer, etc.).
[0052] In some embodiments, the electronic device is a mobile electronic device.Sony Semiconductor Solutions Corporation et al.
[0053] The mobile electronic device may be a smartphone, a tablet, a laptop, a virtual reality device or the like.
[0054] The circuitry includes typical electronic and / or optoelectronic and / or optical components configured to achieve the functions as described herein.
[0055] The circuitry may include a color camera and a time-of-flight sensor. The color camera is configured to capture a color image to acquire the color information. The time-of-flight sensor includes a plurality of time-of-flight pixels, and sensor is configured to acquire histogram data with each of the plurality of time-of-flight pixels to acquire the depth information.
[0056] The circuitry may include one or more further cameras and / or one or more further time-of-flight sensors. The circuitry may include one or more processors, storage, a network interface, an input / output interface, one or more data buses, etc. The circuitry may include an electronic visual display (e.g., a liquid crystal display (“LCD”), a light-emitting diode (“LED”) display, etc.). The electronic visual display may be configured as a touch display (touchscreen) such that a user can interact with displayed content via, for example, a finger or an electronic pen. The circuitry may include input devices such as a keyboard, a computer mouse, a joystick, a microphone, etc. The circuitry may include input devices such as a speaker and the above-mentioned electronic visual display. The electronic visual display may be configured as a stereoscopic display to convey depth to a user or viewer.
[0057] As mentioned above, the circuitry is configured to generate material information of the real object to identify texture information.
[0058] The material information may be generated based on the captured color image of the real object, for example, the captured color image may be processed to localize and classify the real object represented in the captured color image. This may be done by using, for instance, a two-dimensional CNN (“Convolutional Neural Network”), which includes an input layer, one or more two-dimensional convolutional layers, each with one or more filters, one or more fully connected layers and an output layer to process the output of the last two-dimensional convolutional layer to predict a bounding box for the real object represented in the captured color image and a classification of the real object represented in the captured color image.
[0059] However, in such embodiments, the material information is the same for the whole real object and, thus, also the texture information is the same for the whole real object.
[0060] It has been recognized that another approach may be used to generate material information for different spatial parts of the real object depending on the spatial resolution of the time-of-flightSony Semiconductor Solutions Corporation et al.
[0061] sensor, since it has been recognized that the histogram data acquired by each time-of-flight pixel should be processed separately with a deep neural network.
[0062] In some embodiments, the circuitry is configured to input the histogram data of a time-of-flight pixel into a neural network and the neural network is configured to predict a peak position and a peak contrast for each of a preset number of peaks.
[0063] The neural network includes, as the histogram data is one-dimensional data, an input layer, one or more one-dimensional convolutional layers, each with one or more one-dimensional filters, one or more fully connected layers and an output layer to process the output of the last convolutional layer to predict the peak position and the peak contrast for each of the preset number of peaks.
[0064] Thus, in some embodiments, the neural network includes a one-dimensional convolutional layer which uses one or more one-dimensional filters.
[0065] As mentioned above, the histogram data acquired by each time-of-flight pixel should be processed separately or independently by the neural network.
[0066] In some embodiments, the neural network is configured to process the histogram data acquired with each time-of-flight pixel sequentially. For example, first histogram data acquired with a first time-of-flight pixel is input to the neural network and processed by the neural network, then second histogram data acquired with a second time-of-flight pixel is input to the neural network and processed by the neural network, and so on until all datasets are processed.
[0067] In some embodiments, the neural network is configured to process a group or all of the histogram data acquired with each time-of-flight pixel in parallel as a batch, however, the neural network is configured to process the group or all of the histogram data independently. Thus, the neural network includes a plurality of separate or independent processing paths, each processing path including an input layer to accept the respective one-dimensional histogram data, one or more one-dimensional convolutional layers, each with one or more one-dimensional filters, one or more fully connected layers and an output layer to process the output of the last convolutional layer to predict the peak position and the peak contrast for each of the preset number of peaks. In some embodiments, the neural network is configured to classify each peak for generating the material information.
[0068] In some embodiments, the material information is a peak class indicating one of not a peak, cover glass and a material of a target.Sony Semiconductor Solutions Corporation et al.
[0069] In some embodiments, the circuitry is configured to select the peaks with a peak class corresponding to the material of the target.
[0070] In some embodiments, the circuitry is configured to determine, based on the peak position, a depth for each selected peak with a signal-to-noise ratio above a preset threshold, wherein the signal-to-noise ratio is based on the peak contrast, and assign the material of the target to the depth.
[0071] As mentioned above, the circuitry is configured to generate a colored virtual object in a virtual environment and to display the colored virtual object in the virtual environment in accordance with the texture information.
[0072] The circuitry generates, based on the depth information, a virtual representation of the real object in the virtual environment. Generally, depth information represents a point cloud and the circuitry converts point cloud data to a three-dimensional virtual representation of the real object in the virtual environment. The color information is then associated with the three-dimensional virtual representation of the real object to form the colored virtual object.
[0073] The virtual environment may also be referred to as a virtual scene or a virtual setting or a virtual world or a virtual space or the like.
[0074] The circuitry may run or execute a (software) application configured for user-assisted virtual or digital content generation and / or editing. The application may provide a user interface via the circuitry to receive a user input such that a user can interact with the application to generate or edit virtual or digital content, for example, a graphical user interface (“GUI”) may be provided to receive a user input, for instance, via a touchscreen.
[0075] The user can input commands via the user interface to define or set up a virtual environment in which the colored virtual object is displayed. The virtual environment is a three-dimensional virtual space in which positions of virtual light sources and virtual objects are defined by three-dimensional coordinates (e.g., Cartesian coordinates with respect to a set virtual origin). The virtual environment may include one or more virtual light sources. The virtual environment may further include one or more (further) virtual objects.
[0076] As mentioned above, the circuitry is configured to display the colored virtual object in the virtual environment in accordance with the texture information.
[0077] Typically, the colored virtual object is displayed on an electronic visual display of the circuitry which is two-dimensional display such that the three-dimensional coordinates of the virtualSony Semiconductor Solutions Corporation et al.
[0078] environment are projected on the two-dimensional plane to display the colored virtual object in the virtual environment in accordance with the texture information. However, the electronic visual display may also provide a three-dimensional impression, for example, the electronic visual display is configured as a stereoscopic display to display the colored virtual object in the virtual environment in accordance with the texture information such that an impression of depth is conveyed to the user.
[0079] The texture information represents a surface structure and organization and, as such, influences the visual appearance of an object depending on the illumination conditions. The texture information may thus include, for example, a bidirectional reflectance distribution function (“BRDF”) to account for the material indicated in or by the material information. Such functions may be predefined and stored by the circuitry. The texture information may include a description of a typical surface profile of the material indicated by the material information.
[0080] The circuitry calculates the (spatial) virtual light distribution in accordance with the one or more virtual light sources and the colored virtual object (and possibly further virtual objects) in the virtual environment. In particular, the circuitry calculates the spatial virtual light distribution on the colored virtual object to calculate a virtual light scattering in accordance with the texture information, for example, in accordance with the BRDF of the material indicated in or by the material information. Thus, the colored virtual object has an improved natural and realistic appearance in the virtual environment.
[0081] The circuitry may further mix or overlay the displayed colored virtual object in the virtual environment with a two-dimensional image or two-dimensional video. In this way an image or video is mixed with computer graphics, and the two-dimensional image or the two-dimensional video may represent a background image or video, respectively. The image or video itself may be represented by means of computer graphics. The video may correspond to screen content. As mentioned above, the circuitry and the application may provide a user interface to allow the user to edit the virtual content in the virtual environment. The user can input commands via the user interface to interact with the virtual environment, in particular, with the virtual light sources and virtual objects present in the virtual environment. For instance, the received user input may indicate to move or rotate the displayed colored virtual object (or another displayed virtual object or virtual light source). The received user input may indicate a sequence of movements or rotations to generate a video.Sony Semiconductor Solutions Corporation et al.
[0082] A change of the position and / or orientation of the colored virtual object in the virtual environment may lead to a different visual appearance, since the colored virtual object is displayed in accordance with the texture information.
[0083] In some embodiments, the circuitry is configured to receive a user input indicating a change of a position and / or orientation of the colored virtual object in the virtual environment and to display the colored virtual object in the virtual environment in accordance with the change in position and / or orientation of the colored virtual object in the virtual environment and the texture information.
[0084] In some embodiments, the circuitry is configured to receive a user input indicating a change of color information for at least a part of the colored virtual object and to display the colored virtual object in the virtual environment in accordance with the change of color information.
[0085] In some embodiments, the circuitry is configured to receive a user input indicating a change of material information for at least a part of the colored virtual object and to display the colored virtual object in the virtual environment in accordance with the change of material information. Some embodiments pertain to a method, wherein the method includes:
[0086] receiving or acquiring depth information and color information of a real object to generate a colored virtual object in a virtual environment;
[0087] generating material information of the real object to identify texture information; and displaying the colored virtual object in the virtual environment in accordance with the texture information.
[0088] The method may be performed by the electronic device as described herein, in particular, the method may be performed by a mobile electronic device.
[0089] Some embodiments pertain to an electronic device, including circuitry configured to:
[0090] receive or acquire depth information and color information of a real object; generate a colored virtual object in a virtual environment by estimating a material by neural network circuitry operating on the depth information of the real object to identify texture information; and
[0091] assign texture information to the colored virtual object in the virtual environment.
[0092] In some embodiments, the circuitry is configured to process the colored virtual object to change color properties based on the texture information.
[0093] Some embodiments pertain to a method, including:Sony Semiconductor Solutions Corporation et al.
[0094] receiving or acquiring depth information and color information of a real object; generating a colored virtual object in a virtual environment by estimating a material by neural network circuitry operating on the depth information of the real object to identify texture information; and
[0095] assigning texture information to the colored virtual object in the virtual environment. The method may be performed by the electronic device as described herein, in particular, the method may be performed by a mobile electronic device.
[0096] Some embodiments pertain to an electronic device, wherein the mobile electronic device includes circuitry configured to:
[0097] receive or acquire histogram data representing depth information of a scene;
[0098] input the histogram data into a neural network, wherein the neural network is configured to predict a peak position and a peak contrast for each of a preset number of peaks; and determine a depth for each peak with a signal-to-noise ratio above a preset threshold, wherein the signal-to-noise ratio is based on the peak contrast.
[0099] Some embodiments pertain to a method, wherein the method includes:
[0100] receiving or acquiring histogram data representing depth information of a scene; inputting the histogram data into a neural network, wherein the neural network is configured to predict a peak position and a peak contrast for each of a preset number of peaks; and
[0101] determining a depth for each peak with a signal-to-noise ratio above a preset threshold, wherein the signal-to-noise ratio is based on the peak contrast.
[0102] The method may be performed by the electronic device as described herein, in particular, the method may be performed by a mobile electronic device.
[0103] The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer-readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.
[0104] Returning to Fig. 1, there is schematically illustrated in a block diagram an embodiment of an electronic device 50, which is discussed in the following.Sony Semiconductor Solutions Corporation et al.
[0105] The electronic device 50 is a mobile electronic device 50 which includes a ToF sensor 51, an RGB (“Red-Green-Blue”) camera 52, a processor 53, a storage 54 and an electronic visual display 55.
[0106] The mobile electronic device 50 is, for example, a smartphone in this embodiment and the electronic visual display 55 is a touchscreen.
[0107] The mobile electronic device 50 performs the method 70 of Fig. 2, which schematically illustrates in a flow diagram the method 70.
[0108] At 71, the mobile electronic device 50 acquires depth information of a real object 61 to generate a colored virtual object 62 in a virtual environment 63.
[0109] The stripes of the real object 61 in Fig. 1 represent color information and / or material information of the real object 61 for the sake of illustration only.
[0110] The mobile electronic device 50 uses the ToF sensor 50 to acquire the depth information.
[0111] The ToF sensor 51 has a plurality of ToF pixels, such as SPAD (“Single-Photon Avalanche Diode”) pixels, to acquire histogram data with each ToF pixel by performing a ToF measurement. The ToF sensor further has one or more light sources, e.g. an edge emitting laser or a Vertical-Cavity Surface-Emitting Laser (“VCSEL”), wherein each is configured to emit a train of temporally separated light pulses to a different spatial part of the scene, for example, to a different spatial part of the real object 61.
[0112] The ToF measurement includes emitting one or more spatially separated trains of temporally separated light pulses to a scene, wherein each light pulse emission triggers a light detection period with the SPAD pixels. Thus, with the emission of a light pulse, the light detection period with the SPAD pixels is started and synchronized, for example, the SPAD pixels may be activated in response to the trigger. The light detection of each SAPD pixel is performed for a preset time interval (the light detection period) which is divided into a plurality of shorter time intervals and each shorter time interval is associated with a histogram bin. The light detection events generated by the respective SPAD pixel during each of the shorter time intervals are counted, thereby histogram data is acquired with each SPAD pixel representing a histogram of a plurality of bins (i.e. time axis) and associated counts of light detection events (i.e. signal intensity axis). In other words, one histogram is generated for each SPAD pixel, based on the generated light detection events of the respective SPAD pixel.Sony Semiconductor Solutions Corporation et al.
[0113] A peak in the histogram may indicate presence of an object at a certain depth indicated by the peak position on the time axis due to the relationship of depth and signal runtime of the emitted light pulse. The ToF measurement may be performed for a number of light pulses of the train to increase the SNR of the histogram. After the ToF measurement, the histogram data acquired with each SPAD pixel is processed by a logic circuit (e.g., processor 53). The logic circuit may be part of the ToF sensor or may be separated from the ToF sensor such that the histogram data is read out from the ToF sensor before being processed.
[0114] The histogram data acquired with each ToF pixel represents depth information and may then be processed using a neural network (circuitry) to generate a depth value for each part of the real object 61 that is in the field-of-view of the ToF sensor 51.
[0115] Each spatial part of the real object 61 is associated with a different portion of the depth information corresponding to the spatial resolution of the ToF sensor 51. In other words, each histogram, in which the real object 61 is represented as a peak, represents depth information of a different spatial part of the real object 61.
[0116] The processor 53 processes the histograms to generate the depth values of the real object 61. The processing using the neural network will be discussed in more detail under references of Fig. 4, 5 and 6. As a part of the processing, the processor 53 only selects a peak in each histogram to determine the depth, i.e. generate a depth value, that is classified as belonging to the real object 61, for example, the peak class may indicate one of: not a peak, cover glass and a material of the real object 61. Then, only peaks indicating materials of the real object 61 are selected. The granularity of the classification of the material(s) of the real object 61 may depend on the application such that it may be chosen coarser or finer. A rough granularity may include a classification according to larger material classes such as metal, stone or mineral, skin, plant, plastic or woven thread. In other applications, classification is focused more on surface properties, such as highly reflective surface, low-reflectivity one, coarse one, smooth one, and other types of surface properties.
[0117] Generally, the classification result of the neural network represents the major material of the real object 61 for the respective spatial part that is sensed by the respective SPAD pixel. For example, tiny or small objects interspersed with another larger or major material may be difficult to detect such that classification result may correspond to the major material. A finer granularity may also include a classification regarding the reflectivity properties of the material, for example, scratched or polished may be attributes reflected in the material class such that a metal may be classified as scratched metal or polished metal. The neural network may be tailored orSony Semiconductor Solutions Corporation et al.
[0118] trained to distinguish such cases. Moreover, with a finer granularity, the metal may be identified as gold, silver, copper, etc.
[0119] At 72, the mobile electronic device 50 generates material information of the real object 61 to identify texture information.
[0120] The material information is generated in this embodiment when the histogram data is processed as will be discussed in more detail under references of Fig. 4, 5 and 6. As discussed above, the neural network processes the histogram data to generate the material information by classifying the peaks in the histogram and by assigning for some peaks a material class indicating a material of the real object 61, which may also account for the reflectivity properties which may also depend on the application.
[0121] The material information of the real object 61 is assigned to the depth information. In other words, the material information represents material information of each different spatial part of the real object 61 corresponding to the spatial resolution of the ToF sensor 51. In other words, each histogram has associated material information.
[0122] At 73, the RGB camera 52 captures a color image to acquire color information of the real object 61 to generate a colored virtual object 62 in a virtual environment 63.
[0123] The processor 53 executes various procedures stored in the storage 54 to implement the functions as described herein.
[0124] At 74, the processor 53 generates, based on the depth information and the color information, a colored virtual object 62 in a virtual environment 63. The processor 53 associates the color information with a virtual representation of the real object to form the colored virtual object. The virtual environment 63 includes one or more virtual light sources and may include further virtual objects. The processor 53 fuses the depth information and the color information to generate the colored virtual object 62.
[0125] Moreover, at 74, the processor 53 performs texture mapping on the generated colored virtual object 62. The processor 53 identifies, based on the material information of each of the different spatial parts of the real object 61, texture information. The texture information represents a surface structure and organization and, as such, it influences the visual appearance of an object depending on the illumination conditions. The texture information may thus include a bidirectional reflectance distribution function (“BRDF”) to account for the material indicated in the material information. Such functions may be predefined and stored in the storage 54. TheSony Semiconductor Solutions Corporation et al.
[0126] texture information may include a description of a typical surface profile of the material indicated by the material information.
[0127] Thus, at 74, the processor 53 generates the colored virtual object 62 in the virtual environment 63 in accordance with the texture information and instructs the electronic display 55 to display the colored virtual object 62 in the virtual environment in accordance with the texture information.
[0128] Accordingly, realistic 3D rendering is achieved by considering surface material profile. For example, a statue made of stone has a rough surface. A plastic plate has a smooth or flat surface. A metallic object has a shiny surface. These surface profiles can be classified by the processing of the histogram data. By taking the surface profiles or BRDF into account by using texture information, the realistic or natural appearance of the virtual representation 62 of the real object 61 in the virtual environment 63 is improved.
[0129] At 75, the electronic display 55 receives a user input indicating a change of color information and / or material information for at least a part of the colored virtual object 62.
[0130] For example, the electronic display 55 displays a tool button 64 which the user may press using a finger or an electronic pen to open a menu for interacting with the colored virtual object 62 using predetermined tools. A tool may be a color pen with which the user can manipulate or change the color information and / or material information for at least a part of the colored virtual object 62. Then, the processor 53 receives the information of the user input and generates a colored virtual object 65 in the virtual environment 63 in accordance with the changed color information and / or material information. Of course, the processor 53 uses the texture information, in particular the texture information associated with the changed part of the colored virtual object 65 to display the colored virtual object 65 with a realistic appearance in the virtual environment 63. If the material information has changed, of course, new texture information is identified for that part based on the new material information and the new texture information is used.
[0131] Then, the processor 53 instructs the electronic display 55 to display the colored virtual object 65 in the virtual environment 63 in accordance with the change of color information and / or material information.
[0132] The changed color information and / or material information is illustrated in Fig. 1 by the dotted region 66.Sony Semiconductor Solutions Corporation et al.
[0133] For example, drawing some marks on a stone statue and metallic surface will look different because of the difference of the surface texture. The drawings on the stone statue will look rough. In contrast, the drawings on the metallic surface will look much cleaner and smoother. Fig. 3 A schematically illustrates a situation 1 in which it may be difficult to separate a target peak from other peaks, wherein two target peaks 2a and 3a overlap in a histogram.
[0134] Referring now to Fig. 3A, the situation 1 shows a first object 2 adjacent to a second object 3, however, the first object 2 and the second object 3 have slightly different depths. Moreover, the light pulse 4 hits at the edge of the first object 2 and the second object 3.
[0135] Thus, the acquired histogram shows only one peak 4a, since the two peaks 2a and 3a overlap in time such that the resulting measured histogram data indicates only one broader peak 4a.
[0136] Fig. 3B schematically illustrates a situation 5 in which it may be difficult to separate a target peak from other peaks, wherein two separated peaks are present in a histogram.
[0137] Referring now to Fig. 3B, the situation 5 shows a ToF sensor 6 which emits light to acquire depth information about an object 8, however, a glass window 7 is between the ToF sensor 6 and the object 8 such that two separated peaks are present in a histogram, and it is not clear which peak corresponds to a target.
[0138] The situations 1 and 5 of Fig. 3 A and 3B, respectively, can lead to double-peak histograms. Depending on the spatial separation between the foreground (object 2 or window 7) and background (object 3 or object 8) objects, these double peaks are temporally separated in histograms or overlapping with each other as if there is only one peak in a histogram.
[0139] It has been recognized that the situations 1 and 5 may require a different histogram data processing than the known histogram data processing.
[0140] As discussed above, in dToF (“direct time-of-fhght”), counts of detected photons at certain arrival times are described as histograms, where the x axis is the bin related to the time window and the y axis is the number of counted photons.
[0141] If there are photons coming from only one target, a single peak is observed in a histogram. Its peak position is the ToF and related to a target distance.
[0142] In some dToF sensors (e.g., used in mobile electronic devices), the peak resolution (i.e. the number of data samples to represent the peak) may be comparable low.
[0143] As a result, some peak detection algorithms may be, in some cases, less reliable. One of the most commonly used methods for such a case is a center-of-mass (“CoM”) method.Sony Semiconductor Solutions Corporation et al.
[0144] Multiple peaks can be observed under some conditions. For example, there are cases where light partially hits to edge of an object (see Fig. 3 A) or light partially goes through transparent objects (see Fig. 3B). With CoM, it may be difficult to detect accurate peak positions (indicating the depth) in multi-peak histograms, especially in a case in which peaks are overlapping with each other.
[0145] In general, there are several approaches to determine peak positions: 1) CoM is often used. In mobile electronic devices, this method is typically preferred for single-peak analysis because of its simpleness. 2) Gaussian mixture model (“GMM)” is an iterative algorithm to reproduce multiple peaks with several Gaussian distributions. 3) Derivative or wavelet analysis can be used to reveal overlapping peaks.
[0146] However, there may be some characteristics of the known methods, in some cases, related to computational speed, practical aspect (e.g., reliability, a priori knowledge about the number of peaks, and so on), and peak resolution (i.e., the number of bins to express a peak), which may be improved.
[0147] It has thus been recognized that a deep neural network (“DNN”) may be used to improve processing of the histogram data. As will be discussed in the following, the processing allows peak position prediction and peak-wise material classification from multi-peak histograms and may be applied to system-related signals (e.g., stray light from Tx to Rx, cover glass signal). Fig. 4 schematically illustrates in a flow diagram an embodiment of a method 10, which is discussed in the following.
[0148] The method 10 may be performed by the mobile electronic device or the circuitry thereof as described herein. For example, the processor 53 of Fig. 1 may perform the method 10.
[0149] At 11, input data in form of one-dimensional histogram data is acquired.
[0150] At 12, the input data is normalized to 0-1 scale.
[0151] At 13, the normalized histogram data is input to a DNN.
[0152] At 14, the DNN outputs a matrix in which xi, X2, X3 and X4 represent peak positions.
[0153] In the present embodiment, the preset number of peak positions is four, however, the present disclosure is not limited to such a case. The preset number of peak positions represents the maximum number of peak positions which can be detected in the present embodiment. In other words, zero or one or two or three peaks may be detected instead of four.Sony Semiconductor Solutions Corporation et al.
[0154] Further, si, S2, S3 and S4 of the output matrix represent peak contrasts of the peak at peak position xi, X2, X3 and X4, respectively.
[0155] Further, ci, C2, C3 and C4 of the output matrix represent peak classes of the peak at peak position xi, X2, X3 and X4, respectively.
[0156] At 15, the peak position xi, X2, X3 and X4 are converted to depth.
[0157] It has been recognized that the shape of the output matrix is very similar to object detection in two-dimensional images.
[0158] Thus, the DNN to which the histogram data is input may be based on any DNN architectures that are developed for object detection in two-dimensional images, however, the DNN should be modified to one-dimensional data because the histogram data are one-dimensional. This may be done by using one-dimensional convolutional layers with one or more one-dimensional filters, e.g., a vector or array may be used as kernel instead of a matrix typically used for the kernels. In object detection in two-dimensional images typically the bounding box position (x, y) is output to answer the question where the target is. The bounding box size (h, w) is output to answer the question how large the target is. The label is output to answer the question what the target is.
[0159] Similarly, in one-dimensional histogram analysis, the peak position corresponds to the bounding box position, the peak contrast corresponds to the bounding box size and the peak class corresponds to the label.
[0160] In the following, the output of the matrix will be discussed in more detail.
[0161] Peak position:
[0162] Determination of the peak positions is a regression problem. A proper activation function is selected for the regression problem. The activation function is properly selected for the output layer to generate real numbers ranging from zero to one.
[0163] The predicted peak positions (0-to-l scale) are converted to the unit of bins. Using system parameters and calibration parameters, the peak positions in the unit of bins can be converted to depth.
[0164] The calibration parameters (e.g., depth shift, temperature shift, and so on) are determined with a reference method (e.g., CoM) that is used for determining ground-truth data to train the DNN.Sony Semiconductor Solutions Corporation et al.
[0165] When the signal of returned light is strong, SPAD cannot react further to incoming photons because of a dead time. As a result, histogram peaks coming from the strong returned lights are deformed, which is known as a pile-up effect
[0166] It is also possible to include compensation for the pile-up effect in the DNN to get accurate peak positions that are expected from pile-up-free because histograms are deformed by the pile-up-effect in a specific way which can be learned by the DNN through proper training.
[0167] Peak contrast.
[0168] Determination of the peak contrasts is a regression problem. A proper activation function is selected for the regression problem. As mentioned above, the activation function is properly selected for the output layer to generate real numbers ranging from zero to one.
[0169] The peak contrasts (0-to-l scale) are predicted by the DNN and converted later to signal-to-noise ratio (“SNR”) by using parameters obtained in the histogram normalization process (see Fig. 5 below).
[0170] If the SNR calculated from the predicted peak contrast is lower than a preset threshold, that peak is not considered as a peak. Because of the SNR thresholding, up to N peaks, wherein N is the number of columns in the output matrix, can be simultaneously analyzed. For example, if there are two peaks in a histogram and N = 4, only si and S2 can have certain values while S3 and S4 are below a preset threshold.
[0171] The definition of the peak contrast and SNR can be changed depending on user preference or use-cases.
[0172] Peak class:
[0173] The determination of the peak classes is a multi-class classification problem. A proper activation function is selected for the classification problem. A proper activation function is selected for the output layer to generate probabilities of classes.
[0174] The histogram peak shapes (e.g., rising and falling edges of peaks) carry information about a target material.
[0175] Using the aforementioned background knowledge, peak classes (i.e. target object / material types) can be determined. The peak class shows the type of material.
[0176] Below are examples:Sony Semiconductor Solutions Corporation et al.
[0177] In a simple and general use-case, three classes may be used: not a peak (ci = 0), cover glass and stray light from Tx to Rx (ci = 1), scene or target (ci = 2).
[0178] In more advanced applications (e.g., highly secure face ID), different material classes may be used: human skin (ci = 0), rubber (ci = 1), silicon (ci = 2), not a peak (ci = 3).
[0179] The method 10 allows to accurately get one or multiple depths together with material information from dToF histograms. It provides an all-in-one solution as DNN for dToF signal processing (e.g., pile-up correction, noise elimination, multi-peak detection, material sensing). Possible applications of the method 10 include an accurate depth determination, even for challenging scenes (e.g., through transparent window), which may be achieved in any kind of general three-dimensional applications such as dToF-assisted autofocus, SLAM (“Simultaneous Localization and Mapping”), AR (“Augmented Reality”), VR (“Virtual Reality”), and so on. The material sensing, based on histogram data, enables to eliminate system-related noise (e.g., stray light and cover glass signal) and achieve advanced three-dimensional applications such as material-aware three-dimensional rendering and editing, realistic photogrammetry, secure face ID, and so on.
[0180] Fig. 5 schematically illustrates an embodiment of a peak contrast, which is discussed in the following.
[0181] The mobile electronic device as described herein may be perform the peak contrast determination as discussed in the following.
[0182] First, hmnxand hmiriare saved from a raw histogram 40, H . The maximum and minimum values in a histogram, which are denoted as hmaxand hmin, respectively. If there are K histograms, Kx2 unsigned integer data is saved. The K histograms may then be processed one after the other. Then, the histogram is normalized to get H by using hmaxand hminand give the normalized histogram 41, H, to the DNN:
[0183] fj > H hmin
[0184] H
[0185]
[0186] ~IftLmax - h,Lmi ■n '
[0187] The DNN directly outputs the peak contrast which is defined as ct= ht— hnf. It is between 0 and 1.
[0188] Then, in this embodiment, the SNR is calculated using the peak contrast, for example, as:Sony Semiconductor Solutions Corporation et al.
[0189] CJVD_ ^i(.^-max hmin)
[0190] JNK — - .
[0191]
[0192] V "■max
[0193] Fig. 6 schematically illustrates in a flow diagram an embodiment of a method 20, which is discussed in the following.
[0194] The method 20 may be performed by the mobile electronic device or the circuitry thereof as described herein. For example, the processor 53 of Fig. 1 may perform the method 20. The method 20 is based on the method 10 of Fig. 4.
[0195] At 21, a histogram, H, with nb bins is obtained.
[0196] At 22, if nb is equal to ninput (i.e. the input shape of the DNN), the histogram is normalized at 22.
[0197] At 23, if nb is lower than ninput, the histogram is up-sampled at 23 and then normalized at 22. At 24, if nb is greater than ninput, the histogram is down-sampled at 24 and then normalized at 22.
[0198] At 25, the normalized histogram H is obtained with ninput bins.
[0199] At 26, hmnxand hmiriare stored.
[0200] At 27, the normalized histogram H is input into the DNN.
[0201] At 28, the DNN outputs a matrix with peak positions xi, X2, ... , XN (N being an integer), associated peak contrasts si, S2, ... , SN and peak classes ci, C2, ... , CN.
[0202] At 29, only peaks with peak classes corresponding to one or more target classes are selected for further processing. For example, the peaks with class indicating not a peak and cover glass may be removed, while the other peak classes which indicate a material of the target are selected for further processing.
[0203] At 30, the SNR for each peak is calculated using hmaxand hminand the peak contrast of the respective peak.
[0204] At 31, only peaks with a SNR above a preset threshold are selected for further processing.
[0205] At 32, a matrix is output with peak positions xi, X2, ... , XM (M equal to or lower than N), associated peak contrasts si, S2,..., SM and peak classes ci, C2,..., CM- The matrix includes all peaks with SNR above a preset threshold and with certain target peak classes.
[0206] At 33, the peak positions xi are converted to a corresponding bin.
[0207] At 34, the bin is converted to a ToF.Sony Semiconductor Solutions Corporation et al.
[0208] At 35, the ToF is converted to a depth.
[0209] Fig. 7 schematically illustrates in a block diagram an application example 80 of the method of Fig. 4 or Fig. 6, which is discussed in the following.
[0210] In the application example 80, a transmitter 81 of the ToF sensor emits a light pulse towards an object 84 through the cover glass 83 of the ToF sensor such that a reflection from the cover glass 83 and from the object 84 reaches the receiver of the ToF sensor, thereby two peaks are present in the histogram.
[0211] The histogram is input, at 85, into the DNN which outputs a matrix indicating one peak with class of cover glass (ci = 1) and one peak with class of target (c2 = 2), thereby a method for cover glass removal is provided.
[0212] In more detail:
[0213] In some ToF sensors, a cover glass is placed to cover both Tx 81 (transmitter) and Rx 82 (receiver). The cover glass 83 creates a visible signal in a dToF histogram. This signal may interfere with an actual signal coming from a target 84 of interest, resulting in a multi-peak histogram.
[0214] The DNN can detect multiple peaks at the same time. Moreover, with a material sensing technique, the cover glass 84 can be easily identified by letting the DNN learn how the cover glass 83 alters the peak shapes in the histograms. The peak class shows which peak comes from the cover glass 83, and which one is from the actual target 84 in a scene. In most cases, the signal from the actual target 84 is of interest. With the help of the peak class, only the signal from the actual target 84 can be selected.
[0215] Only a peak with non-zero peak contrast (SNR) and peak class of target (ci = 2) is selected. The cover glass signal (ci = 1) can be ignored. In the example above, only X2 will be converted to depth. The obtained depth can be used for any three-dimensional application.
[0216] Fig. 8 schematically illustrates in a block diagram an application example 90 of the method of Fig. 4 or Fig. 6, which is discussed in the following.
[0217] The application example 90 is an application of dToF-assisted RGB camera auto-focus. The RGB camera 300 has an image sensor 301 and a lens 302.
[0218] In the application example 90, a transmitter 91 of the ToF sensor emits a light pulse towards an object 94 through a glass window 93 such that a reflection from the glass window 93 and fromSony Semiconductor Solutions Corporation et al.
[0219] the object 94 reaches the receiver of the ToF sensor 92, thereby two peaks are present in the histogram and may overlap.
[0220] The histogram is input, at 95, into the DNN which outputs a matrix indicating one peak with class of glass window (ci = 1) and one peak with class of target (c2 = 2).
[0221] At 96, only the peak with target peak class (c2 = 2) is converted to depth.
[0222] At 97, the depth is converted to a position of the lens 302.
[0223] At 98, the target position of the lens 302 is selected, thereby providing a method for improved auto-focus (“AF”). The RGB camera lens 302 is quickly moved to the determined target position without scanning the position of the RGB camera lens 302.
[0224] In more detail:
[0225] When there is a scene where a target object 94 is placed behind transparent window, two peaks appear in a dToF histogram, i.e. one from the window and one from target. In some cases, they could possibly overlap each other. It may be difficult for some algorithms to analyze multi-peak histograms, especially overlapping peaks.
[0226] However, the determination of the correct depths is crucial to estimate an accurate lens position of an RGB camera. A failure to estimate the correct RGB camera lens position results in a blurry image.
[0227] In the same scene, there is a possibility that a strong signal from the glass window 93 is measured at a certain incident angle. The strong signal may distort a histogram shape because of a pile-up effect which may prevent from an accurate depth determination.
[0228] Even if succeeded in obtaining two correct depths in the same scene, it would not be clear which one to select for the AF application.
[0229] Hence, depth values of the glass window 93 and the target 94 are accurately determined by the DNN in real-time even when two peaks overlap with each other. A peak deformation caused by the pile-up effect can be also taken into account and compensated. With the obtained depth values, correct RGB camera lens positions for focusing on the glass window 93 and the target 94 behind it can be determined via a thin-lens equation. With the peak class in the DNN output, it is possible to distinguish the glass window 93 and the target 94 behind it. Peak data with the peak class of the target 94 behind the glass window 93 will be considered (ci = 2).Sony Semiconductor Solutions Corporation et al.
[0230] Fig. 9 schematically illustrates in a flow diagram an embodiment of a method 100, which is discussed in the following.
[0231] The method 100 may be performed by the mobile electronic device as described herein, such as the mobile electronic device of Fig. 1.
[0232] At 101, histogram data is acquired representing depth information of a scene, as discussed herein.
[0233] At 102, the histogram data is input into a neural network, wherein the neural network is configured to predict a peak position, a peak contrast and a peak class for each of a preset number of peaks, as discussed herein.
[0234] At 103, a depth is determined for each peak with a peak contrast above a preset threshold, as discussed herein.
[0235] At 104, the peak class is assigned to the respective depth, as discussed herein.
[0236] Fig. 10 schematically illustrates in a flow diagram an embodiment of a method 200, which is discussed in the following.
[0237] The method 200 may be performed by the mobile electronic device as described herein, such as the mobile electronic device of Fig. 1.
[0238] At 201, depth information and color information of a real object is acquired to generate a colored virtual object in a virtual environment, as discussed herein.
[0239] At 202, material information of the real object is generated to identify texture information, as discussed herein.
[0240] At 203, the colored virtual object is displayed in the virtual environment in accordance with the texture information, as discussed herein.
[0241] At 204, a user input is received indicating a change of color information and / or material information for at least a part of the colored virtual object, as discussed herein.
[0242] At 205, the colored virtual object is displayed in the virtual environment in accordance with the change of color information and / or material information, as discussed herein.
[0243] Fig. 11 schematically illustrates in a block diagram an embodiment of a multi-purpose computer 130 which can be used for implementing a mobile electronic device.Sony Semiconductor Solutions Corporation et al.
[0244] The computer 130 can be implemented such that it can basically function as any type of mobile electronic device as described herein. The computer has components 131 to 142, which can form a circuitry, such as any one of the circuitries of the mobile electronic device as described herein. Embodiments which use software, firmware, programs or the like for performing the methods as described herein can be installed on computer 130, which is then configured to be suitable for the concrete embodiment.
[0245] The computer 130 has a CPU 131 (Central Processing Unit), which can execute various types of procedures and methods as described herein, for example, in accordance with programs stored in a read-only memory (ROM) 132, stored in a storage 137 and loaded into a random-access memory (RAM) 133, stored on a medium 140 which can be inserted in a respective drive 139, etc.
[0246] The computer 130 has a GPU 141 (Graphical Processing Unit), which can execute various types of procedures and methods as described herein, for example, in accordance with programs stored in a read-only memory (ROM) 132, stored in a storage 137 and loaded into a random-access memory (RAM) 133, stored on a medium 140 which can be inserted in a respective drive 139, etc.
[0247] The CPU 131 and the GPU 141 can commonly execute various types of procedures and methods as described herein, for example, in accordance with programs stored in a read-only memory (ROM) 132, stored in a storage 137 and loaded into a random-access memory (RAM) 133, stored on a medium 140 which can be inserted in a respective drive 139, etc.
[0248] The CPU 131, the ROM 132, the RAM 133 and the GPU 141 are connected with a bus 142, which in turn is connected to an input / output interface 134. The number of CPUs, GPUSs, memories and storages is only exemplary, and the skilled person will appreciate that the computer 130 can be adapted and configured accordingly for meeting specific requirements which arise, when it functions as a mobile electronic device.
[0249] At the input / output interface 134, several components are connected: an input 135, an output 136, the storage 137, a communication interface 138 and the drive 139, into which a medium 140 (compact disc, digital video disc, compact flash memory, or the like) can be inserted.
[0250] The input 135 can be a pointer device (mouse, graphic table, or the like), a keyboard, a microphone, a camera, a time-of-flight device, an event sensor, a touchscreen, etc.Sony Semiconductor Solutions Corporation et al.
[0251] The output 136 can have a display (liquid crystal display, cathode ray tube display, light emittance diode display, etc.), loudspeakers, etc.
[0252] The storage 137 can have a hard disk, a solid-state drive and the like.
[0253] The communication interface 138 can be adapted to communicate, for example, via a local area network (LAN), wireless local area network (WLAN), mobile telecommunications system (GSM, UMTS, LTE, NR etc.), Bluetooth, infrared, etc. It should be noted that the description above only pertains to an example configuration of computer 130.
[0254] It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding.
[0255] All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.
[0256] In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure.
[0257] Note that the present technology can also be configured as described below.
[0258] (1) An electronic device, wherein the electronic device includes circuitry configured to: receive or acquire depth information and color information of a real object to generate a colored virtual object in a virtual environment;
[0259] generate material information of the real object to identify texture information; and display the colored virtual object in the virtual environment in accordance with the texture information.
[0260] (2) The electronic device of (1), wherein the circuitry is configured to receive a user input indicating a change of color information for at least a part of the colored virtual object and display the colored virtual object in the virtual environment in accordance with the change of color information.
[0261] (3) The electronic device of (1) or (2), wherein the circuitry is configured to receive a user input indicating a change of material information for at least a part of the colored virtual objectSony Semiconductor Solutions Corporation et al.
[0262] and display the colored virtual object in the virtual environment in accordance with the change of material information.
[0263] (4) The electronic device of any one of (1) to (3), wherein the circuitry includes a time-of-flight senor which includes a plurality of time-of-flight pixels, and wherein the time-of-flight sensor is configured to acquire histogram data with each of the plurality of time-of-flight pixels to acquire the depth information.
[0264] (5) The electronic device of (4), wherein the circuitry is configured to input the histogram data of a time-of-flight pixel into a neural network, wherein the neural network is configured to predict a peak position and a peak contrast for each of a preset number of peaks.
[0265] (6) The electronic device of (5), wherein the neural network is configured to classify each peak for generating the material information.
[0266] (7) The electronic device of (6), wherein the material information is a peak class indicating one of: not a peak, cover glass and a material of the real object.
[0267] (8) The electronic device of (7), wherein the circuitry is configured to select the peaks with a peak class corresponding to the material of the real object.
[0268] (9) The electronic device of (8), wherein the circuitry is configured to determine, based on the peak position, a depth for each selected peak with a signal-to-noise ratio above a preset threshold, wherein the signal-to-noise ratio is based on the peak contrast, and assign the material of the real object to the depth.
[0269] (10) The electronic device of any one of (5) to (9), wherein the neural network includes a onedimensional convolutional layer which uses one or more one-dimensional filters.
[0270] (11) A method, wherein the method includes:
[0271] receiving or acquiring depth information and color information of a real object to generate a colored virtual object in a virtual environment;
[0272] generating material information of the real object to identify texture information; and displaying the colored virtual object in the virtual environment in accordance with the texture information.
[0273] (12) The method of (11), including receiving a user input indicating a change of color information for at least a part of the colored virtual object and displaying the colored virtual object in the virtual environment in accordance with the change of color information.Sony Semiconductor Solutions Corporation et al.
[0274] (13) The method of (11) or (12), including receiving a user input indicating a change of material information for at least a part of the colored virtual object and displaying the colored virtual object in the virtual environment in accordance with the change of material information. (14) The method of any one of (11) to (13), including acquiring histogram data with each of a plurality of time-of-flight pixels to acquire the depth information.
[0275] (15) The method of (14), including inputting the histogram data of a time-of-flight pixel into a neural network, wherein the neural network is configured to predict a peak position and a peak contrast for each of a preset number of peaks.
[0276] (16) The method of (15), wherein the neural network is configured to classify each peak for generating the material information.
[0277] (17) The method of (16), wherein the material information is a peak class indicating one of: not a peak, cover glass and a material of the real object.
[0278] (18) The method of (17), including selecting the peaks with a peak class corresponding to the material of the real object.
[0279] (19) The method of (18), including determining, based on the peak position, a depth for each selected peak with a signal-to-noise ratio above a preset threshold, wherein the signal-to-noise ratio is based on the peak contrast, and assign the material of the real object to the depth.
[0280] (20) The method of any one of (15) to (19), wherein the neural network includes a onedimensional convolutional layer which uses one or more one-dimensional filters.
[0281] (21) A computer program comprising program code causing a computer to perform the method according to anyone of (11) to (20), when being carried out on a computer.
[0282] (22) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (11) to (20) to be performed.
[0283] (23) An electronic device, wherein the mobile electronic device includes circuitry configured to:
[0284] receive or acquire histogram data representing depth information of a scene;
[0285] input the histogram data into a neural network, wherein the neural network is configured to predict a peak position and a peak contrast for each of a preset number of peaks; and determine a depth for each peak with a signal-to-noise ratio above a preset threshold, wherein the signal-to-noise ratio is based on the peak contrast.Sony Semiconductor Solutions Corporation et al.
[0286] (24) The electronic device of (23), wherein the neural network is configured to classify each peak, and wherein the circuitry is configured to assign the peak class to the respective depth. (25) The electronic device of (24), wherein the peak class indicates one of: not a peak, cover glass and target.
[0287] (26) The electronic device of (24), wherein the peak class indicates one of: not a peak, glass window and target.
[0288] (27) The electronic device of (24), wherein the peak class indicates one of: not a peak, cover glass and a material of a target.
[0289] (28) The electronic device of any one of (23) to (27), wherein the neural network includes a one-dimensional convolutional layer which uses one or more one-dimensional filters.
[0290] (29) The electronic device of any one of (23) to (28), wherein the circuitry is configured to normalize the histogram data before input into the neural network.
[0291] (30) The electronic device of (29), wherein the peak contrast is given by a difference between a height of the respective peak and a noise floor of the normalized histogram data.
[0292] (31) A method, wherein the method includes:
[0293] receiving or acquiring histogram data representing depth information of a scene; inputting the histogram data into a neural network, wherein the neural network is configured to predict a peak position and a peak contrast for each of a preset number of peaks; and
[0294] determining a depth for each peak with a signal-to-noise ratio above a preset threshold, wherein the signal-to-noise ratio is based on the peak contrast.
[0295] (32) The method of (31), wherein the neural network is configured to classify each peak, and wherein the circuitry is configured to assign the peak class to the respective depth.
[0296] (33) The method of (32), wherein the peak class indicates one of: not a peak, cover glass and target.
[0297] (34) The method of (32), wherein the peak class indicates one of: not a peak, glass window and target.
[0298] (35) The method of (32), wherein the peak class indicates one of: not a peak, cover glass and a material of a target.Sony Semiconductor Solutions Corporation et al.
[0299] (36) The method of any one of (31) to (35), wherein the neural network includes a onedimensional convolutional layer which uses one or more one-dimensional filters.
[0300] (37) The method of any one of (31) to (36), wherein the circuitry is configured to normalize the histogram data before input into the neural network.
[0301] (38) The method of (37), wherein the peak contrast is given by a difference between a height of the respective peak and a noise floor of the normalized histogram data.
[0302] (39) A computer program comprising program code causing a computer to perform the method according to any one of (31) to (38), when being carried out on a computer.
[0303] (40) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to any one of (31) to (38) to be performed.
[0304] (41) A method, including:
[0305] receiving or acquiring depth information and color information of a real object; generating a colored virtual object in a virtual environment by estimating a material by neural network circuitry operating on the depth information of the real object to identify texture information; and
[0306] assigning texture information to the colored virtual object in the virtual environment. (42) The method of (41), further including processing the colored virtual object to change color properties based on the texture information.
[0307] (43) A computer program comprising program code causing a computer to perform the method according to (41) or (42), when being carried out on a computer.
[0308] (44) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to (41) or (42) to be performed.
[0309] (45) An electronic device, including circuitry configured to:
[0310] receive or acquire depth information and color information of a real object; generate a colored virtual object in a virtual environment by estimating a material by neural network circuitry operating on the depth information of the real object to identify texture information; and
[0311] assign texture information to the colored virtual object in the virtual environment.Sony Semiconductor Solutions Corporation et al.
[0312] (46) The electronic device of (45), wherein the circuitry is configured to process the colored virtual object to change color properties based on the texture information.
Claims
Sony Semiconductor Solutions Corporation et al.CLAIMS1. An electronic device, comprising circuitry configured to:acquire depth information and color information of a real object to generate a colored virtual object in a virtual environment;generate material information of the real object to identify texture information; and display the colored virtual object in the virtual environment in accordance with the texture information.
2. The electronic device of claim 1, wherein the circuitry is configured to receive a user input indicating a change of color information for at least a part of the colored virtual object and display the colored virtual object in the virtual environment in accordance with the change of color information.
3. The electronic device of claim 1, wherein the circuitry is configured to receive a user input indicating a change of material information for at least a part of the colored virtual object and display the colored virtual object in the virtual environment in accordance with the change of material information.
4. The electronic device of claim 1, wherein the circuitry includes a time-of-flight senor which includes a plurality of time-of-flight pixels, and wherein the time-of-flight sensor is configured to acquire histogram data with each of the plurality of time-of-flight pixels to acquire the depth information.
5. The electronic device of claim 4, wherein the circuitry is configured to input the histogram data of a time-of-flight pixel into a neural network, wherein the neural network is configured to predict a peak position and a peak contrast for each of a preset number of peaks.
6. The electronic device of claim 5, wherein the neural network is configured to classify each peak for generating the material information.
7. The electronic device of claim 6, wherein the material information is a peak class indicating one of not a peak, cover glass and a material of the real object.
8. The electronic device of claim 7, wherein the circuitry is configured to select the peaks with a peak class corresponding to the material of the real object.
9. The electronic device of claim 8, wherein the circuitry is configured to determine, based on the peak position, a depth for each selected peak with a signal-to-noise ratio above a presetSony Semiconductor Solutions Corporation et al.threshold, wherein the signal-to-noise ratio is based on the peak contrast, and assign the material of the real object to the depth.
10. The electronic device of claim 5, wherein the neural network includes a one-dimensional convolutional layer which uses one or more one-dimensional filters.
11. A method, comprising:acquiring depth information and color information of a real object to generate a colored virtual object in a virtual environment;generating material information of the real object to identify texture information; and displaying the colored virtual object in the virtual environment in accordance with the texture information.
12. The method of claim 11, comprising receiving a user input indicating a change of color information for at least a part of the colored virtual object and displaying the colored virtual object in the virtual environment in accordance with the change of color information.
13. The method of claim 11, comprising receiving a user input indicating a change of material information for at least a part of the colored virtual object and displaying the colored virtual object in the virtual environment in accordance with the change of material information.
14. The method of claim 11, comprising acquiring histogram data with each of a plurality of time-of-flight pixels to acquire the depth information.
15. The method of claim 14, comprising inputting the histogram data of a time-of-flight pixel into a neural network, wherein the neural network is configured to predict a peak position and a peak contrast for each of a preset number of peaks.
16. The method of claim 15, wherein the neural network is configured to classify each peak for generating the material information.
17. The method of claim 16, wherein the material information is a peak class indicating one of: not a peak, cover glass and a material of the real object.
18. The method of claim 17, comprising selecting the peaks with a peak class corresponding to the material of the real object.
19. The method of claim 18, comprising determining, based on the peak position, a depth for each selected peak with a signal-to-noise-ratio above a preset threshold, wherein the signal-to-noise ratio is based on the peak contrast, and assign the material of the real object to the depth.Sony Semiconductor Solutions Corporation et al.
20. The method of claim 15, wherein the neural network includes a one-dimensional convolutional layer which uses one or more one-dimensional filters.