Information processing device and information processing method

The method addresses inefficiencies in depth estimation by modeling error distribution for sparse depth data, enhancing training efficiency and noise robustness in learning models.

WO2026009705A1PCT designated stage Publication Date: 2026-01-08SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/021890
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-05
Filing Date
2025-06-18
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

In depth estimation using machine learning, the inefficiency of large network sizes and the unclear relationship between noise in RGB images and correct depth information make it difficult to determine appropriate augmentation, leading to suboptimal training results.

Method used

An information processing device and method that generate learning images with added errors by modeling the error distribution of sparse depth information using an error modeling function, allowing for efficient data augmentation and training of a noise-robust learning model.

Benefits of technology

The method efficiently generates training data with added errors, enabling the training of a smaller network size learning model that is robust to noise, improving the accuracy of depth estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025021890_08012026_PF_FP_ABST
    Figure JP2025021890_08012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an information processing device and an information processing method that make it possible to generate a training image to which an error has been efficiently added. The information processing device comprises: an error modeling unit for generating, from sparse depth information having a depth value or parallax stored in some pixels and ground-truth depth information representing a ground truth value thereof, an error modeling function modeling the error distribution of the depth value or parallax included in the sparse depth information; and an error addition unit that generates a random error by using the error modeling function as an error adding function and generates new depth information including the random error. The technology of the present disclosure is applicable to, for example, an information processing device for training a learning model using training data.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device and information processing method

[0001] The present disclosure relates to an information processing device and an information processing method, and more particularly to an information processing device and an information processing method that are capable of efficiently generating learning images to which errors are added.

[0002] In depth estimation using machine learning, for example, using two stereo images (RGB images) as inputs for a learning model results in inefficiency due to the large network size. Furthermore, the learning results are often influenced by the texture of the stereo images.

[0003] For example, there is a technique for generating dense depth information from sparse depth information by generating a learning model that estimates dense depth information from sparse depth information (see, for example, Patent Document 1).

[0004] In training a learning model, noise is often added to an RGB image to perform augmentation (data expansion) of the training image in order to increase noise resistance. Patent Literature 2 discloses a method for learning a noise-adding function from a ground truth depth map and a noisy depth map.

[0005] JP 2020-123114 A JP 2018-109976 A

[0006] However, in depth estimation using machine learning, the relationship between noise in RGB images and correct depth information is unclear, making it difficult to determine what kind of augmentation should be performed.

[0007] The present disclosure has been made in consideration of such circumstances, and makes it possible to efficiently generate learning images to which errors have been added.

[0008] An information processing device according to one aspect of the present disclosure includes an error modeling unit that generates an error modeling function that models the error distribution of the depth values ​​or disparities contained in the sparse depth information, based on sparse depth information in which depth values ​​or disparities are stored in some pixels and correct depth information that is the correct value of the sparse depth information, and an error assigning unit that generates a random error using the error modeling function as an error assigning function and generates new depth information that includes the random error.

[0009] An information processing method according to one aspect of the present disclosure includes generating an error modeling function that models the error distribution of the depth values ​​or disparities contained in sparse depth information, from the sparse depth information in which depth values ​​or disparities are stored in some pixels and correct depth information that is the correct value of the sparse depth information; and generating new depth information that includes the random errors by using the error modeling function as an error assignment function.

[0010] In one aspect of the present disclosure, an error modeling function that models the error distribution of the depth values ​​or disparities contained in the sparse depth information is generated from sparse depth information in which depth values ​​or disparities are stored in some pixels and correct depth information that is the correct value of the sparse depth information, and a random error is generated using the error modeling function as an error assignment function, and new depth information containing the random error is generated.

[0011] The information processing device according to one aspect of the present disclosure can be realized by causing a computer to execute a program. The program to be executed by the computer can be provided by transmitting it via a transmission medium or by recording it on a recording medium.

[0012] The information processing device may be an independent device or an internal block constituting a single device.

[0013] FIG. 1 is a block diagram illustrating a configuration example of an information processing device according to a first embodiment of the present disclosure. FIG. 2 is a diagram illustrating input learning data input to a learning image input unit. FIG. 3 is a diagram illustrating the processing of an error modeling unit. FIG. 4 is a diagram summarizing the relationship between the correct distance and the standard deviation of the error, with the focal length being the parameter. FIG. 5 is a diagram illustrating an error modeling function that models the error distribution for each focal length. FIG. 6 is a flowchart illustrating error modeling processing by the error modeling unit. FIG. 7 is a flowchart illustrating data extension processing by the error assigning unit. FIG. 8 is a block diagram illustrating a hardware configuration example of a computer as the information processing device of FIG. 1. FIG. 9 is a block diagram illustrating a configuration example of an information processing system according to a second embodiment of the present disclosure.

[0014] Hereinafter, modes for carrying out the technology of the present disclosure (hereinafter referred to as embodiments) will be described with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations will be assigned the same reference numerals to avoid redundant description. The description will be given in the following order: 1. Configuration example of the first embodiment 2. Details of the learning data generation process 3. Computer configuration example 4. Configuration example of the second embodiment

[0015] 1. Configuration Example of First Embodiment FIG. 1 shows a configuration example of an information processing device according to a first embodiment of the present disclosure.

[0016] 1 is a device that performs a process of learning a learning model that estimates dense depth information from sparse depth information. More precisely, learning the learning model is a process of obtaining parameters of a deep neural network (DNN) as the learning model.

[0017] The information processing device 1 includes a learning image input unit 21 , an error modeling unit 22 , an error assigning unit 23 , and a learning unit 24 .

[0018] The training image input unit 21 accepts training data input from the outside and outputs it to the error modeling unit 22 and the error assignment unit 23. Hereinafter, the training data input to the training image input unit 21 will also be referred to as input training data to distinguish it from training data (extended training data) generated by a data extension process described later.

[0019] The training data input to the training image input unit 21 is data including a plurality of pairs of depth information composed of a parallax image or a depth image and ground truth depth information, which is the ground truth of the depth information. A parallax image is an image in which the parallax calculated from a right RGB image and a left RGB image, which are stereo images, is stored as a pixel value for each pixel. A depth image is an image in which a depth value calculated from the parallax is stored as a pixel value for each pixel. However, the depth information input to the training image input unit 21 has the same resolution as the right RGB image and the left RGB image, but is sparse depth information in which depth values ​​or parallax are stored only in some pixels. Specifically, highly reliable parallax or depth values, in which the reliability of the calculated parallax or depth value is equal to or greater than a predetermined value, are stored as pixel values, and low-reliability pixels are stored with a predetermined pixel value (e.g., zero). Note that in this embodiment, for convenience, pixels in which a predetermined pixel value is stored as a low reliability pixel are referred to as pixels in which no pixel value is stored.

[0020] The training image input unit 21 accepts input of sparse depth information and correct depth information and outputs them to the error modeling unit 22 and the error assigning unit 23. Alternatively, the training image input unit 21 may accept input of the right RGB image, the left RGB image, and correct depth information, calculate depth values ​​or disparities and reliability by itself, generate sparse depth information, and output it together with the correct depth information to the error modeling unit 22 and the error assigning unit 23. Alternatively, depth information in which depth values ​​or disparities are stored for all pixels, including pixels with low reliability, and correct depth information may be input to the training image input unit 21, which may generate sparse depth information based on the reliability and output training data consisting of pairs of sparse depth information and correct depth information to the error modeling unit 22 and the error assigning unit 23. Well-known methods can be used to calculate the depth values ​​or disparities and reliability. For example, a method of calculating depth values ​​or disparities using block matching may be used. The confidence level is generally higher in the edge region of the object and lower in the flat region.

[0021] The error modeling unit 22 analyzes errors in depth values ​​or disparities included in the sparse depth information using input training data supplied from the training image input unit 21. As a result of the error analysis, the error modeling unit 22 generates an error modeling function that models the error distribution of depth values ​​or disparities. The error modeling unit 22 outputs the generated error modeling function and parameters used in the error modeling function to the error assigning unit 23.

[0022] The error assigning unit 23 performs augmentation (data augmentation) using the input training data supplied from the training image input unit 21 and the error modeling function from the error modeling unit 22. More specifically, the error assigning unit 23 generates a random error by using the error modeling function generated by the error modeling unit 22 as an error assigning function that assigns an error during data augmentation. The error assigning unit 23 then assigns the generated error to the correct depth information of the input training data supplied from the training image input unit 21, thereby generating new depth information containing a random error (hereinafter referred to as new depth information). This generates pairs of new depth information and correct depth information. The multiple pairs of new depth information and correct depth information generated here are referred to as augmented training data to distinguish them from the input training data. The new depth information is also sparse depth information in which depth values ​​or disparities are stored only in highly reliable pixels. The error assigning unit 23 outputs the learning data including the input learning data and the extended learning data to the learning unit 24 .

[0023] The learning unit 24 uses the learning data supplied from the error assigning unit 23 to learn a learning model that estimates dense depth information from sparse depth information. The learning model may be, for example, a convolutional neural network (CNN), but other models may also be used. Any method may also be used for learning the learning model. For example, the method disclosed in Patent Document 1, a prior art document, may be used to learn the learning model that estimates dense depth information from sparse depth information.

[0024] When the right RGB image and the left RGB image have been input to the training image input unit 21, the training unit 24 may add the right RGB image and the left RGB image as input data for training. In addition, the reliability of the depth information may be added as input data for training.

[0025] As described above, the information processing device 1 generates expanded training data consisting of new depth information and correct depth information by expanding the data from the sparse depth information and correct depth information that are input training data using the error modeling function generated by the error modeling unit 22 as an error assignment function.The information processing device 1 then uses the input training data and the expanded training data to train a training model that estimates dense depth information from the sparse depth information.The information processing device 1 analyzes the error distribution from a small amount of training data set and models an optimal error distribution, thereby efficiently assigning errors and generating a large amount of training data set, thereby enabling efficient training.

[0026] 2. Details of Learning Data Generation Process The process performed by the information processing device 1 will now be described in detail.

[0027] FIG. 2 is a diagram illustrating input learning data input to the learning image input unit 21. As shown in FIG.

[0028] The training data input to the training image input unit 21 is composed of data having a plurality of pairs of sparse depth information 32 and ground truth depth information 33, which is the ground truth of the sparse depth information 32, as shown in FIG. 2. The sparse depth information 32 is generated from a right RGB image 31R and a left RGB image 31L, which are stereo images. The training image input unit 21 may also generate the sparse depth information 32 from the right RGB image 31R and the left RGB image 31L. Pixels represented in black in the sparse depth information 32 shown in FIG. 2 are pixels with low reliability and represent pixels for which no pixel value is stored.

[0029] The right RGB image 31R and the left RGB image 31L shown in FIG. 2 are data obtained from “Dataset” at <URL: https: / / github.com / abhijithpunnappurath / dual-pixel-defocus-disparity?tab=readme-ov-file>.

[0030] FIG. 3 is a diagram for explaining the processing of the error modeling unit 22. As shown in FIG.

[0031] The error modeling unit 22 calculates an error from a pair of sparse depth information 32 and correct depth information 33, and generates an error map 34. The error modeling unit 22 calculates the difference between the sparse depth information 32 and the correct depth information 33 only for pixels in which pixel values ​​of the sparse depth information 32 are stored, and generates the error map 34. In the example of FIG. 3 , an error map 34-1 is generated from a pair of sparse depth information 32-1 and correct depth information 33-1. An error map 34-2 is generated from a pair of sparse depth information 32-2 and correct depth information 33-2. An error map 34-3 is generated from a pair of sparse depth information 32-3 and correct depth information 33-3. Similarly, error maps 34 are generated for all pairs of sparse depth information 32 and correct depth information 33 supplied from the training image input unit 21.

[0032] The error modeling unit 22 generates an error modeling function that models the error distribution using the generated multiple error maps 34 (error maps 34-1, 34-2, 34-3, ...). First, the error modeling unit 22 uses the multiple error maps 34 to calculate the standard deviation by assuming a predetermined error distribution for each predetermined parameter, and summarizes the relationship between the correct distance and the standard deviation of the error distribution.

[0033] FIG. 4 is a diagram summarizing the relationship between the correct distance and the standard deviation of the error for each focal length using multiple error maps 34, with the parameter being the focal length. In FIG. 4, fd represents the focal length, and FIG. 4 summarizes the relationship between the correct distance and the standard deviation of the error for seven focal lengths: fd = 1000 [mm], fd = 1250 [mm], fd = 2000 [mm], fd = 2500 [mm], fd = 5000 [mm], fd = 7500 [mm], and fd = 10000 [mm]. The horizontal axis of FIG. 4 represents the correct distance [mm], and the vertical axis represents the standard deviation of a predetermined distribution assumed as the error distribution. Examples of the assumed error distribution include a Gaussian distribution and a Laplace distribution.

[0034] Next, the error modeling unit 22 generates an error modeling function that models the error distribution.

[0035] Figure 5 shows the error modeling function y f is generated and superimposed on the error data shown in Figure 4. Note that the unit of the correct distance on the horizontal axis z in Figure 4 is expressed in [mm], but in Figure 5 it is expressed in [m]. The vertical axis y in Figure 5 represents the standard deviation of a predetermined distribution assumed as the error distribution.

[0036] Error modeling function y f can be generated using the method disclosed in Patent Document 2, for example. Alternatively, the error modeling function y can be generated by a function identification process called PhySO (Physical Symbolic Optimization). f PhySO is a program that infers a free-form symbolic analytical function that fits y=f(x) when given data (x, y). It eliminates physically impossible solutions and limits the degrees of freedom, resulting in good performance. PhySO is disclosed in "https: / / github.com / WassimTenachi / PhySO" and "https: / / arxiv.org / pdf / 2303.03192.pdf".

[0037] For example, the error modeling function y , which represents the standard deviation of the error for each focal length shown in FIG. f The following equation (1) was generated: In equation (1), C1, C2, and C3 are predetermined constants, and f is a fixed value representing the focal length (fd). z is the error modeling function y f is a variable that represents the correct distance.

[0038] Finally, the error modeling unit 22 calculates the generated error modeling function y f and its error modeling function y f 4 and 5, the error modeling function y f The focal length value, which is a parameter summarizing the error distribution, is output to the error assigning unit 23 .

[0039] In the above example, the error assigning unit 23 calculated an error modeling function using focal length as an error distribution parameter as the error assigning function for assigning an error. However, an error modeling function using data other than focal length as an error distribution parameter may also be calculated. For example, optical system setting values ​​and sensor setting values ​​can be used as parameters for summarizing the error distribution. Examples of optical system setting values ​​that can be used as error distribution parameters include the focal length described above, as well as f-number and lens aperture. Examples of sensor setting values ​​that can be used as error distribution parameters include the pixel pitch or pixel size, sensitivity characteristics, and spectral characteristics of the image sensor. The error modeling function may be a function using one parameter, as in Equation (1), or a function using two or more parameters.

[0040] <Error Modeling Process> Next, the error modeling process by the error modeling unit 22 will be described with reference to the flowchart in Fig. 6. This process starts, for example, when input training data including a plurality of pairs of sparse depth information and correct depth information is supplied from the training image input unit 21.

[0041] First, in step S1, the error modeling unit 22 calculates the difference between the sparse depth information storing pixel values ​​and the correct depth information for each pixel, and generates an error map. By generating an error map for each pair of the sparse depth information and the correct depth information, multiple error maps are generated.

[0042] In step S2, the error modeling unit 22 uses the generated error maps to calculate the standard deviation by assuming a predetermined error distribution for each predetermined parameter, and summarizes the relationship between the correct distance and the standard deviation of the error distribution.

[0043] In step S3, the error modeling unit 22 generates an error modeling function that models the error distribution.

[0044] In step S4, the error modeling unit 22 outputs the generated error modeling function and the parameters summarizing the error modeling function to the error assigning unit 23, and the error modeling process ends.

[0045] <Data Expansion Process> Next, with reference to the flowchart in Fig. 7 , a data expansion process for increasing the training data using the error modeling function generated by the error modeling unit 22 will be described. This process is started, for example, when the error modeling function and parameters are supplied from the error modeling unit 22 to the error assigning unit 23. In the process in Fig. 7 , the error modeling function of equation (1) using focal length as a parameter will be described as being supplied from the error modeling unit 22.

[0046] First, in step S21, the error assigning unit 23 selects a predetermined pair from among a plurality of pairs of sparse depth information and correct depth information, which are input training data supplied from the training image input unit 21.

[0047] In step S22, the error assigning unit 23 sets a predetermined pixel, the pixel value of which is stored in the selected set of sparse depth information, as a pixel of interest.

[0048] In step S23, the error assigning unit 23 acquires the correct distance of the pixel of interest from the correct depth information.

[0049] In step S24, the error assigning unit 23 randomly selects the focal length f, which is a parameter, from six types, for example, fd=1.0 [m], fd=1.25 [m], fd=2.0 [m], fd=2.5 [m], fd=5 [m], and fd=10 [m], and calculates the error modeling function y f The correct distance of the pixel of interest is substituted into the variable z of f In addition, in a use case where the focal length that can be taken by the correct depth information is given in the range of, for example, fd = 1.0 to 2.5 [m], the focal length f is randomly selected from the range of fd = 1.0 to 2.5 [m], and the error assignment function y f The standard deviation can be calculated.

[0050] In step S25, the error assigning unit 23 calculates the correct distance of the pixel of interest as the average μ and the error assigning function y f The probability distribution (μ, σ) with the standard deviation of 2 ) is used to calculate a pixel value including a random error, and set it as the pixel value of the pixel of interest in the new depth information. 2 ) is a probability distribution such as a Gaussian distribution or a Laplace distribution assumed as the error distribution.

[0051] In step S26, the error assigning unit 23 determines whether all pixels whose pixel values ​​are stored in the selected sparse depth information have been set as pixels of interest. If it is determined in step S26 that all pixels have not yet been set as pixels of interest, the process proceeds to step S27, where the error assigning unit 23 sets the next pixel in the sparse depth information as the pixel of interest. Then, the process returns to step S23, and the processes of steps S23 to S26 described above are repeated.

[0052] If it is determined in step S26 that all pixels have been set as pixels of interest, this means that one piece of new depth information has been completed. This new depth information is also sparse depth information, since it is an image that stores pixel values ​​including random errors for pixels in which pixel values ​​of the selected sparse depth information are stored. If it is determined in step S26 that all pixels have been set as pixels of interest, the process proceeds to step S28, and the error assigning unit 23 stores a pair of the generated new depth information and the corresponding normal depth information in a predetermined storage unit such as an internal memory.

[0053] In step S29, the error adding unit 23 determines whether a specified number of new depth information pieces have been generated for the selected sparse depth information. The number of new depth information pieces to be generated for one piece of sparse depth information is determined in advance. For example, if a setting is made to generate 10 new depth information pieces for one piece of sparse depth information, the error adding unit 23 determines whether 10 new depth information pieces have been generated.

[0054] If it is determined in step S29 that the specified number of new depth information pieces have not yet been generated for the selected sparse depth information pieces, the process returns to step S22, and the above-described processes of steps S22 to S29 are repeated. By the processes of steps S22 to S29, one more piece of new depth information piece is generated.

[0055] Then, if it is determined in step S29 that the specified number of new depth information has been generated, the processing proceeds to step S30, and the error assignment unit 23 determines whether new depth information has been generated for all pairs of sparse depth information and correct depth information, which are the input learning data supplied from the learning image input unit 21.

[0056] If it is determined in step S30 that new depth information has not yet been generated for all pairs of sparse depth information and correct depth information supplied from the training image input unit 21, the process returns to step S21, and the processes of steps S21 to S30 described above are repeated. By the processes of steps S21 to S30, another pair of sparse depth information and correct depth information is selected, and a specified number of new depth information are further generated.

[0057] On the other hand, if it is determined in step S30 that new depth information has been generated for all pairs of sparse depth information and correct depth information supplied from the training image input unit 21, the processing proceeds to step S31, and the error assignment unit 23 outputs the newly generated training data (extended training data) and the training data (input training data) supplied from the training image input unit 21 to the training unit 24, thereby completing the data extension processing.

[0058] The data extension process by the error adding unit 23 is performed as described above.

[0059] In the error modeling process, when the error model is formulated using two error modeling functions, for example, a first error modeling function using a first parameter and a second error modeling function using a second parameter, the error assigning unit 23 sets the error obtained by integrating the first error by the first error modeling function using the first error modeling function and the second error by the second error modeling function using the second error modeling function as the pixel value of the pixel of interest in the new depth information. The error obtained by integrating the first error and the second error can be, for example, the average value or maximum value of the first error and the second error.

[0060] According to the information processing device 1 described above, augmentation (data expansion) can be performed using input training data supplied from the training image input unit 21 to generate expanded training data. The error modeling unit 22 of the information processing device 1 analyzes errors contained in the sparse depth information of the input training data supplied from the training image input unit 21 and generates an error modeling function that models the error distribution. The error assignment unit 23 of the information processing device 1 generates new depth information, which is new depth information obtained by efficiently assigning errors to correct depth information. Therefore, the information processing device 1 can efficiently generate training images to which errors have been assigned. The learning unit 24 of the information processing device 1 can generate a noise-robust learning model (deep neural network model) using the generated training images. The learning model of the learning unit 24, which inputs depth information in the form of disparity or depth values, can have a smaller network size than a learning model that inputs right RGB images and left RGB images.

[0061] 3. Computer Configuration Example The series of processes executed by the information processing device 1 can be executed by hardware or software. When the series of processes are executed by software, the programs that make up the software are installed in a computer. Here, the computer includes a microcomputer built into dedicated hardware, and a general-purpose personal computer, for example, that can execute various functions by installing various programs.

[0062] FIG. 8 is a block diagram showing an example of the hardware configuration of a computer serving as the information processing device 1. As shown in FIG.

[0063] The computer 100 includes a CPU (Central Processing Unit) 101, a ROM (Read Only Memory) 102, and a RAM (Random Access Memory) 103. The CPU 101, the ROM 102, and the RAM 103 are interconnected by a bus 104.

[0064] An input / output interface 105 is further connected to the bus 104. An input unit 106, an output unit 107, a storage unit 108, a communication unit 109, and a drive 110 are connected to the input / output interface 105.

[0065] The input unit 106 includes a keyboard, mouse, microphone, touch panel, input terminal, etc. The output unit 107 includes a display, speaker, output terminal, etc. The storage unit 108 includes a hard disk, SSD (Solid State Drive), RAM disk, non-volatile memory, etc. The communication unit 109 includes a network interface, etc. The drive 110 drives removable media 111 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.

[0066] In the computer 100 configured as above, the CPU 101 performs the above-described series of processes by, for example, loading a program stored in the storage unit 108 into the RAM 103 via the input / output interface 105 and the bus 104 and executing the program. The RAM 103 also stores data necessary for the CPU 101 to execute various processes as appropriate.

[0067] The program executed by the CPU 101 of the computer 100 can be provided by being recorded on a removable medium 111 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0068] In the computer 100, the program can be installed in the storage unit 108 via the input / output interface 105 by inserting the removable medium 111 into the drive 110. The program can also be received by the communication unit 109 via a wired or wireless transmission medium and installed in the storage unit 108. Alternatively, the program can be installed in the ROM 102 or the storage unit 108 in advance.

[0069] 4. Configuration Example of Second Embodiment FIG. 9 is a block diagram showing a configuration example of an information processing system according to a second embodiment of the present disclosure.

[0070] The information processing system shown in FIG. 9 is a system that uses a learning model learned by the information processing device 1 in FIG. 1, and is configured from the information processing device 1 and an imaging device 300.

[0071] The imaging device 300 is configured with a digital camera, an imaging sensor such as a CMOS image sensor or a CCD, an imaging module, etc. The imaging device 300 has an imaging unit 321, a signal processing unit 322, and a storage unit 323, and the signal processing unit 322 has a parallax calculation unit 341 and an inference unit 342.

[0072] As described above, the information processing device 1 learns a learning model that estimates dense depth information from sparse depth information, using the learning data supplied from the error adding unit 23. The information processing device 1 outputs parameters of the learning model as the learning result to the imaging device 300. The parameters of the learning model supplied from the information processing device 1 are stored in the storage unit 323 of the imaging device 300.

[0073] For example, if the imaging device 300 is a digital camera, the imaging unit 321 is configured as a stereo camera having a first imaging element that generates an image for the right eye (right-eye image) and a second imaging element that generates an image for the left eye (left-eye image). In this case, the imaging unit 321 generates a right-eye image and a left-eye image and outputs them to the parallax calculation unit 341. Also, for example, if the imaging device 300 is an image sensor, the imaging unit 321 is configured as a pixel array in which pixels having a dual pixel structure in which two photodiodes are aligned in the left-right direction are two-dimensionally arranged in a matrix. In this case, the imaging unit 321 outputs the signals of each of the two photodiodes in the pixel to the parallax calculation unit 341 as a phase difference signal.

[0074] For example, if the imaging device 300 is a digital camera, the parallax calculation unit 341 calculates a parallax or depth value and reliability for each pixel from the right-eye image and left-eye image supplied from the imaging unit 321. The parallax calculation unit 341 generates sparse depth information consisting only of the parallax or depth values ​​of pixels with high reliability, in other words, pixels with reliability equal to or greater than a predetermined value, from the calculated parallax or depth values ​​of each pixel, and outputs the sparse depth information to the inference unit 342. For example, if the imaging device 300 is an imaging sensor, the parallax calculation unit 341 calculates a parallax or depth value and reliability for each pixel based on the phase difference signal supplied from the imaging unit 321, generates sparse depth information consisting only of pixels with high reliability, and outputs the sparse depth information to the inference unit 342.

[0075] The inference unit 342 estimates and outputs dense depth information using a learning model that estimates dense depth information from the generated sparse depth information. Parameters of the learning model are acquired from the storage unit 323. The dense depth information may be output to an external device, the storage unit 323 within the imaging device 300, or a display (not shown). The signal processing unit 322 is configured with, for example, an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), a microprocessor, etc.

[0076] The storage unit 323 is configured, for example, with a solid state drive (SSD), a hard disk drive (HDD), or a non-volatile memory, and stores parameters of the learning model supplied from the information processing device 1. The information processing device 1 and the imaging device 300 may transmit and receive data directly via wired communication conforming to standards such as HDMI (registered trademark) (High-Definition Multimedia Interface) and USB (Universal Serial Bus), or via short-range wireless communication such as Bluetooth (registered trademark) or NFC (Near Field Communication), or may transmit and receive data via a network line such as the Internet or a local area network (LAN). In addition, the imaging device 300 may acquire, from the server device, parameters of the learning model uploaded to the server device by the information processing device 1.

[0077] The imaging device 300 configured as described above can generate and output high-resolution (high-density) and high-accuracy depth information using (parameters of) the learning model generated by the information processing device 1.

[0078] A configuration equivalent to the imaging device 300 can be incorporated into electronic devices with a distance measurement function, such as mobile terminals such as smartphones, tablets, and game consoles with a distance measurement function, head-mounted displays, etc. This allows electronic devices such as smartphones to generate and output high-resolution and high-precision depth information.

[0079] The embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the technology of the present disclosure.

[0080] For example, it is possible to adopt a form in which all or part of the above-described embodiments are combined as appropriate.

[0081] For example, the technology of the present disclosure can be configured as a cloud computing system in which a single function is shared and processed collaboratively by multiple devices via a network.

[0082] Furthermore, each step described in the above flowchart can be executed by one device or can be shared and executed by multiple devices. When one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0083] In this specification, the steps described in the flowcharts may be performed in chronological order in the order described, but they do not necessarily have to be processed in chronological order, and may be performed in parallel or at any necessary timing, such as when a call is made.

[0084] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0085] The effects described in this specification are merely examples and are not intended to be limiting, and there may be effects other than those described in this specification.

[0086] The technology disclosed herein may employ the following configurations. (1) An information processing device including: an error modeling unit that generates, from sparse depth information in which depth values ​​or disparities are stored in some pixels and correct depth information that is the correct value of the sparse depth information, an error modeling function that models an error distribution of the depth values ​​or disparities included in the sparse depth information; and an error assigning unit that generates a random error using the error modeling function as an error assigning function to generate new depth information including the random error. (2) The information processing device described in (1), in which the error modeling unit generates an error map from the sparse depth information and the correct depth information, and generates the error modeling function using the error map. (3) The information processing device described in (2), in which the error modeling unit uses the error map to calculate a standard deviation by assuming a predetermined error distribution for each predetermined parameter, and generates the error modeling function. (4) The information processing device described in (3), in which a setting value of an optical system or a setting value of a sensor is used as the predetermined parameter. (5) The information processing device according to any one of (3) to (4), wherein a focal length is used as the predetermined parameter. (6) The information processing device according to any one of (1) to (5), wherein the error modeling function is a function representing the relationship between a correct distance and a standard deviation of an error. (7) The information processing device according to any one of (1) to (6), wherein the error assigning unit generates the new depth information by assigning the random error to the correct depth information. (8) The information processing device according to any one of (1) to (7), wherein the error assigning unit calculates a value of the error assigning function by substituting pixel values ​​of the correct depth information into the error assigning function, and generates a random error whose average is the pixel values ​​of the correct depth information and whose standard deviation is the value of the error assigning function.(9) The information processing device according to any one of (1) to (8), wherein the error modeling unit generates the error modeling function for each predetermined parameter, and the error assigning unit generates a random error using the error modeling function of a parameter randomly selected from a range of parameters that the correct depth information can take as the error assigning function, and generates new depth information including the random error. (10) The information processing device according to any one of (1) to (9), further comprising a learning unit that uses the sparse depth information, the new depth information, and the correct depth information to train a learning model that estimates dense depth information from the sparse depth information. (11) An information processing method comprising: generating an error modeling function that models an error distribution of the depth values ​​or disparities included in the sparse depth information, from sparse depth information in which depth values ​​or disparities are stored in some pixels and correct depth information that is the correct value of the sparse depth information; and generating a random error using the error modeling function as an error assigning function, and generating new depth information including the random error.

[0087] 1 Information processing device, 21 Learning image input unit, 22 Error modeling unit, 23 Error assignment unit, 24 Learning unit, 100 Computer, 101 CPU, 102 ROM, 103 RAM, 106 Input unit, 107 Output unit, 108 Storage unit, 109 Communication unit, 110 Drive, 111 Removable media, 300 Imaging device, 321 Imaging unit, 322 Signal processing unit, 323 Storage unit, 341 Parallax calculation unit, 342 Inference unit

Claims

1. An information processing device comprising: an error modeling unit that generates an error modeling function that models the error distribution of the depth values ​​or disparities contained in the sparse depth information from sparse depth information in which depth values ​​or disparities are stored in some pixels and correct depth information that is the correct value of the sparse depth information; and an error assignment unit that generates a random error using the error modeling function as an error assignment function and generates new depth information containing the random error.

2. The information processing device according to claim 1, wherein the error modeling unit generates an error map from the sparse depth information and the ground truth depth information, and generates the error modeling function using the error map.

3. The information processing device according to claim 2, wherein the error modeling unit uses the error map to calculate a standard deviation by assuming a predetermined error distribution for each predetermined parameter, and generates the error modeling function.

4. The information processing device according to claim 3, wherein the predetermined parameters are set values ​​of an optical system or set values ​​of a sensor.

5. The information processing device according to claim 3, wherein the predetermined parameter is a focal length.

6. The information processing device according to claim 1, wherein the error modeling function is a function that represents the relationship between the correct distance and the standard deviation of the error.

7. The information processing device according to claim 1, wherein the error adding unit generates the new depth information by adding the random error to the correct depth information.

8. The information processing device described in claim 1, wherein the error assignment unit assigns pixel values ​​of the correct depth information to the error assignment function to calculate the value of the error assignment function, and generates a random error with the pixel values ​​of the correct depth information as the average and the value of the error assignment function as the standard deviation.

9. The information processing device according to claim 1, wherein the error modeling unit generates the error modeling function for each predetermined parameter, and the error assignment unit generates a random error using the error modeling function of a parameter randomly selected within the range of parameters that the correct depth information can take as the error assignment function, and generates new depth information including the random error.

10. The information processing device according to claim 1, further comprising a learning unit that uses the sparse depth information, the new depth information, and the correct depth information to learn a learning model that estimates dense depth information from the sparse depth information.

11. An information processing method comprising: generating an error modeling function that models the error distribution of the depth values ​​or disparities contained in sparse depth information from sparse depth information in which depth values ​​or disparities are stored in some pixels and correct depth information that is the correct value of the sparse depth information; generating random errors using the error modeling function as an error assignment function, and generating new depth information containing the random errors.

Citation Information

Patent Citations

  • Gesture estimation method based on parallel convolution neural network

    CN107423698A

  • Depth sensor noise

    JP2018109976A

  • Depth super resolution device, depth super resolution method, and program

    JP2020123114A

  • Information processing device, information processing method, and program

    JP2022074731A