Information processing device, imaging device, method, and program

By setting up a 3D model of the camera and theme in the virtual three-dimensional space, and generating a depth map and a focal length map, the problem of target switching in multi-person tracking is solved, and learning data containing depth information is achieved efficiently, and learning ability of NN in three-dimensional context is improved.

JP2025073076APending Publication Date: 2025-05-12CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024175265
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-25
Filing Date
2024-10-04
Publication Date
2025-05-12

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of target switching when using in-depth information for multi-person tracking, especially when two people with similar characteristics intersect, NN cannot learn a three-dimensional context to eliminate switching.

Method used

By setting up a 3D model of the camera and the subject in a virtual three-dimensional space, a depth map and a focal length map containing depth information within the camera's field of view are generated, and these images are used to generate a focal length map containing focal length information in part of the area.

Benefits of technology

It realizes efficient acquisition of learning data containing deep information in multi-person tracking scenarios, effectively reducing the occurrence of target switching and improving the learning ability of NN in three-dimensional contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025073076000001_ABST
    Figure 2025073076000001_ABST
Patent Text Reader

Abstract

To provide a technique for efficiently acquiring learning data having depth information.SOLUTION: An information processing device includes: setting means configured to arrange a three-dimensional model of a subject and a camera in a virtual three-dimensional space; first generation means configured to generate a depth map including at least a depth value of a partial region of a region around the subject based on an image in which the subject rendered based on an imaging field of view of the camera is captured and distance information corresponding to the imaging field of view; and second generation means configured to generate a defocus map including a defocus amount of the partial region based on the depth value of the partial region and an imaging parameter of the camera.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing device, an imaging device, a method, and a program. [Background technology]

[0002] Cameras equipped with a focus adjustment function that automatically adjusts the focal position of the photographic lens are becoming widespread. Phase-difference AF and contrast AF methods have been put to practical use as camera focus adjustment methods (AF methods). Phase-difference AF has the advantage of being able to focus more quickly than contrast AF, because it can directly calculate the amount of focus misalignment from two images with parallax.

[0003] In recent years, a method has been proposed to detect object regions in an image using a neural network (NN). A major challenge in this field is to detect objects while tracking the subject with high accuracy in captured images acquired in real time. The parameters input to the NN for image recognition are RGB color images, but by inputting depth information to the NN in addition to the color image, object recognition that takes into account three-dimensional context can be realized. Here, the amount of shift in the focal plane mentioned above can become depth information.

[0004] To improve the generalization performance of a NN, a large amount of training data is required. Data augmentation (DA) is used as a method to improve the generalization performance of a NN even with a small amount of training data. DA is a method to artificially expand the training data (e.g., images) by processing them with blurring, shaking, image synthesis, rotation, translation, enlargement / reduction, up / down / left / right inversion, noise addition, color tone change, brightness change, etc.

[0005] Patent Document 1 proposes a method of increasing learning data by using three-dimensional computer graphics (hereinafter referred to as 3DCG) to change drawing parameters such as lighting for a three-dimensional recognition target model and using the rendered image as training data. Patent Document 2 proposes a method of increasing learning data by superimposing a first image, which is a rendered 3D human model, on a background image, missing pixels along the contours of the human body, and adding noise. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] JP 2018-163554 A [Patent Document 2] Patent Publication No. 2021-43839 Summary of the Invention [Problem to be solved by the invention]

[0007] For example, when two people with similar features cross while tracking multiple people, the problem of switching between tracking targets occurs. However, since Patent Document 1 does not consider the depth information of the superimposed object, it is not possible to make the NN learn a three-dimensional context to eliminate switching. Patent Document 2 causes pixel loss at the edge of the superimposed depth image or the contour of the human body part in the depth image. Therefore, even if the learning data is increased based on images acquired by a camera using a phase difference AF method that cannot obtain dense depth information, it is not effective in eliminating switching.

[0008] Therefore, an object of the present invention is to provide a technique for efficiently acquiring learning data having depth information. [Means for solving the problem]

[0009] In order to achieve the object of the present invention, an information processing device according to one embodiment of the present invention comprises: a setting means for arranging a three-dimensional model of a subject and a camera in a virtual three-dimensional space; a first generation means for generating a depth map including depth values ​​of at least a partial region around the subject based on an image of the subject rendered based on the field of view of the camera and distance information corresponding to the field of view; and a second generation means for generating a defocus map including a defocus amount of the partial region based on the depth value of the partial region and the shooting parameters of the camera. Effect of the Invention

[0010] According to the present invention, it is possible to provide a technique for efficiently acquiring learning data having depth information. [Brief description of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram showing a hardware configuration of an information processing apparatus according to a first embodiment. [Diagram 2] FIG. 1 is a diagram showing the functional configuration of an information processing apparatus according to a first embodiment. [Diagram 3] 1A to 1C are diagrams illustrating the arrangement of 3D models in a three-dimensional space according to the first embodiment. [Figure 4] 5 is a flowchart illustrating a process of generating a defocus map according to the first embodiment. [Diagram 5] 4A to 4C are views for explaining distance information and a depth value calculation region according to the first embodiment. [Figure 6] FIG. 11 is a diagram showing the functional configuration of an information processing device according to a second embodiment. [Figure 7] 13A to 13C are diagrams illustrating the arrangement of 3D models in a three-dimensional space according to the second embodiment. [Figure 8(a)] 13A to 13C are views for explaining alignment of defocus amount calculation regions according to the second embodiment. [Figure 8(b)] 13A to 13C are views for explaining alignment of defocus amount calculation regions according to the second embodiment. [Figure 8(c)] 13A to 13C are views for explaining alignment of defocus amount calculation regions according to the second embodiment. [Figure 9] 10 is a flowchart illustrating a process of generating a defocus map according to a second embodiment. [Figure 10] FIG. 2 is a diagram showing the functional configurations of a learning device and an imaging device. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.

[0013] First Embodiment FIG. 1 is a diagram showing a hardware configuration of an information processing device according to the first embodiment.

[0014] The information processing device 10 is a device that generates learning data for a NN that performs defocus inference. The information processing device 10 includes a CPU 100, a ROM 110, a RAM 120, a HDD 130, an input unit 140, a display unit 150, and a communication unit 160. The information processing device 10 is, for example, a general-purpose PC.

[0015] A CPU (Central Processing Unit) 100 is a central processing unit, and performs calculations and logical determinations for various processes.

[0016] A ROM (Read-Only-Memory) 110 stores the control program executed by the CPU 100 .

[0017] A RAM (Random Access Memory) 120 is the main memory of the CPU 100, and provides a temporary storage area such as a work area.

[0018] The HDD (Hard Disk Drive) 130 is a hard disk that stores data and programs according to this embodiment. An external storage device (not shown) may be used to fulfill the same role as the HDD 130. Here, the external storage device includes, for example, a medium (recording medium) and an external storage drive for realizing access to the medium. Examples of the medium include a flexible disk (FD), a CD-ROM, a DVD, a USB memory, an MO, and a flash memory. The external storage device may also be a server device connected via a network.

[0019] The input unit 140 is configured with a keyboard, a touch panel, and the like, and is a device for accepting input from a user.

[0020] The display unit 150 is configured with a liquid crystal display or the like, and can display various data and processing results to the user. The display unit 150 can also communicate with another device (not shown) via the communication unit 160. The other device may receive instructions from the user via the communication unit 160, or may output processing results to the display unit 150. The other device is, for example, a PC, a smartphone, or a tablet terminal.

[0021] FIG. 2 is a diagram showing the functional configuration of the information processing device according to the first embodiment.

[0022] The information processing device 10 includes a modeling unit 201, a camera setting unit 203, a distance information acquisition unit 204, a depth map generation unit 205, a defocus map generation unit 206, and a rendering unit 207. The 3D model DB 202 has three-dimensional models (also called 3D models) of people and objects. The learning data 208 includes the output (defocus map) of the defocus map generation unit 206 and the output (rendered image) of the rendering unit 207.

[0023] The modeling unit 201 can set three-dimensional models of a camera 301, a subject (e.g., people 302a to 302c), and a background (object 303, object 304) in a virtual three-dimensional space 300 in Fig. 3 described below. The modeling unit 201 can select three-dimensional models of the subject and background from a 3D model DB 202. Furthermore, the modeling unit 201 can determine external shooting parameters, which are settings of the shooting direction of the camera 301.

[0024] FIG. 3 is a diagram illustrating the arrangement of 3D models in a three-dimensional space according to the first embodiment.

[0025] FIG. 3 shows the arrangement of a camera 301, a person 302a, a person 302b, a person 302c, an object 303, and an object 304 in a three-dimensional space 300. The camera 301 is a 3D model of a camera. The people 302a to 302c are 3D models of people. By replacing the texture, the people 302a to 302c can be arranged as different people. The object 303 (illustrated as a cylinder) and the object 304 (illustrated as a cube) are 3D models of objects. At least one of the object 303 and the object 304 can be arranged as an obstacle between the camera 301 and the person 302. In addition, by arranging the object 303 behind the person 302a (upward in FIG. 3), the background of the person 302a can be produced.

[0026] The camera setting unit 203 sets internal shooting parameters of the camera 301 arranged in the three-dimensional space 300. Here, the internal shooting parameters include settings such as the sensor size, the focal length of the lens, the focus position, the aperture value, the shutter speed, and the ISO sensitivity.

[0027] Based on the field of view (FOV) of the camera 301, the distance information acquisition unit 204 calculates the distance from the camera 301 to the subjects (persons 302a to 302c) and the distance from the camera 301 to the background (objects 303 and 304) over the entire field of view.

[0028] The depth map generating unit 205 determines a depth value calculation region in the rendered image based on the field of view of the camera 301. The depth value calculation region is, for example, a region including all of the 12×16 divided cells (partial regions). The depth map generating unit 205 calculates a depth value for each cell (partial region) based on the distance information acquired by the distance information acquiring unit 204. The depth value is a value obtained by aggregating the depth information in the cell (partial region) into one value, and in this embodiment, is an average value of the distance information in the cell (partial region).

[0029] The defocus map generating unit 206 calculates the amount of defocus, which is an index of the amount of focus deviation, for each cell from the depth value for each cell calculated by the depth map generating unit 205, and generates a defocus map.

[0030] The rendering unit 207 stores learning data 208 that associates an image rendered based on the internal and external shooting parameters of the camera 301 with a defocus map generated by the defocus map generation unit 206. Although it has been described that the rendering unit 207 performs the storage process of the learning data 208, the defocus map generation unit 206 may perform the same storage process as the rendering unit 207. Note that the rendering unit 207 may appropriately assign annotation information acquired based on CG (computer graphics) information that reproduces a 3D model to the learning data 208 according to the machine learning task. For example, when the machine learning task is an object detection task, the rendering unit 207 assigns annotation information of the type (e.g., person), coordinates, and size of the subject for each subject of the rendered image.

[0031] 4 is a flowchart illustrating the process of generating a defocus map according to the first embodiment. The process of generating a defocus map is realized by the CPU 100 executing a program in the ROM 110.

[0032] In S401, the distance information acquisition unit 204 acquires distance information from the camera 301 to the subjects (persons 302a to 302c) and distance information from the camera 301 to the background (objects 303 and 304) based on the field of view (FOV) of the camera 301. For example, the distance information acquisition unit 204 can acquire each of the above distance information based on CG (computer graphics) information that reproduces a 3D model. Here, the distance information is an image in which the distances to the subjects and background in the entire field of view of the camera 301 are quantified and visualized, as shown in FIG. 5(a).

[0033] 5A and 5B are diagrams illustrating distance information and a depth value calculation region according to the first embodiment.

[0034] Distance information 500 indicates the distance to person 302a, person 302b, and object 303 with shades of color. The darker the color in distance information 500, the closer the subject is to camera 301. On the other hand, the lighter the color in distance information 500, the farther the subject is from camera 301. Now, return to the description of FIG. 4.

[0035] In S402, the depth map generation unit 205 determines a depth value calculation region 510 in distance information 500 (eg, an image).

[0036] 5B is a diagram for explaining a depth value calculation region on the distance information. Note that the depth value calculation region 510 is used as a defocus amount calculation region, which will be described later.

[0037] Depth value calculation area 510 on distance information 500 in FIG. 5(b) is an area divided into 12×16 cells. Depth value calculation area 510 also refers to an area around the main subject. Depth map generation unit 205 sets depth value calculation area 510 so that one cell (partial area) is large enough to sample the depth value of the main subject (for example, the size of one cell (partial area) is half the size of the face of the main subject). Here, we return to the explanation of FIG. 4.

[0038] In S403 , the defocus map generation unit 206 calculates the average depth value of each cell in the depth value calculation region 510 .

[0039] In S404, the defocus map generating unit 206 calculates the defocus amount of each cell by subtracting the distance to the focus position of the camera 301 (that is, the distance to the focal plane of the camera) from the average depth value of each cell.

[0040] In S405, the defocus map generating unit 206 generates a defocus map by dividing the defocus amount of each cell acquired in S404 by the aperture value, which is an internal shooting parameter of the camera 301. The phase difference AF method can directly calculate the amount of deviation of the focal plane from two images with parallax. However, when the aperture value of the camera 301 is large, the base line length is short, so that sufficient parallax is not obtained and the defocus amount is relatively small. As a result, the image rendered based on the shooting field of the camera 301 becomes an image that is in focus overall. Therefore, when the aperture value of the internal shooting parameter of the camera 301 is set to a large value, a gain is applied to the defocus amount calculated in S405 so that the defocus amount becomes small, thereby simulating a defocus amount that is more in line with the actual shooting environment. Then, the defocus map generating unit 206 saves the learning data 208 that associates the defocus map generated in S405 with the image rendered by the rendering unit 207, and ends the process.

[0041] In S405, the defocus map generating unit 206 calculates the defocus amount by dividing the defocus amount acquired in S404 by the internal shooting parameters of the camera 301 (specifically, the aperture value), but this is not limited to this. For example, the defocus map generating unit 206 may acquire in advance a table that specifies the relationship between the aperture value and the depth value for a specific lens, and determine the final defocus amount based on the table. This makes it possible to more faithfully reproduce the defocus amount of the actual camera lens.

[0042] Also, although it has been described in S403 that the average value is used as the representative value of the depth of each cell in the depth value calculation area 510, the most frequent value may be used. In a cell in which the distribution of distance information is multi-peaked, due to the characteristics of the correlation calculation of the phase difference AF method, there are multiple defocus amounts with high correlation. However, the average of these defocus amounts results in a defocus amount that does not focus on any of the subjects. By using the most frequent value, the defocus map generation unit 206 can determine one defocus amount from among the defocus amounts with high correlation.

[0043] As described above, according to the first embodiment, a defocus map that takes into account the actual shooting environment of the camera can be obtained by using a 3D model of the camera and the subject arranged in a virtual 3D space. This makes it possible to efficiently obtain learning data in which an image of an arbitrary subject is associated with a defocus map.

[0044] <Image capture device that infers distance> An imaging device incorporating a trained model trained based on the training data generated in the first embodiment will be described.

[0045] FIG. 10 shows an example of the configuration of a learning device 1001 and an image capture device 1008 that are trained using the training data 208.

[0046] In the learning device 1001 , the learning data acquisition unit 1002 receives the learning data 208 .

[0047] The distance to an object in an image is inferred using an inference unit 1003. The inference target is an example and may vary depending on the application.

[0048] A loss calculation unit 1004 calculates a loss by comparing the inference result output from the inference unit 1003 with the correct answer value acquired by the learning data acquisition unit 1002. As a loss function, an L1 loss that is generally used in regression tasks is used.

[0049] The weight update unit 1005 updates the weights of the network used in machine learning from the loss calculated by the loss calculation unit 1004. After that, the weights are output as a trained model 1007, and at the same time, the weight information is stored in the parameter storage unit 1006, and the weights are used by the inference unit 1003 at the next training. At this time, the output destination is not limited to a specific format, and may be a memory of a general-purpose computer or a control circuit inside a camera. In the description of this embodiment, the output is assumed to be output to a storage unit 1009 of an imaging device 1008 capable of acquiring a defocus map. The storage unit 1009 is a recording medium such as a memory card.

[0050] The imaging device 1008 reads out the trained model 1007 stored in the memory unit 1009 by the model reading unit 1010.

[0051] Inference unit 1013 inputs image 1011 and defocus map 1012 to trained model 1007 to obtain distance inference results to objects in the image.

[0052] The shooting parameter inference unit 1003 may be implemented using various models, such as a neural network like a covolutional neural network (CNN), a vision transformer (ViT), or a support vector machine (SVM) combined with a feature extractor.

[0053] <Second embodiment> In the second embodiment, a defocus map is generated by arranging three-dimensional models of three cameras in a virtual three-dimensional space so that parallax information can be obtained. Since the defocus map can be generated by a method similar to the imaging surface phase difference AF, learning data close to a defocus map acquired by a camera in real space can be obtained. Hereinafter, a method of generating learning data (defocus map) will be described with reference to Figs. 6 to 9. Note that in the second embodiment, differences from the first embodiment will be described.

[0054] Fig. 6 is a diagram showing the functional configuration of an information processing device according to the second embodiment. In Fig. 6, the same functional blocks as in Fig. 2 are denoted by the same reference numerals as in Fig. 2.

[0055] The information processing device 600 includes a modeling unit 201, a camera setting unit 203, a defocus map generating unit 206, and a rendering unit 207. Compared to the first embodiment, the information processing device 600 does not include the distance information acquiring unit 204 and the depth map generating unit 205, but is not limited to this.

[0056] FIG. 7 is a diagram illustrating the arrangement of 3D models in a three-dimensional space according to the second embodiment.

[0057] As shown in FIG. 7, the rendering unit 207 renders three images based on three types of shooting fields arranged at equal intervals in the base line direction between the cameras 3012 and 3013. The first shooting field is a shooting field based on an arbitrary shooting direction of the camera 3011. The camera 3011 is disposed at the same position as the camera 301 in FIG. 3, for example. The second shooting field is a shooting field of the camera 3012 disposed at a distance of −b / 2 in the longitudinal direction (base line direction) of the first shooting field. The third shooting field is a shooting field of the camera 3013 disposed at a distance of +b / 2 in the longitudinal direction (base line direction) of the first shooting field. The internal shooting parameters of the cameras 3011 to 3013 are set to be the same. The external shooting parameters of the cameras 3011 to 3013 are set to be the same except for the camera coordinates for creating parallax. As a result, parallax information based on the condition of the base line length b (not shown) between the camera 3012 and the camera 3013 can be obtained.

[0058] 9 is a flowchart illustrating the process of generating a defocus map according to the second embodiment. The process of generating a defocus map is realized by the CPU 100 executing a program in the ROM 110.

[0059] In S901, the rendering unit 207 renders a "left-eye image" based on the shooting parameters of the camera 3012. The rendering unit 207 renders a "right-eye image" based on the shooting parameters of the camera 3013. The rendering unit 207 renders a "both-eye image" that is an image of an intermediate shooting field of view between the left-eye image and the right-eye image based on the shooting parameters of the camera 3011.

[0060] In S902, the defocus map generation unit 206 determines a defocus amount calculation region (not shown) in both eye images, and divides the defocus amount calculation region into cells of, for example, 12 × 16. Note that this defocus amount calculation region (first region) is the same as the depth value calculation region in FIG. 5(b).

[0061] In S903, the defocus map generating unit 206 calculates an area (the second area in the left eye image and the third area in the right eye image) corresponding to the defocus amount calculation area calculated in S902 in each of the left eye image and the right eye image, and calculates a defocus amount for each corresponding cell. Hereinafter, a method of aligning the defocus amount calculation area calculated in S902 for the left eye image and the right eye image will be described with reference to FIG.

[0062] With reference to Fig. 8(a) to Fig. 8(c), alignment of the defocus amount calculation region according to the second embodiment will be described. Fig. 8(a) is a diagram showing cameras 3011 to 3013 viewed from the short side direction of the sensor. Here, θ represents the horizontal angle of view of the sensor.

[0063] The distance from the sensor center of cameras 3011 to 3013 to the focal plane is assumed to be Zo. In this case, the photographing field of view in the horizontal (long axis) direction of the sensor on the focal plane is expressed as 2Zotan(θ / 2). This photographing field of view is the range recorded in the horizontal pixels of the image. Therefore, the amount of deviation g between the optical center of camera 3011 and the optical center of camera 3012 for the left eye image is calculated by formula (1) with the horizontal resolution of the camera being H. b / 2:2Zotan(θ / 2)=g:H (1)

[0064] That is, in the left eye image of FIG. 8(b), if the defocus amount calculation region having the same angle of view as the both-eye image of FIG. 8(c) is shifted to the right by g, the region of the left eye image corresponding to the defocus amount calculation region of the both-eye image can be obtained. Similarly, in the right eye image (not shown), if the defocus amount calculation region having the same angle of view as the both-eye image is shifted to the left by g, the region of the right eye image corresponding to the defocus amount calculation region of the both-eye image can be obtained. Hereinafter, the calculation method of the defocus amount will be described. Returning to the description of FIG. 9.

[0065] In S904, the defocus map generating unit 206 calculates the amount of image shift (image shift amount) in corresponding cells of both eye images by performing correlation calculation processing between corresponding cells of the left-eye image and the right-eye image.

[0066] Specifically, corresponding cells in the left-eye image and right-eye image are averaged in the column direction to obtain left-eye and right-eye signals. Corresponding cells are shifted relative to each other in the row direction to calculate the correlation amount COR(s), which indicates the degree of signal matching.

[0067] The left-eye signal of the kth column in a certain cell is denoted by A(k), the right-eye signal is denoted by B(k), and the range of k corresponding to the cell is denoted by W. The shift amount by the shift process is denoted by s, and the shift range of the shift amount s is denoted by Γ. In this case, the correlation amount COR(s) is calculated by Equation (2). TIFF2025073076000002.tif15106...(2)

[0068] By shifting by the shift amount s, the left eye signal A(k) of the kth column and the right eye signal B(ks) of the ksth column are matched and subtracted to generate a shifted subtraction signal. The absolute value of the generated shifted subtraction signal is calculated and the sum is taken within the range W corresponding to the cell area to calculate the correlation amount COR(s). Since the correlation amount COR(s) is found for each column, the real-valued shift amount s that minimizes the correlation amount COR(s) is calculated using three-point interpolation or the like to determine the image shift amount.

[0069] In S905, the defocus map generation unit 206 calculates the defocus amount d by multiplying the image shift amount by a conversion coefficient K. Here, the conversion coefficient K is a parameter that changes according to the base length b between the positions of the cameras 3012 and 3013. The conversion coefficient K corresponds to a conversion coefficient that converts the image shift amount in pixel units to a defocus amount in the image plane phase difference AF.

[0070] The defocus map generation unit 206 performs the same calculation as described above for all 12×16 cells, and finally obtains a defocus map of the images for both eyes.

[0071] The defocus map generation unit 206 saves the learning data 208 that associates the defocus map of the both-eye images generated in S905 with the both-eye images rendered by the rendering unit 207, and ends the process.

[0072] As described above, according to the second embodiment, it is possible to calculate a defocus map based on parallax information between a left-eye image and a right-eye image. Also, it is possible to efficiently acquire learning data in which an image of an arbitrary subject is associated with a defocus map that simulates the sensor characteristics and optical characteristics of the image plane phase difference AF.

[0073] <Other embodiments> The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0074] The disclosure of this specification includes the following information processing device, imaging device, method, and program. (Item 1) A setting means for arranging a three-dimensional model of a subject and a camera in a virtual three-dimensional space; a first generating means for generating a depth map including depth values ​​of at least a partial area of ​​an area around the subject based on an image including the subject that is rendered based on a field of view of the camera and distance information corresponding to the field of view; and a second generating means for generating a defocus map including a defocus amount of the partial region based on a depth value of the partial region and an imaging parameter of the camera. Information processing device. (Item 2) The defocus amount of the partial region is a value obtained by subtracting the distance to the focal plane of the camera from the depth value of the partial region. Item 1. An information processing device according to item 1. (Item 3) The second generation means adjusts the subtracted value in accordance with the magnitude of the shooting parameter of the camera. Item 3. An information processing device according to item 2. (Item 4) The second generation means determines a defocus amount of the partial region based on information associating an aperture value corresponding to a predetermined lens of the camera with a depth value of the partial region. 4. The information processing device according to any one of items 1 to 3. (Item 5) The depth value of the subregion is an average value. 5. The information processing device according to any one of items 1 to 4. (Item 6) The depth value of the subregion is a mode value. 6. The information processing device according to any one of items 1 to 5. (Item 7) The partial region has a size that covers at least a part of the face of the subject. 7. The information processing device according to any one of items 1 to 6. (Item 8) The shooting parameter of the camera is an aperture value. 8. The information processing device according to any one of items 1 to 7. (Item 9) A storage means for storing the image and the defocus map in association with each other. 9. The information processing device according to any one of items 1 to 8. (Item 10) Obtaining an inference result for a subject in a captured image based on a trained model trained using the defocus map generated by the second generation means of the information processing device according to any one of items 1 to 9, 1. An imaging device comprising: (Item 11) A method executed by an information processing device, comprising: A setting process in which a 3D model of the subject and the camera are placed in a virtual 3D space; A first generation step of generating a depth map including depth values ​​of at least a partial area of ​​an area around the subject based on an image including the subject rendered based on a field of view of the camera and distance information corresponding to the field of view; and a second generation step of generating a defocus map including a defocus amount of the partial region based on a depth value of the partial region and an imaging parameter of the camera. method. (Item 12) A program for causing a computer to function as each of the means of the information processing device according to any one of items 1 to 9. (Item 13) a setting means for arranging a three-dimensional model of a subject in a virtual three-dimensional space, and arranging the three-dimensional models of the first camera, the second camera, and the third camera at intervals so that the optical axes of the first camera, the second camera, and the third camera are parallel to each other; a determining means for determining at least a first region around the object in a binocular image of the object rendered based on a first field of view of the first camera; a generating means for generating a defocus map including a defocus amount of a partial area of ​​the first area based on parallax information between a partial area of ​​a second area corresponding to the first area in a left-eye image in which the subject is depicted and rendered based on a second shooting field of the second camera, and a partial area of ​​a third area corresponding to the first area in a right-eye image in which the subject is depicted and rendered based on a third shooting field of the third camera, Information processing device. (Item 14) The setting means disposes the first camera at a midpoint of a base line length between the position of the second camera and the position of the third camera. Item 14. An information processing device according to item 13. (Item 15) The generating means determines a defocus amount of a partial region of the first region based on the length of the base line. Item 15. An information processing device according to item 14. (Item 16) A method executed by an information processing device, comprising: a setting step of arranging a three-dimensional model of a subject in a virtual three-dimensional space, and arranging the three-dimensional models of the first camera, the second camera, and the third camera at intervals so that the optical axes of the first camera, the second camera, and the third camera are parallel to each other; determining a first region around at least the object in a binocular image of the object rendered based on a first field of view of the first camera; and generating a defocus map including a defocus amount of a partial region of the first region based on parallax information between a partial region of a second region corresponding to the first region in a left-eye image in which the subject is depicted and rendered based on a second shooting field of the second camera, and a partial region of a third region corresponding to the first region in a right-eye image in which the subject is depicted and rendered based on a third shooting field of the third camera. method. (Item 17) A program for causing a computer to function as each of the means of the information processing device according to any one of items 13 to 15.

[0075] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0076] 10. Information processing device 100 CPU 110 ROM 120 RAM 130 HDD 140 Input section 150 Display section 160 Communications Department

Claims

1. A setting means for arranging a three-dimensional model of a subject and a camera in a virtual three-dimensional space; a first generating means for generating a depth map including depth values ​​of at least a partial area of ​​an area around the subject based on an image including the subject that is rendered based on a field of view of the camera and distance information corresponding to the field of view; and a second generating means for generating a defocus map including a defocus amount of the partial region based on a depth value of the partial region and an imaging parameter of the camera. Information processing device.

2. The defocus amount of the partial region is a value obtained by subtracting the distance to the focal plane of the camera from the depth value of the partial region. The information processing device according to claim 1 .

3. The second generation means adjusts the subtracted value in accordance with the magnitude of the shooting parameter of the camera. The information processing device according to claim 2 .

4. the second generation means determines a defocus amount of the partial region based on information associating an aperture value corresponding to a predetermined lens of the camera with a depth value of the partial region; The information processing device according to claim 1 .

5. The depth value of the subregion is an average value. The information processing device according to claim 1 .

6. The depth value of the subregion is a mode value. The information processing device according to claim 1 .

7. The partial region has a size that covers at least a part of the face of the subject. The information processing device according to claim 1 .

8. The shooting parameter of the camera is an aperture value. The information processing device according to claim 1 .

9. A storage means for storing the image and the defocus map in association with each other. The information processing device according to claim 1 .

10. obtaining an inference result for a subject in a captured image based on a trained model trained using the defocus map generated by the second generation means of the information processing device according to claim 1; 1. An imaging device comprising:

11. A method executed by an information processing device, comprising: A setting process of placing a three-dimensional model of a subject and a camera in a virtual three-dimensional space; a first generation step of generating a depth map including depth values ​​of at least a partial area of ​​an area around the object based on an image including the object rendered based on a field of view of the camera and distance information corresponding to the field of view; and a second generation step of generating a defocus map including a defocus amount of the partial region based on a depth value of the partial region and an imaging parameter of the camera. method.

12. A program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 9.

13. a setting means for arranging a three-dimensional model of a subject in a virtual three-dimensional space, and arranging the three-dimensional models of the first camera, the second camera, and the third camera at intervals so that the optical axes of the first camera, the second camera, and the third camera are parallel to each other; a determining means for determining at least a first region around the object in a binocular image of the object rendered based on a first field of view of the first camera; a generating means for generating a defocus map including a defocus amount of a partial area of ​​the first area based on parallax information between a partial area of ​​a second area corresponding to the first area in a left-eye image in which the subject is depicted and rendered based on a second shooting field of the second camera, and a partial area of ​​a third area corresponding to the first area in a right-eye image in which the subject is depicted and rendered based on a third shooting field of the third camera. Information processing device.

14. the setting means disposes the first camera at a midpoint of a base line length between the position of the second camera and the position of the third camera; The information processing device according to claim 13.

15. The generating means determines a defocus amount of a partial region of the first region based on the length of the base line. The information processing device according to claim 14.

16. A method executed by an information processing device, comprising: a setting step of arranging a three-dimensional model of a subject in a virtual three-dimensional space, and arranging the three-dimensional models of the first camera, the second camera, and the third camera at intervals so that the optical axes of the first camera, the second camera, and the third camera are parallel to each other; determining a first region around at least the object in a binocular image of the object rendered based on a first field of view of the first camera; and generating a defocus map including a defocus amount of a partial region of the first region based on parallax information between a partial region of a second region corresponding to the first region in a left-eye image in which the subject is depicted and rendered based on a second shooting field of the second camera, and a partial region of a third region corresponding to the first region in a right-eye image in which the subject is depicted and rendered based on a third shooting field of the third camera. method.

17. A program for causing a computer to function as each of the means of the information processing device according to any one of claims 13 to 15.

Citation Information

Patent Citations

  • Image processing device, image processing method, image processing program, and teacher data generating method

    JP2018163554A

  • Learning system, analysis system, method for learning, method for analysis, program, and storage medium

    JP2021043839A