Information processing apparatus, information processing system, method for controlling information processing apparatus, and program

The system addresses bandwidth issues in MR and AR by generating depth meta information to reduce the amount of data transmitted, effectively managing occlusion and reducing bandwidth strain.

JP2025175920APending Publication Date: 2025-12-03CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024144262
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-20
Filing Date
2024-08-26
Publication Date
2025-12-03

AI Technical Summary

Technical Problem

The transmission of depth maps in mixed reality (MR) and augmented reality (AR) systems strains communication bandwidth due to the high amount of data required for representing occlusion between real and virtual objects.

Method used

An information processing system that reduces the amount of depth information transmitted by generating depth meta information to represent only the depth range necessary for occlusion, using a client terminal to estimate depth maps and a server terminal to generate simplified depth maps based on this meta information.

Benefits of technology

This approach effectively suppresses the communication bandwidth shortage by reducing the amount of data transmitted, ensuring seamless integration of virtual objects into real space while maintaining occlusion representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025175920000001_ABST
    Figure 2025175920000001_ABST
Patent Text Reader

Abstract

To prevent shortage of bands for communication when achieving occlusion of one object to the other object in MR or AR.SOLUTION: An information processing apparatus that is a first apparatus has: first acquisition means that acquires a first image; first map acquisition means that acquires a first depth map indicating depth information corresponding to the first image; reduction means that reduces the amount of information on the first depth map on the basis of meta information so as to represent the range of depth corresponding to the meta information and thereby generates a second depth map; and first transmission means that transmits the first image and the second depth map to a second apparatus.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing system, a control method for an information processing device, and a program. [Background technology]

[0002] There is a technology that synthesizes virtual objects expressed by computer graphics (CG) into real space and presents them to the user. This technology is called mixed reality (MR) and augmented reality (AR).

[0003] MR and AR are technologies that combine virtual objects into real space. For this reason, it is best for users to use highly portable display devices to view various scenes in real space. However, highly portable devices have limitations in terms of computing resources and power supply. As a result, methods have emerged in which the rendering of 3D models, which is a heavy load, is processed by a server terminal on the cloud, and the resulting virtual image is sent to the client terminal on the display side (such as cloud rendering). Considering the portability of the client terminal, it is preferable for the communication between the client terminal to be wireless.

[0004] However, wireless communication is subject to bandwidth limitations. Mixed reality and augmented reality (AR) also require the transmission and reception of depth maps that represent the depth information of virtual images. For example, when a real object exists in front of a virtual object, it is necessary to represent occlusion, where the real object obscures the virtual object and makes it invisible. To represent occlusion, depth information for both the real object and the virtual object is required so that the depth of the real object can be compared. However, transmitting depth maps generally significantly reduces bandwidth.

[0005] In Patent Document 1, for an area where a real object is in front of a virtual object and occlusion occurs (an occluded area), a client terminal generates a mask or a depth map of real space and transmits it to a server terminal. The server terminal generates a CG image that omits the rendering of the occluded area and transmits the CG image to the client terminal. The client terminal then combines the real space with the CG image. This eliminates the need to transmit a depth map from the server terminal to the client terminal. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Publication No. 2021-140539 Summary of the Invention [Problem to be solved by the invention]

[0007] In Patent Document 1, a depth map needs to be transmitted from a client terminal to a server terminal, which puts a strain on the communication bandwidth when realizing occlusion between a real object and a virtual object in MR or AR.

[0008] Therefore, an object of the present invention is to provide a technique for suppressing pressure on the communication capacity when realizing occlusion of one object by another object in MR or AR. [Means for solving the problem]

[0009] One aspect of the present invention is a method for producing a medicament for the treatment of a pulmonary arthritis. An information processing device that is a first device, a first acquisition means for acquiring a first image; a first map acquisition means for acquiring a first depth map indicating depth information corresponding to the first image; a reduction means for generating a second depth map by reducing an amount of information of the first depth map based on the meta information so as to represent a depth range corresponding to the meta information; and first transmitting means for transmitting the first image and the second depth map to a second device. The information processing device is characterized by the above.

[0010] One aspect of the present invention is a method for producing a medicament for the treatment of a pulmonary arthritis. A control method for an information processing device that is a first device, comprising: a first acquisition step of acquiring a first image; a first map acquisition step of acquiring a first depth map indicating depth information corresponding to the first image; a reduction step of generating a second depth map by reducing an amount of information of the first depth map based on the meta information so as to represent a depth range corresponding to the meta information; a first transmitting step of transmitting the first image and the second depth map to a second device; having The present invention relates to a method for controlling an information processing apparatus. [Effects of the Invention]

[0011] According to the present invention, when realizing occlusion of one object by another object in MR or AR, it is possible to suppress a shortage of communication bandwidth. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a configuration diagram of an information processing system according to a first embodiment. [Figure 2] 4 is a flowchart of processing of the information processing system according to the first embodiment. [Figure 3] 10 is a flowchart of a process for generating depth meta information according to the first embodiment. [Figure 4] 1 is a diagram showing the relationship between depth and depth according to the first embodiment. FIG. [Figure 5] 10 is a diagram showing the relationship between the depth and the depth according to Modification 1. FIG. [Figure 6] 10 is a flowchart of a process for generating depth meta information according to Modification 1. [Figure 7] FIG. 10 is a diagram illustrating a plurality of physical objects according to Modification 2. [Figure 8] 10 is a flowchart of a process for generating depth meta information according to Modification 2. [Figure 9] FIG. 11 is a configuration diagram of an information processing system according to a third modification. [Figure 10] FIG. 2 is a hardware configuration diagram of a first information processing apparatus according to the first embodiment. [Figure 11] FIG. 3 is a diagram illustrating depth meta information according to the first embodiment. [Figure 12] FIG. 10 is a diagram illustrating a simple map according to Modification 2. [Figure 13] 13 is a flowchart of a process for generating depth meta information according to Modification 4. [Figure 14] 13 is a flowchart of a simple map generation process according to a fourth modification. [Figure 15] 13 is a diagram illustrating a process of predicting a target range of depth meta information according to Modification 4. FIG. [Figure 16] FIG. 13 is a diagram illustrating the amount of information in a simple map according to Modification 5. DETAILED DESCRIPTION OF THE INVENTION

[0013] <Embodiment 1> In the first embodiment, an example of a video see-through MR system that can express reality and virtuality in a fusion manner will be described. In the first embodiment, it is possible to achieve both "high portability of the display device" and "high load CG rendering in the entire system." For this purpose, a head-mounted display device is used. The head-mounted display (HMD) is connected to the server that performs the CG rendering via wireless communication. Hereafter, the head-mounted display device will be referred to as "HMD."

[0014] Note that AR, which uses an optical see-through display device to superimpose a virtual image on real space, has a configuration different from the "image capture unit and synthesis unit" described below. However, the technology described in the first embodiment (a method for reducing the amount of information in a depth map) can also be applied to AR. Furthermore, the display device may not be an HMD, but may be a handheld device (such as a smartphone or tablet PC).

[0015] Representing occlusion is important to seamlessly combine virtual objects into real space. Representing occlusion generally requires depth information for both real and virtual objects. When performing CG rendering, a depth map is used to determine the front and back of a 3D model. These depth maps are generally expressed in 16 to 32 bits. Therefore, if the depth map generated during rendering is sent directly to the client device, it significantly strains the communication bandwidth.

[0016] On the other hand, if you only want to achieve occlusion, you may need a smaller amount of information than is required for CG rendering. For example, if you only want to achieve occlusion between a hand (or a real object held in the hand) and a virtual object, depth information within a range of about one meter from the viewpoint (HMD) is enough to achieve occlusion. This allows you to reduce the number of bits in the depth information.

[0017] Therefore, an information processing system (MR system) that reduces the amount of depth information used for transmission and reception is described below. Note that, below, the distance in the optical axis direction from the viewpoint (= the reference position for imaging) in three-dimensional space is called "depth." The value obtained by converting "depth" into a pixel value on a depth map (pixel value of a distance image) is also called "depth."

[0018] FIG. 1 shows an overall block diagram of an information processing system according to the first embodiment. The information processing system includes a first information processing device 101 and a second information processing device 102. The first information processing device 101 is a client terminal for presenting an MR space to a user. In the first embodiment, the first information processing device 101 is an HMD. The second information processing device 102 is a rendering server (server terminal) for generating a virtual image in the MR space.

[0019] (Regarding the first information processing device) The first information processing device 101 includes an imaging unit 103, a display unit 104, a position and orientation estimation unit 105, a depth estimation unit 106, a meta information generation unit 107, a transmission unit 108, a reception unit 109, a depth conversion unit 110, and a synthesis unit 111. Each component of the first information processing device 101 may be implemented in the same housing of an HMD. Alternatively, the display unit 104 and the imaging unit 103 may be configured in the HMD, and the other blocks may be implemented in a portable housing such as a smartphone.

[0020] The imaging unit 103 has a camera that captures a color image (captured image) of the real space. The imaging unit 103 has a camera for capturing images used for position and orientation estimation and real depth estimation. The imaging unit 103 may have a separate camera for each purpose, or may achieve two or more purposes with one camera. Furthermore, each of the cameras for these purposes may have left and right stereo cameras. In the first embodiment, the imaging unit 103 is fixed to the HMD and moves in accordance with the movement of the HMD. In the first embodiment, the imaging unit 103 is , has a stereo camera to show different images to each of the user's left and right eyes.

[0021] The display unit 104 is a device for displaying an image. In the first embodiment, the display unit 104 uses a display mounted on an HMD worn by the user. The HMD requires an eyepiece display lens for observation. In addition, in a handheld device such as a smartphone, the display unit 104 may be the main display. In a handheld device, an eyepiece display lens is not required.

[0022] The position and orientation estimation unit 105 estimates the position and orientation (position and orientation of the viewpoint) of the imaging unit 103. The position and orientation estimation unit 105 estimates the position and orientation of the imaging unit 103, for example, by using Visual SLAM or Visual Odometry that uses images acquired by the imaging unit 103. At this time, the position and orientation estimation unit 105 may estimate the position and orientation of the imaging unit 103 by using an inertial measurement unit (IMU) in combination.

[0023] Furthermore, the position and orientation estimation unit 105 may estimate the position and orientation by other methods that do not use images acquired by the imaging unit 103. For example, the position and orientation estimation unit 105 may use an outside-in method such as motion capture. In the first embodiment, the virtual image generation unit 113 performs stereo rendering based on the estimated position and orientation. However, if the relative position and orientation between the stereo cameras is known in advance as a parameter, it is sufficient to estimate the position and orientation of either one of the cameras.

[0024] The depth estimation unit 106 estimates a depth map indicating depth information corresponding to a captured image (real image). Specifically, the depth estimation unit 106 estimates the depth of an object existing in the real space in front of the eyes of the HMD (user) and acquires a depth map of the real object (hereinafter referred to as a "real map"). The depth map has the same resolution as the captured image and is an image in which each pixel indicates a depth.

[0025] The meta information generation unit 107 generates depth meta information based on the reality map acquired by the depth estimation unit 106. The depth meta information is information necessary for the depth information reduction unit 114 to reduce the amount of information in the depth map of a virtual object. Hereinafter, the depth map of a virtual object will be referred to as a "virtual map."

[0026] The transmission unit 108 transmits the information on the position and orientation estimated by the position and orientation estimation unit 105 (position and orientation information) and the depth meta information to the second information processing apparatus 102.

[0027] The receiving unit 109 receives the virtual image and the virtual map with the amount of information reduced from the second information processing device 102. Hereinafter, the virtual map after the amount of information has been reduced by the second information processing device 102 will be referred to as a "simple map."

[0028] The depth conversion unit 110 converts at least one of the depths indicated by the real map and the simple map using depth meta information so that the depths indicated by the real map and the simple map can be compared, because the real map and the simple map are expressed in different formats.

[0029] The synthesis unit 111 synthesizes the image of the real space and the virtual image based on the depth (the depth of the real object and the depth corresponding to the virtual image) after processing by the depth conversion unit 110. In this way, the synthesis unit 111 generates a mixed reality image (synthetic image) in which the virtual object appears to exist consistently in the real space.

[0030] The configuration of the second information processing device 102 will be described. The image processing unit 110 includes a receiving unit 112, a virtual image generating unit 113, a depth information reducing unit 114, a transmitting unit 115, and a data holding unit 116.

[0031] The receiving unit 112 receives the position and orientation information and the depth meta information from the first information processing apparatus 101.

[0032] The data storage unit 116 includes a hard disk or a solid state drive. The data storage unit 116 stores data (shape data, material data, and animation data) for generating a virtual image. The stored data is loaded into a memory used by the CPU or a memory on the GPU when generating the virtual image.

[0033] The virtual image generation unit 113 performs rendering (generation) of a virtual image based on the position and orientation information, the virtual camera parameters, and the CG model data. In the first embodiment, the virtual image generation unit 113 generates a virtual image for the right eye and a virtual image for the left eye based on the relative position and orientation between the stereo cameras that are determined in advance to realize stereoscopic vision. When generating the virtual images, the virtual image generation unit 113 also generates a virtual map for the right eye and a virtual map for the left eye. Therefore, the virtual image generation unit 113 also serves as a map generation unit (map acquisition unit) that generates a virtual map indicating depth information corresponding to the virtual image. Note that, hereinafter, the terms "virtual image" and "virtual map" will be used without distinguishing between those for the right eye and those for the left eye. The virtual map is passed to the depth information reduction unit 114.

[0034] The depth information reduction unit 114 generates a simple map by reducing the amount of information in the virtual map acquired from the virtual image generation unit 113 based on the depth meta information. A specific method for generating the simple map will be described later.

[0035] The transmission unit 115 transmits the stereo virtual image (color information) generated by the virtual image generation unit 113 and the simple map generated by the depth information reduction unit 114 to the first information processing device 101.

[0036] 10 is a block diagram showing an example of the hardware configuration of a computer applicable to the first information processing apparatus 101. In addition to the imaging unit 103 and the display unit 104, the computer has a CPU 1001, a RAM 1002, a ROM 1003, a keyboard 1004, a mouse 1005, and a monitor 1006. The computer also has an external storage device 1007, a storage medium drive 1008, and an interface 1009. The second information processing apparatus 102 may also have the same hardware configuration as the first information processing apparatus 101.

[0037] The CPU 1001 is a control unit that controls the entire computer using programs and data stored in the RAM 1002 or the ROM 1003. The CPU 1001 executes each process that the first information processing apparatus 101 performs.

[0038] The RAM 1002 has an area for temporarily storing programs and data loaded from the external storage device 1007 or the storage medium drive 1008. The RAM 1002 also has an area for temporarily storing data received from the outside via the interface 1009. The data received from the outside may be, for example, a captured image. The RAM 1002 also has a work area used by the CPU 1001 when executing various processes. In other words, the RAM 1002 can provide various areas as needed.

[0039] The ROM 1003 stores the computer's setting data, boot program, and the like.

[0040] The keyboard 1004 and the mouse 1005 are input devices (operation members) that accept user operations. A computer user can input various instructions to the CPU 1001 by operating at least one of the keyboard 1004 and the mouse 1005.

[0041] The monitor 1006 is a display device different from the display unit 104, and has a CRT or LCD screen, etc. The monitor 1006 can display the results of processing executed by the CPU 1001 as images or text.

[0042] The external storage device 1007 is a large-capacity information storage device, such as a hard disk drive. The external storage device 1007 stores programs (such as an operating system (OS)) and data for causing the CPU 1001 to execute the processes described above as being performed by each information processing device. Such programs include programs executed by each component of each information processing device in the information processing system. Such data also includes data of the virtual space and what has been described above as known information.

[0043] The programs and data stored in the external storage device 1007 are loaded into the RAM 1002 as appropriate under the control of the CPU 1001. The CPU 1001 executes processing using the loaded programs and data, thereby executing the processes described above as being performed by each information processing device.

[0044] The storage medium drive 1008 reads out programs and data recorded on a storage medium (such as a CD-ROM or DVD-ROM). The storage medium drive 1008 also writes programs and data to the storage medium. Note that some or all of the programs and data described as being stored in the external storage device 1007 may be recorded on this storage medium. The programs and data read out by the storage medium drive 1008 from the storage medium are output to the external storage device 1007 or RAM 1002.

[0045] The interface 1009 is an interface for connecting to the imaging unit 103. The interface 1009 has an analog video port or a digital input / output port (such as IEEE1394). The interface 1009 may also have an Ethernet (registered trademark) port or the like for outputting to the display unit 104. Data input via the interface 1009 is output to the RAM 1002 or the external storage device 1007. If a sensor system is used to acquire position and orientation information, the sensor system is connected to the interface 1009.

[0046] A bus 1010 connects the above-mentioned components.

[0047] The processing of the first information processing apparatus 101 and the second information processing apparatus 102 according to the first embodiment will be described with reference to the flowchart of FIG.

[0048] In step S2010, depth meta information for reducing the amount of information in the virtual map is generated. Details of the processing in step S2010 will be described with reference to the flowchart in FIG.

[0049] In step S3010, the depth estimation unit 106 determines the area of ​​the real object to be occluded (hereinafter referred to as the "target area") based on the image (real image) acquired by the imaging unit 103. The real object that expresses occlusion with the virtual object can be specified (limited) depending on the purpose of use. For example, in verifying the workability of product assembly in the manufacturing industry, the parts to be assembled are expressed as virtual objects, and the user wearing the HMD can see the virtual objects by themselves. The user's real body interferes with the virtual object. In this case, it is sufficient to represent the occlusion between the user's body and the virtual object. Therefore, the user's body region is determined as the target region. Semantic segmentation using machine learning may be used to determine the body region. Furthermore, a method of "estimating skeletal information and defining a region based on the estimation result" may be used to determine the body region. Furthermore, a specified color range may be determined as the body region. Note that the method of determining the target region is not limited to the above.

[0050] In step S3020, the depth estimation unit 106 calculates the depth of the target region (region of the physical object) determined in step S3010. The depth may be calculated using semi-global matching or deep learning from a physical image (stereo image) obtained by the imaging unit 103, or may be calculated using a distance sensor or the like. An IR camera stereo provided in the imaging unit 103 may capture an image of a pattern projected onto a physical object by an IR dot projector, and the depth estimation unit 106 may calculate the depth by stereo matching of the captured image of the pattern.

[0051] Furthermore, if noise occurs in the depth estimation and an extreme value is generated, the depth estimation unit 106 may remove the value or perform a smoothing process.

[0052] The order of the processes in steps S3010 and S3020 may be reversed. Therefore, the depth estimation unit 106 may determine the target region based on the shape or depth information of the three-dimensionally reconstructed physical object.

[0053] The depth estimation unit 106 generates a depth map by converting the depth of each pixel calculated in this way into a depth. For the conversion between depth and depth, a conversion described later in step S2030 (the same conversion as that used to generate a depth map after deleting information using depth meta information) may be used, or an arbitrary conversion formula between near-far values, depths, and depths may be used.

[0054] In step S3030, the meta information generation unit 107 calculates the near value and the far value of the physical object. Then, the meta information generation unit 107 generates information summarizing the "near value, far value, depth expression format, and depth expression bit number" as depth meta information, as shown in Figures 11A and 11B. The meta information generation unit 107 may include the format of the depth conversion formula (information on the formula for converting depth to depth) in the depth meta information, as necessary.

[0055] Here, the near value is the depth value (minimum depth value) of the nearest part in the target area (physical object) calculated in steps S3010 and S3020. The far value is the depth value (maximum depth value) of the farthest part in the target area. In this case, if the depth value of the nearest part in the target area is smaller than the first threshold, the meta information generation unit 107 may determine the first threshold as the near value. Also, if the depth value of the farthest part in the target area is larger than the second threshold, the meta information generation unit 107 may determine the second threshold as the far value.

[0056] The depth representation format indicates whether depth is represented by an integer or a floating point in the simple map. The depth representation format also indicates whether the near value is associated with the greatest depth in the simple map or the far value is associated with the greatest depth. The depth representation format may be a predetermined format, or may be determined to match the characteristics of a renderer that generates a virtual image.

[0057] The number of bits to represent depth indicates how many bits are used to represent the depth of each pixel of the simple map. A preset value may be used as the number of bits to represent depth. The number of bits to represent depth may be transmitted by the transmitting unit 108. Alternatively, it may be a value according to the state of communication detected by the transmitting unit 115.

[0058] The format of the depth conversion formula is the format of the relational expression between the depth of the virtual image and the depth in the simple map. The format of the depth conversion formula indicates whether it is an inverse proportional formula or a linear formula. Specific conversions related to the inverse proportional formula will be described later with reference to step S2030.

[0059] In step S2020, the transmission unit 108 transmits the depth meta information to the second information processing device 102.

[0060] In step S2030, the depth information reduction unit 114 reconstructs, based on the depth meta information, the virtual map generated by the virtual image generation unit 113. In this way, the depth information reduction unit 114 reduces the amount of information in the virtual map and generates a simplified map.

[0061] Figure 4 shows the depth d of the virtual map (depth map before information reduction) for each pixel of the depth map, relative to the depth Z from the viewpoint (reference position). CG and the depth d of the simplified map (depth map after information reduction) R The relationship between the depth d on the vertical axis and the CG, d R The scales in indicate the depths that can be represented in each depth map. Note that in Figure 4, the intervals between the scale marks on the vertical axis are equal. However, when the depth representation format is a floating-point representation format, the exponent and mantissa are each fixed length, and the range of change of one bit at the end of the mantissa depends on the exponent, so the intervals between the scale marks on the vertical axis are not equal.

[0062] Generally, CG rendering takes into consideration the correspondence with the calculations that convert three-dimensional space into two-dimensional images through perspective transformation. Therefore, the relationship between depth Z and depth is such that depth corresponds linearly to the inverse of depth Z (1 / Z). Note that there are no particular restrictions on which depth, near or far, corresponds to the maximum depth that can be expressed in the depth map, as this depends on the renderer, etc.

[0063] 4, for simplicity, the relationship curve between depth Z and depth is expressed by the same curve before and after the reduction of the amount of information. However, the depth expression format may be different before and after the reduction of the amount of information, or the parameters for linear correspondence with the inverse of depth Z (1 / Z) may be different. Furthermore, the relationship curve between depth Z and depth may be expressed by a function other than 1 / Z.

[0064] In the first embodiment, the near value (near R ) and far value (far R ) The depth representation format and depth representation bit number indicated by the depth meta information can represent the depth d R The maximum or minimum value of the object in the virtual image is assigned. CG ) and far value (far CG ) and the depth d CG Within the range of R ) and far value (far R ) is determined. Then, the range is determined based on the depth d R can be converted to be represented by the following representation:

[0065] For example, for a virtual map, the near value (near CG ) is 0.01m, and the far value (far CG ) is 100m. d R is an unsigned 24-bit integer value. For simple maps, near R is 0.3m, and far R Let us assume that the depth is 0.5 m. We will explain the case where the depth representation is an 8-bit unsigned integer and the depth and the depth are inversely proportional to each other.

[0066] First, before reducing the amount of information, the depth Z and the depth d CG The relation between d CG = a1 / Z+b1, and after the amount of information is reduced, the depth Z and the depth d R The relation between d R= a2 / Z+b2. For each of the virtual map and the simple map, the minimum and maximum depth values ​​are The coefficients a1, b1, a2, and b2 can be calculated based on the correspondence between " and " near value and far value. " Specifically, before the amount of information is reduced, a1 = -167788.9 and b1 = 16778892.9 hold, and after the amount of information is reduced, a2 = -191.25 and b2 = 637.5 hold.

[0067] Using this relational expression, the depth can be calculated from the depth of each pixel in the virtual map, and by substituting this depth into the "relationship between depth and depth after information volume reduction," it can be converted to the depth after information volume reduction. At this time, the information volume is reduced by converting the depth to an integer and by truncating values ​​outside the range (0 to 255) to 255 or 0.

[0068] The equation for the relationship between depth Z and depth d is not limited to the above equation, and may be expressed as, for example, a linear equation using Z. In addition, although the minimum depth value is associated with the near value and the maximum depth value is associated with the far value in the above description, the association may be reversed. Note that the equation for the relationship between depth Z and depth d needs to be shared between the depth conversion unit 110 and the depth information reduction unit 114. For this reason, it is preferable that the depth conversion unit 110 and the depth information reduction unit 114 each have information on a predetermined equation. If this is not the case, information on the type of equation for the relationship between depth Z and depth d may also be included in the depth meta information.

[0069] In step S2040, the transmission unit 115 transmits the simple map together with the virtual image to the first information processing device 101. At this time, the transmission unit 115 may acquire information on the communication status (such as delay, loss rate, and amount of information waiting for communication) via hardware that realizes communication or an API of a communication library so that the degree of information reduction can be adjusted depending on the communication status. Then, the transmission unit 115 may transmit the acquired information to the first information processing device 101 so that the information can be referenced when the meta information generation unit 107 generates depth meta information. In other words, the meta information generation unit 107 may control (change) information on the depth representation format or information on the depth representation bit value depending on the communication status between the first information processing device 101 and the second information processing device 102.

[0070] In step S2050, the depth conversion unit 110 converts the depth of the simple map into a value comparable to the depth of the real map based on the depth meta information. However, if the correspondence relationship between the depth and the depth in the simple map is the same as the correspondence relationship between the depth and the depth in the real map, the depth conversion unit 110 does not convert the depth.

[0071] When the correspondence between depth and depth differs between the real map and the simple map, the depth conversion unit 110 performs conversion so that the correspondence between depth and depth in the two depth maps matches. Here, the correspondence between depth and depth in the real map is defined in step S3020. The correspondence between depth and depth in the simple map is also defined based on depth meta information in step S2030. Therefore, the depth conversion unit 110 can mutually convert the correspondence between the two depth maps via depth. There is no limitation to converting the correspondence between the real map and the simple map to match the correspondence between the real map and the simple map.

[0072] In step S2060, the synthesis unit 111 compares the depth of the reality map with the depth of the simple map, and synthesizes the reality image with the virtual image so that the context is consistently represented. Specifically, the synthesis unit 111 achieves occlusion by not rendering a virtual object in an area where a reality object is located in front of a virtual object. The reality image used in step S2060 may be the same image as the reality image used to generate the depth meta information, or may be an image acquired (captured) at a later time than the reality image used to generate the depth meta information. The reality map used in step S2060 may be estimated based on either image.

[0073] According to the first embodiment, the amount of information of the virtual map is reduced based on the depth meta information so as to represent only the depth range corresponding to the depth meta information, thereby reducing the amount of information transmitted from the second information processing device 102 to the first information processing device 101. Therefore, it is possible to suppress a shortage of communication bandwidth between the first information processing device 101 and the second information processing device 102.

[0074] <Variation 1> In the first embodiment, a real object (such as the body of a user wearing an HMD) with continuous depth is assumed as the target of occlusion. However, there are cases where not only the user's body or an object being held but also walls or furniture present in the space are targeted for occlusion. In such cases, when using the method described in the first embodiment, the near and far values ​​can be calculated by depth estimation over the entire field of view. However, when there is a certain depth between the "user's body or an object being held" and the "wall, furniture, etc.", the depth information indicating the depth therebetween is often unnecessary for implementing occlusion. Therefore, in the first modification, a method for suppressing such unnecessary information will be described. The overall block diagram of the information processing system according to the first modification is shown in FIG. 1, and the overall processing flowchart is shown in FIG. 2, and therefore, description thereof will be omitted.

[0075] In Modification 1, when there is a certain depth between the user's body or an object being held (hereinafter referred to as the "foreground") and a wall, furniture, or the like (hereinafter referred to as the "background"), the meta information generation unit 107 calculates near and far values ​​for the foreground and background, respectively. The meta information generation unit 107 does not express, as depth, a range between the foreground and background where no object exists.

[0076] As shown in Figure 5, the foreground near value is set to "Z Rn1 " and the foreground far value is "Z Rf1 " Also, the near value of the background is set to "Z Rn2 " and set the background far value to "Z Rf2 In this case, after reducing the amount of information in the depth map, Z Rn1 The depth corresponding to Rn1 " and Z Rf1 The depth corresponding to Rf1 " and Z Rn2 The depth corresponding to Rn2 " and Z Rf2 The depth corresponding to Rf2 Here, when the depth is expressed by the depth expression format and depth expression bit number indicated by the depth meta information, the depth d Rf1 and depth d Rn2 According to the adjacency, there is a depth Z where no real object exists. Rf1 and depth Z Rn2 The depth information between Z and Z can be omitted. Rn1 ,Z Rf1 ,Z Rn2 ,Z Rf2 and their corresponding depths d Rn1 ,d Rf1 ,d Rn2 ,d Rf2 The information is represented in a table.

[0077] In this case, "Z Rn1 ~Z Rf1 and d Rn1 ~d Rf1 "The relationship between" and "Z Rn2 ~Z Rf2 and d Rn2 ~d Rf2The relationship between d and d does not have to be expressible in the same formula. R Within the range of values ​​that can be expressed as a whole, the foreground d Rn1 ~d Rf1 The ratio of the depth information to the depth information may be increased to increase the resolution. The above table is shared between the first information processing device 101 and the second information processing device 102 as depth meta information, thereby reducing the depth information and ensuring consistency in the conversion.

[0078] Next, a specific method for separating the foreground and background and calculating the depth meta information described above will be described. This process corresponds to step S2010 shown in Fig. 2. Details of step S2010 will be described with reference to the flowchart in Fig. 6.

[0079] In step S6010, the depth estimation unit 106 defines (separates) a foreground physical object and a background physical object, and calculates depth maps for the foreground physical object and the background physical object. The foreground physical object can be determined by the method described in the first embodiment.

[0080] Note that the foreground and background real objects can be determined based on the assumption that "foreground real objects are often dynamic objects, and background real objects are often static objects." The depth estimation unit 106 reprojects the depth map of the previous frame from the viewpoint of the current frame and calculates the difference between the reprojected depth map and the depth map of the current frame. The depth estimation unit 106 may then determine an area where the difference occurs and that is within a predetermined depth range as the area of ​​the foreground real object (hereinafter referred to as the "foreground target area"). The depth estimation unit 106 may also determine an area other than the foreground target area as the area of ​​the background real object (hereinafter referred to as the "background target area").

[0081] When calculating a depth map of a background object region, the depth estimation unit 106 may reconstruct three-dimensional information from multi-viewpoint information accumulated between frames for non-foreground regions using a structure-from-motion technique or a machine learning technique. When three-dimensional reconstruction is required, any of the methods described in the first embodiment can be used. Furthermore, when noise occurs in the estimation of depth information and extremely different depths are generated, the depth estimation unit 106 may exclude the values ​​or perform smoothing processing.

[0082] In step S6020, the depth estimation unit 106 calculates the minimum depth value Z Rn1 and the maximum depth value Z Rf1 The depth estimation unit 106 calculates the minimum depth value Z Rn2 and the maximum depth value Z Rf2 Calculate.

[0083] In step S6030, the meta information generating unit 107 generates depth meta information. The depth meta information includes a depth representation format and a depth representation bit number. The depth meta information also includes a depth Z Rn1 ,Z Rn2 ,Z Rf1 ,Z Rf2 and the corresponding depth d Rn1 ,d Rn2 ,d Rf1 ,d Rf2 The depth meta information includes information on the type of "formula for converting depth to depth" for each of the foreground and background as needed. Rf1 Depth Z of the closest point Rn2 If the foreground and background indicate a closer depth, the foreground and background may be merged into one depth region.

[0084] In this way, by generating the depth meta information, the depth information reduction unit 114 can reduce the "depth Z Rn1 and depth Z Rn2 and depth Z Rf1 and depth Z Rf2The amount of information in the virtual map can be reduced so that it only represents the "range between . This makes it possible to generate a simplified map with even less information, and further reduces the burden on communication bandwidth.

[0085] In the first modification, the depth information is separated into two levels, foreground and background, in the depth direction, but may be separated into three or more levels. For example, the depth map may be clustered by its depth value, and a table may be generated that associates the "minimum and maximum depth values" with the "corresponding depths" for each class, and this may be used as the depth meta information. Alternatively, segmentation of the color image or the depth map may be used as the division method, and similar processing may be performed for each segment.

[0086] <Variation 2> In the first embodiment and the first modification, the information processing system focuses on the depth direction and removes information outside the depth range of the occlusion target region from the virtual map. In this way, the information processing system minimizes the impact of a decrease in resolution due to information volume reduction, thereby reducing the amount of information in the virtual map. In the second modification, the information processing system also focuses on the image coordinate axis (u, v) direction and extracts the range of only the occlusion target region from the virtual map. Then, the information processing system transmits information on the extracted range of the virtual map from the second information processing device 102 to the first information processing device 101.

[0087] The overall block diagram of the information processing system according to the second modification is shown in Fig. 1, and the overall processing flowchart according to the second modification is the same as the flowchart in Fig. 2. Therefore, the description thereof will be omitted.

[0088] 7 is a schematic diagram illustrating two occluding physical objects 702 and 703 in a depth map 701 estimated by the depth estimation unit 106. The processing of step S2010 in the flowchart of FIG. 2 will be described with reference to the flowchart shown in FIG.

[0089] In step S8010, the depth estimation unit 106 determines a target physical object, and then determines a target image region (hereinafter referred to as a "depth transmission region") whose depth is to be transmitted. The method for determining the target physical object is the same as the method described in the first embodiment and the first modification, and therefore description thereof will be omitted. For example, the depth estimation unit 106 performs a labeling process on a depth map 701 of the physical object to separate it into a physical object 702 and a physical object 703. Then, the depth estimation unit 106 determines a rectangular region 704 that encompasses the physical object 702 and a rectangular region 705 that encompasses the physical object 703. The depth estimation unit 106 determines the region of the multiple rectangular regions 704 and 705 as the depth transmission region. Note that in the second modification, each element of the depth transmission region is defined as a rectangle, but may be defined as a polygon, an ellipse, or another shape.

[0090] In step S8020, the depth estimation unit 106 calculates the depth of each physical object determined in step S8010. The processing of step S8020 can be implemented by processing similar to the processing of steps S3020 and S6020. Note that the order of the processing of steps S8010 and S8020 may be reversed, and the target region may be extracted from the shape and depth information of the three-dimensionally reconstructed physical object.

[0091] In step S8030, the meta information generation unit 107 generates depth meta information. In the second modification, the depth meta information includes information on the position and size of each of the rectangular areas 704 and 705. The depth meta information also includes "near value, far value, number of pixels in the range corresponding to the rectangular area 704 after the amount of information is reduced (number of pixels in the range corresponding to each rectangular area in the simple map), depth representation format, and depth representation bit count" corresponding to the rectangular area 704. The depth meta information also includes "near value, far value, number of pixels in the range corresponding to the rectangular area 705 after the amount of information is reduced, depth representation format, and depth representation bit count" corresponding to the rectangular area 705. Note that the depth meta information may also include information on the type of depth conversion formula for each of the rectangular areas 704 and 705, as necessary.

[0092] In the second modification, when an occlusion region is defined by a plurality of rectangles, the meta information generating unit 107 specifies the position and size of each rectangle by the coordinates of the diagonal vertices (for example, the bottom left and top right) on the depth map. However, the present invention is not limited to this method.

[0093] Next, the meta information generation unit 107 calculates the near value and the far value of the physical object corresponding to each of the rectangular areas 704 and 705 that make up the depth transmission area.

[0094] In the second modification, the size of the depth transmission region changes for each frame, and therefore the communication volume also fluctuates. Therefore, a temporary increase in communication volume is likely to cause communication instability. Therefore, a threshold Th for the amount of information in the simple map is set in advance so that the total amount of information in the simple map does not exceed a certain value. If the meta information generation unit 107 determines that the total amount of information in the simple map exceeds the threshold Th, the meta information generation unit 107 may lower the resolution of the simple map. Therefore, for example, if the meta information generation unit 107 determines that the total amount of information in the simple map exceeds the threshold Th, the meta information generation unit 107 controls the number of pixels in the range corresponding to each rectangular region in the simple map so that the total amount of information in the simple map is equal to the threshold Th.

[0095] Based on the depth meta information generated as described above, the depth information reduction unit 114 performs depth reduction, trimming outside the depth transmission area, down-conversion, and the like on the virtual map. Specifically, the depth information reduction unit 114 generates a simplified map as shown in Figures 12A and 12B by extracting ranges corresponding to each rectangular area from the virtual map. In this way, the depth information reduction unit 114 further reduces the amount of information in the depth map to be transmitted to the first information processing device 101.

[0096] <Variation 3> In the first embodiment, the first modification, and the second modification, the overall blocks of the information processing system have been described with reference to Fig. 2. However, the arrangement of the blocks constituting the first information processing device and the second information processing device is not necessarily limited to this.

[0097] 9, a position and orientation estimation unit 105 and a meta information generation unit 903 may be arranged in a second information processing device 902. In this case, the "reality map estimated by the depth estimation unit 106 of the first information processing device 901" and the "image of real space (captured image)" may be transmitted to the second information processing device 102 via the transmission unit 108. Using this information, the second information processing device 102 may estimate the position and orientation of the imaging unit 103 and generate depth meta information.

[0098] Furthermore, the meta information generation unit 903 does not need to directly use a real map (real depth information) to generate depth meta information. For example, if only a limited number of virtual objects are desired to be occluded and can be specified in advance, depth meta information may be generated based on the depth map (virtual map) of the virtual object in a manner similar to that described above. Specifically, the meta information generation unit 903 may calculate the maximum and minimum depth values ​​of a specific virtual object and generate depth meta information that includes this information. In this case, the depth information reduction unit 114 generates a simplified map from the virtual map based on the generated depth meta information.

[0099] Note that when there are multiple candidates for virtual objects to be occluded, even if depth information is reduced using depth meta information including information on the "maximum and minimum depth values" and "image area range" of all of the objects, the effect of reducing the amount of information is low. In this case, the virtual object to be occluded may be dynamically selected (changed) based on the user's operation state of the virtual object or the movement state of the virtual object. For example, when the user selects or moves a virtual object from among multiple virtual objects, only the selected or moved virtual object may be set as the target of occlusion. Furthermore, when the second information processing device 102 has three-dimensional information of real objects (e.g., hand positions or feature point information in three-dimensional space), virtual objects within a certain distance from these real objects may be set as the target of occlusion.

[0100] The generated depth meta information is transmitted together with the virtual image and the simple map to the first information processing device 901. The depth conversion unit 110 converts the depth of the simple map or the depth of the real map based on the depth meta information.

[0101] At least a part of the depth meta information may be preset information (predetermined set values). For example, the near value and the far value in the depth meta information may be the minimum and maximum depth values ​​in a predetermined depth range.

[0102] <Variation 4> Generally, wireless communication often requires a longer time (latency) than wired communication. In addition, when processing such as compressing and decompressing communication data using a conventional method is required in a limited bandwidth, the processing time is added to the latency. In the first and second variations, the first information processing device 101 generates depth meta information and transmits the depth meta information to the second information processing device 102. The second information processing device 102 generates a simple map that compresses the depth information of the virtual object based on the depth meta information. The second information processing device 102 then transmits this simple map to the first information processing device 101 together with color information of the virtual object. The first information processing device 101 then performs synthesis based on the simple map. That is, all communication time and each processing time for round trips between the first information processing device 101 and the second information processing device 102 are required. Therefore, if a real object to be occluded is moving and real images taken at different times are used for generating the depth meta information and for synthesis, this delay may cause the real object to deviate from the range of the real object described in the depth meta information during synthesis. As a result, appropriate occlusion may not be achieved.

[0103] In Modification 4, when generating a simple map based on depth meta information, the second information processing device 102 estimates the existence range of a real object at the time of acquisition of a real image used for synthesis. This addresses the above-mentioned problem of delay. The basic form of Modification 4 is the same as that described in Embodiment 1, Modification 1, and Modification 2, so only the differences will be described below. Note that the following description will be given on the assumption that the real image used to generate the depth meta information and the real image used for synthesis with the virtual image are acquired at different times (image capture times). Furthermore, the depth map used for synthesis is estimated (generated) based on the real image used for synthesis with the virtual image.

[0104] First, the meta information generation unit 107 of the first information processing device 101 calculates "clue information for estimating the existence range of a real object at the time of synthesis by the second information processing device 102" and adds the clue information to the depth meta information. Here, the processing performed in step S2010 of Fig. 2 will be described with reference to the flowchart of Fig. 13. Furthermore, the details of this processing will be described with reference to the schematic diagram of Fig. 15A.

[0105] Steps S13010 and S13020 are similar to steps S3010 and S3020 in the flowchart of FIG. 3 according to the first embodiment, and therefore a description thereof will be omitted.

[0106] In step S13030, the meta information generation unit 107 calculates (detects) the velocity of a physical object. In FIG. 15A, velocity 1505 of physical object 1503 is calculated. The components of the calculated velocity may differ depending on the base embodiment. For example, in the first embodiment and the first modification, the existence range of a physical object is determined by a near value and a far value, so the depth component of the velocity is necessary. In this case, the velocities of the foremost and farthest parts of the physical object may be calculated based on the "elapsed time from the previous frame to the current frame" and the "changes in the near value and the far value." Only one representative velocity value may be calculated for the physical object based on the average of the near value and the far value. In the first modification, the velocity is determined for each physical object.

[0107] In Modification 2, a presence range 1502 of a physical object 1503 as shown in FIG. 15A is determined by UV components (components of a plane perpendicular to the depth) in addition to depth components. Therefore, the velocity of the UV components is also necessary. In this case, the meta information generation unit 107 may calculate the velocity of each of the diagonal points of a rectangle indicating the presence range 1502 of the physical object, or may calculate the velocity 1505 of the center 1504 or center of gravity of the presence range 1502 as a representative value of the velocity. Furthermore, if noise is added to the velocity, the accuracy of predicting the presence range decreases. Therefore, the velocity may be calculated by smoothing, such as by using a moving average, using values ​​calculated in the past.

[0108] In step S13040, the meta information generation unit 107 calculates a delay time of the depth meta information for synthesis. The depth meta information is a real image (depth of a real object) acquired by the image capturing unit 103. The depth meta information is generated based on the real image used for estimating the depth of the real object (the real image used for estimating the depth of the real object). Therefore, the acquisition time of the real image used for estimating the depth of the real object is time T1 corresponding to the depth meta information. Meanwhile, during synthesis, the depth estimation unit 106 estimates the latest depth of the real object, so the acquisition time of the real image used for depth estimation at this time (= the real image used for synthesis) is set to time T2. Then, the delay time of the depth meta information for synthesis can be calculated as T2-T1 (the difference between time T2 and time T1). The calculated delay time is used to predict the existence range of the real object in the second information processing device 102, assuming that it will not fluctuate suddenly over a short period of time. Furthermore, if noise-like fluctuations are added to the delay time, the existence range will not be predicted correctly. For this reason, the delay time may be calculated by smoothing using a moving average or the like using values ​​calculated in the past.

[0109] In step S13050, the meta information generation unit 107 generates depth meta information so as to include information on the speed of a physical object and information on a delay time (delay time of the depth meta information) in addition to the information described in the first embodiment or each modification.

[0110] The processing performed by the second information processing device 102 will be described with reference to the flowchart of Fig. 14 and the schematic diagram of Fig. 15B. The flowchart of Fig. 14 shows processing that is executed instead of step S2030 in the flowchart of Fig. 2. In the second information processing device 102, the depth information reduction unit 114 predicts the existence range of a physical object at the time of synthesis, based on the depth meta information generated in steps S13010 to S13050. Then, the depth information reduction unit 114 generates a simple map based on the predicted existence range.

[0111] In step S14010, the receiving unit 112 acquires (receives) depth meta information.

[0112] In step S14020, the depth information reduction unit 114 predicts the range of the physical object at the time of synthesis based on information included in the depth meta information (information on the range of the physical object, and information on the speed and delay time). Specifically, the range of the physical object at the time of synthesis is predicted by moving the range of the latest physical object by a distance calculated from the product of the speed and the delay time (linear extrapolation by adding the product of the speed and the delay time to the range of the latest physical object). For example, the "near value, far value" or "depth representative value" of the physical object at the time of synthesis is predicted based on the depth speed. When the depth representative value is predicted, the depth information reduction unit 114 predicts the "near value, far value" at the time of synthesis by adding the amount of change in the predicted depth representative value to the "near value, far value." In configurations such as Modification 2, the speeds corresponding to the two diagonal vertices or the center (or center of gravity) of a rectangle on the UV coordinate system are obtained within the range of the physical object, so the range of the physical object at the time of synthesis can be predicted in a similar manner.

[0113] Note that, because the prediction described above is based on linear extrapolation, a change in the velocity of a physical object can cause an error in the prediction. To reduce the impact of this error, the depth information reduction unit 114 may set the range of the predicted physical object larger by a certain margin. Furthermore, while the above description performs linear prediction, prediction using machine learning may also be performed. If a motion sensor such as an IMU is attached to the target physical object and angular velocity or acceleration information is obtained, these values ​​may be used to update or predict the latest range of the physical object.

[0114] In step S14030, the depth information reduction unit 114 generates a simple map using the same method as in the first embodiment, based on the predicted range of the physical object obtained in step S14020.

[0115] According to the fourth modification, the range of the real object at the time of capturing the real image used for synthesis is predicted, and a simple map is generated based on the prediction result. The real image and the virtual image are synthesized based on the "real map generated based on the real object" and the "simple map." As a result, even if the real object has moved, a simple map is generated that takes the movement into consideration, so that appropriate occlusion of the real object can be achieved.

[0116] <Variation 5> In Modification 2, the second information processing device 102 generates a simplified map having depth information only for an area where a real object, such as a hand, that is subject to occlusion is present. Furthermore, in Modification 3, the second information processing device 102 generates a simplified map having depth information for an area where a virtual object, such as a hand, that is subject to occlusion is present. This method reduces the effect of compressing depth information when objects that are subject to occlusion are ubiquitous within the screen. On the other hand, in Modification 5, the area is divided into an "area corresponding to an important object that requires more accurate occlusion (hereinafter referred to as an "important area")" and an "area (range) other than the important area." The second information processing device 102 then generates a simplified map by further compressing the depth information for the area corresponding to the area other than the important area. This reduces the data size of the simplified map (depth information) transmitted from the second information processing device 102 to the first information processing device 101. This process can also be applied to the configuration of Modification 4.

[0117] FIG. 16A shows a real image 1601 displayed in the configurations of Modifications 2 and 4. A rectangular area 1602 encompasses a real object 1603. FIG. 16B is a diagram schematically illustrating the magnitude relationship of the amount of information in a simple map corresponding to the real image 1601. For example, the hand area is considered important, and the depth information reduction unit 114 increases the amount of depth information per unit area in a rectangular area 1605 (rectangular range) corresponding to an important area encompassing the hand area. The depth information reduction unit 114 reduces the amount of depth information per unit area in an area 1606 (range) other than the rectangular area 1605 compared to the rectangular area 1605. This makes it possible to reduce the amount of information in the entire simple map.

[0118] Similarly, Fig. 16C shows a superimposed real image 1607 in the configurations of Modifications 3 and 4. A virtual object 1608 is an important virtual object designated as important by a preset setting or a user instruction. Fig. 16D is a diagram schematically showing the magnitude relationship of the amount of information in a simple map corresponding to real image 1607. In region 1611 corresponding to rectangular region 1610 that encompasses the important virtual object, the amount of depth information per unit area is large. The amount of depth information per unit area of ​​the other region 1612 (region 1612 corresponding to the region including virtual object 1609) is smaller than that of region 1611.

[0119] The above-mentioned area with a small amount of information is realized by relatively reducing the number of bits representing depth or by reducing the resolution. By including the setting of the method for reducing the amount of information in the depth meta information, it is possible to achieve consistency between the reduction of depth information in the second information processing device 102 and the depth conversion in the first information processing device 101. In other words, it is possible to achieve appropriate occlusion.

[0120] The information included in the depth meta information is described below. When the number of bits representing depth is relatively reduced, the depth meta information includes, in addition to information about the important region, "some or all of the depth representation bit number and depth representation format representing the depth of the region other than the important region, the near value, the far value, and the format of the depth conversion formula." The near value and the far value may be fixed values ​​or may be determined from the upper and lower limits of the depth in real space. When the resolution of the depth map is relatively reduced, the depth meta information includes the number of pixels in width and height indicating the size of the important region, and the number of pixels in width and height of the map representing the depth of the region other than the important region. Note that a process of relatively reducing the resolution and the number of bits may also be performed.

[0121] If you convert a simple map with reduced information directly and use it for occlusion, it will look unrealistic. Jaggies may occur at the boundary between the occlusion of an object and a virtual object. Therefore, to enable comparison between the simple map and the real map, the depth conversion unit 110 of the first information processing device 101 may convert the simple map to have the same amount of information (depth expression bit value or resolution) as the real map, and then perform smoothing processing.

[0122] In order to prevent the amount of information in the simple map from fluctuating significantly depending on whether or not an important area is present, if the important area does not exist in the real image (screen), a rectangular area (rectangular range) corresponding to the center of the real image may be set as the important area. The amount of depth information per unit area may be smaller in areas other than the center than in areas corresponding to the center.

[0123] Furthermore, in the above, "If A is greater than or equal to B, proceed to step S1; if A is less than (lower than) B, proceed to step S2" may be read as "If A is greater than (higher than) B, proceed to step S1; if A is less than or equal to B, proceed to step S2." Conversely, "If A is greater than (higher than) B, proceed to step S1; if A is less than (lower than) B, proceed to step S2" may be read as "If A is greater than (higher than) B, proceed to step S1; if A is less than (lower than) B, proceed to step S2." Therefore, unless a contradiction arises, "greater than or equal to A" may be read as "greater than (higher; longer; more) than A," and "less than or equal to A" may be read as "less than (lower; shorter; fewer) than A." Furthermore, "greater than (higher; longer; more) than A" may be read as "greater than or equal to A," and "less than (lower; shorter; fewer) than A" may be read as "less than or equal to A."

[0124] The various controls described above may or may not be performed by a single piece of hardware (e.g., a processor or circuit). The entire device may be controlled by multiple pieces of hardware (e.g., multiple processors, multiple circuits, or a combination of one or more processors and one or more circuits) sharing the processing.

[0125] The above processor is a processor in the broad sense, and includes general-purpose processors and dedicated processors. General-purpose processors include, for example, CPUs (Central Processing Units), MPUs (Micro Processing Units), and DSPs (Digital Signal Processors). Dedicated processors include, for example, GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), and PLDs (Programmable Logic Devices). Programmable logic devices include, for example, FPGAs (Field Programmable Gate Arrays) and CPLDs (Complex Programmable Logic Devices).

[0126] Although the embodiments of the present invention have been described in detail, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Furthermore, each of the above-described embodiments merely represents one embodiment of the present invention, and each embodiment can be combined as appropriate.

[0127] <Other embodiments> The present invention can also be realized by a process in which a program that realizes one or more functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in the computer of the system or device read and execute the program, or by a circuit that realizes one or more functions.

[0128] The disclosure of the above embodiments includes the following configurations, methods, and programs. (Configuration 1) An information processing device that is a first device, a first acquisition means for acquiring a first image; a first map acquisition means for acquiring a first depth map indicating depth information corresponding to the first image; a reduction means for generating a second depth map by reducing an amount of information of the first depth map based on the meta information so as to represent a depth range corresponding to the meta information; and first transmitting means for transmitting the first image and the second depth map to a second device. 1. An information processing device comprising: (Configuration 2) the meta information includes information on minimum and maximum depth values ​​of the first object; the reduction means generates the second depth map by reducing an amount of information in the first depth map so as to represent a range of depths between the minimum and maximum values ​​of the depths of the first object. 2. The information processing device according to configuration 1, (Configuration 3) the meta information includes information on a depth representation format for each pixel of the second depth map and information on the number of representation bits for the second depth map; 3. The information processing device according to configuration 2. (Configuration 4) The information on the number of representation bits in the meta information changes depending on a communication state between the first device and the second device. 4. The information processing device according to configuration 3. (Configuration 5) The meta information includes information regarding a relational expression for converting depth to depth. 5. The information processing device according to any one of configurations 2 to 4. (Configuration 6) the reduction means does not perform a process of reducing the amount of information of the first depth map when a communication state between the first device and the second device is a specific communication state; 6. The information processing device according to any one of configurations 2 to 5. (Configuration 7) the meta information includes information about the image region of the first object; 7. The information processing device according to any one of configurations 2 to 6. (Configuration 8) The meta information includes information on the number of pixels in a range corresponding to the image region in the second depth map. 8. The information processing device according to configuration 7. (Configuration 9) the information on the number of pixels in the range corresponding to the image area in the meta information changes depending on the communication status between the first device and the second device; 9. The information processing device according to configuration 8. (Configuration 10) the reduction means reduces an amount of information in the first depth map so as to represent a depth range corresponding to the meta information, and generates the second depth map by extracting a range corresponding to the image region from the first depth map based on the meta information. 10. The information processing device according to any one of configurations 7 to 9. (Configuration 11) The meta information includes information about a first region that is an image region and a second region that is not an image region. 10. The information processing device according to any one of configurations 7 to 9. (Configuration 12) The meta information includes information about a first region corresponding to a central portion of an image and information about a second region corresponding to a portion other than the central portion. 7. The information processing device according to any one of configurations 1 to 6. (Configuration 13) the reduction means generates the second depth map by reducing the amount of information of the first depth map so that the amount of depth information per unit area in the range corresponding to the second region is smaller than that in the range corresponding to the first region. 13. The information processing device according to configuration 11 or 12. (Configuration 14) At least a part of the information in the meta information is preset information. 14. The information processing device according to any one of configurations 1 to 13. (Configuration 15) generating means for generating the meta information based on the first depth map; the first image includes the first object; the first transmitting means transmits the meta information, the first image, and the second depth map to the second device; 11. The information processing device according to any one of configurations 2 to 10. (Configuration 16) The first object is designated in response to a user operation or movement of an object. 16. The information processing device according to configuration 15. (Configuration 17) An information processing system including the first device, which is the information processing device according to any one of configurations 1 to 14, and the second device, The second device is a second acquisition means for acquiring a second image; a second map acquisition means for acquiring a third depth map indicating depth information corresponding to the second image; generating means for generating the meta information based on the third depth map; a second transmitting means for transmitting the meta-information to the first device; having An information processing system comprising: (Configuration 18) the second image includes a first object and a second object; the meta information includes information on minimum and maximum depth values ​​of the first object and information on minimum and maximum depth values ​​of the second object; the reduction means generates the second depth map by reducing an amount of information in the first depth map so as to represent a range of depths between the minimum and maximum values ​​of the depth of the first object and a range of depths between the minimum and maximum values ​​of the depth of the second object. 18. The information processing system according to configuration 17. (Configuration 19) the second device further comprises a synthesis means for synthesizing the first image and the second image based on the second depth map and the third depth map. 19. The information processing system according to configuration 17 or 18. (Configuration 20) the second device further comprises a combining means for combining two images; the second acquisition means acquires a third image whose acquisition time is a second time that is later than a first time that is an acquisition time of the second image; the reduction means generates the second depth map based on a prediction result of a range of an object captured in the second image at the second time; the second map acquisition means acquires a fourth depth map indicating depth information corresponding to the third image; the combining means combines the first image and the third image based on the fourth depth map and the second depth map. 19. The information processing system according to configuration 17 or 18. (Configuration 21) the generating means generates the meta information including information on a difference between the second time and the first time and information on a speed of an object captured in the second image; the reduction means predicts the range of the object appearing in the second image at the second time based on the meta information; 21. The information processing system according to configuration 20. (method) A control method for an information processing device that is a first device, comprising: a first acquisition step of acquiring a first image; a first map acquisition step of acquiring a first depth map indicating depth information corresponding to the first image; a reduction step of generating a second depth map by reducing an amount of information of the first depth map based on the meta information so as to represent a depth range corresponding to the meta information; a first transmitting step of transmitting the first image and the second depth map to a second device; having 2. A method for controlling an information processing apparatus comprising: (program) 17. A program for causing a computer to function as each means of the information processing device according to any one of configurations 1 to 16. [Explanation of symbols]

[0129] 101: First information processing device (second device), 102: second information processing device (first device), 113: Virtual image generation unit, 114: Depth information reduction unit, 115: Transmission unit

Claims

1. An information processing device that is a first device, a first acquisition means for acquiring a first image; a first map acquisition means for acquiring a first depth map indicating depth information corresponding to the first image; a reduction means for generating a second depth map by reducing an amount of information of the first depth map based on the meta information so as to represent a depth range corresponding to the meta information; and first transmitting means for transmitting the first image and the second depth map to a second device.

1. An information processing device comprising:

2. the meta information includes information on minimum and maximum depth values ​​of the first object; the reduction means generates the second depth map by reducing an amount of information in the first depth map so as to represent a range of depths between the minimum and maximum values ​​of the depths of the first object.

2. The information processing apparatus according to claim 1, wherein:

3. the meta information includes information on a depth representation format for each pixel of the second depth map and information on the number of representation bits for the second depth map; 3. The information processing apparatus according to claim 2, wherein:

4. the information on the number of representation bits in the meta information changes depending on a communication state between the first device and the second device; 4. The information processing apparatus according to claim 3,

5. The meta information includes information regarding a relational expression for converting depth to depth.

3. The information processing apparatus according to claim 2, wherein:

6. the reduction means does not perform a process of reducing the amount of information of the first depth map when a communication state between the first device and the second device is a specific communication state; 3. The information processing apparatus according to claim 2, wherein:

7. the meta-information includes information about an image region of the first object; 3. The information processing apparatus according to claim 2, wherein:

8. the meta information includes information on the number of pixels in a range corresponding to the image region in the second depth map; 8. The information processing apparatus according to claim 7,

9. the information on the number of pixels in the range corresponding to the image area in the meta information changes depending on the communication status between the first device and the second device; 9. The information processing apparatus according to claim 8,

10. the reduction means reduces an amount of information in the first depth map so as to represent a depth range corresponding to the meta information, and generates the second depth map by extracting a range corresponding to the image region from the first depth map based on the meta information.

8. The information processing apparatus according to claim 7,

11. the meta information includes information about a first region that is an image region and a second region that is not an image region; 8. The information processing apparatus according to claim 7,

12. the meta information includes information on a first region corresponding to a central portion of an image and information on a second region corresponding to a portion other than the central portion of the image; 2. The information processing apparatus according to claim 1, wherein:

13. the reduction means generates the second depth map by reducing the amount of information of the first depth map so that the amount of depth information per unit area in the range corresponding to the second region is smaller than that in the range corresponding to the first region.

12. The information processing apparatus according to claim 11,

14. At least a part of the information in the meta information is preset information.

2. The information processing apparatus according to claim 1, wherein:

15. generating means for generating the meta information based on the first depth map; the first image includes the first object; the first transmitting means transmits the meta information, the first image, and the second depth map to the second device; 3. The information processing apparatus according to claim 2, wherein:

16. the first object is designated in response to a user's operation or movement of an object; 16. The information processing apparatus according to claim 15,

17. An information processing system including the first device, which is the information processing device according to any one of claims 1 to 14, and the second device, The second device is a second acquisition means for acquiring a second image; a second map acquisition means for acquiring a third depth map indicating depth information corresponding to the second image; a generating means for generating the meta information based on the third depth map; second transmitting means for transmitting the meta-information to the first device; having An information processing system comprising:

18. the second image includes a first object and a second object; the meta information includes information on minimum and maximum depth values ​​of the first object and information on minimum and maximum depth values ​​of the second object; the reduction means generates the second depth map by reducing an amount of information in the first depth map so as to represent a range of depths between the minimum and maximum values ​​of the depth of the first object and a range of depths between the minimum and maximum values ​​of the depth of the second object.

18. The information processing system according to claim 17.

19. the second device further comprises a synthesis means for synthesizing the first image and the second image based on the second depth map and the third depth map.

18. The information processing system according to claim 17.

20. the second device further comprises a combining means for combining two images; the second acquisition means acquires a third image whose acquisition time is a second time that is later than a first time that is an acquisition time of the second image; the reduction means generates the second depth map based on a prediction result of a range of an object captured in the second image at the second time; the second map acquisition means acquires a fourth depth map indicating depth information corresponding to the third image; the combining means combines the first image and the third image based on the fourth depth map and the second depth map.

18. The information processing system according to claim 17.

21. the generating means generates the meta information including information on a difference between the second time and the first time and information on a speed of an object captured in the second image; the reduction means predicts a range of the object appearing in the second image at the second time based on the meta information; 21. The information processing system according to claim 20.

22. A control method for an information processing device that is a first device, comprising: a first acquisition step of acquiring a first image; a first map acquisition step of acquiring a first depth map indicating depth information corresponding to the first image; a reduction step of generating a second depth map by reducing an amount of information in the first depth map based on the meta information so as to represent a depth range corresponding to the meta information; a first transmitting step of transmitting the first image and the second depth map to a second device; having 2. A method for controlling an information processing apparatus comprising:

23. A program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Information system, terminal, server, and program

    JP2021140539A