Method for representing a 3D scene and its depth plane data, encoder, and display device.
By correlating depth planes with human visual acuity, the method reduces the number of depth planes required for accurate depth simulation in three-dimensional scenes, optimizing data processing and aligning with human perception capabilities.
Patent Information
- Application Number
- JP2023574350
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-02
- Filing Date
- 2022-06-02
- Publication Date
- 2026-05-18
- Estimated Expiration
- 2042-06-02
AI Technical Summary
Existing methods for representing three-dimensional scenes in stereoscopic, augmented, and virtual reality applications require a large number of depth planes, leading to excessive data processing, which is inefficient and not optimized for human visual perception capabilities.
The method reduces the number of depth planes by utilizing a 'Depth Perceptual Quantization' function that correlates depth with human visual acuity, determining depth planes based on the minimum discernible difference and incorporating spatial visual acuity, allowing for accurate depth simulation with fewer planes.
This approach reduces data processing requirements while ensuring accurate depth simulation that matches or exceeds human visual perception, making it more efficient and effective in rendering three-dimensional scenes.
Smart Images

Figure 0007861026000027 
Figure 0007861026000028 
Figure 0007861026000029
Abstract
Description
[Technical Field]
[0001] [Related applications] This application claims priority to U.S. Provisional Application No. 63 / 195898 and European Patent Application No. 21177381.7, both filed on 2 June 2021, which are incorporated herein by reference in their entirety. [Background technology]
[0002] Some stereoscopic, augmented reality, and virtual reality applications represent a three-dimensional scene as a series of images at different distances (depth planes) relative to the viewer. To render such a scene from a desired viewpoint, each depth plane can be processed sequentially and combined with other planes to simulate a two-dimensional projection of the three-dimensional scene at the desired viewer position. This two-dimensional projection can then be displayed on a head-mounted device, mobile phone, or other flat screen. By dynamically adjusting the two-dimensional projection based on the viewer's position, it is possible to simulate an experience of being inside the three-dimensional scene. [Overview of the Initiative]
[0003] Reducing the number of depth planes required to accurately represent a 3D scene is valuable because such a reduction reduces the amount of data that needs to be processed. In the embodiments disclosed herein, the reduction in the number of depth planes is achieved while ensuring that accurate simulations can be rendered that match or slightly exceed the human visual system's ability to perceive depth. The embodiments disclosed herein utilize a "Depth Perceptual Quantization" function or D that correlates the physical distance of depth (depth plane) with the capabilities of the human visual system, such as visual acuity. PQ Includes. D PQ Each depth plane calculated by this method is a constant "just noticeable difference" from adjacent planes.
[0004] In a first embodiment, a method for representing a three-dimensional scene stored as a three-dimensional dataset is disclosed. This method includes determining P depth plane depths along a first field of view direction relative to a first viewpoint. The separation ΔD between each proximal depth D of the depth plane depths and an adjacent distal depth (D+ΔD) is the minimum discernible difference determined by (i) the proximal depth D, (ii) a lateral offset Δx perpendicular to the first field of view direction and between the first and second viewpoints, and (iii) a field of view angle Δφ tilted by separation ΔD when viewed from the second viewpoint. The method also extracts P proxy images I from the three-dimensional dataset. k This includes generating a proxy 3D dataset containing the following: Generating a proxy 3D dataset includes, for each of the P depth plane depths, (i) constructing a 3D dataset, and (ii) generating a proxy image from at least one cross-sectional image of a plurality of cross-sectional images, each representing a cross-section of a 3D scene at each of the plurality of scene depths, thereby generating a proxy 3D dataset containing P proxy images.
[0005] In a second embodiment, the encoder includes a processor and memory. The memory stores machine-readable instructions. When the machine-readable instruction is executed by the processor, it controls the processor to perform the method according to any of the first embodiments.
[0006] In a third embodiment, the display device includes an electronic visual display, a processor, and memory. The memory stores machine-readable instructions, which, when executed by the processor, result in each of the P proxy images I k For k=0, 1, ..., (P-1), (i) The proxy image I is a linear function of the following equation. k Depth of each scene D k Decide:
number
[0007] In a fourth aspect, the method of representing depth plane data is, for each of a plurality of 2D images corresponding to each of a plurality of depths D in a 3D scene, (i) determining a normalized depth D' from the depth D, and (ii) calculating a normalized perceived depth D PQ equal to the following formula:
Equation
Brief Description of the Drawings
[0008] [Figure 1] It is a schematic diagram of a viewer looking at a 3D scene rendered by a display of a device.
[0009] [Figure 2] It is a schematic diagram showing a geometric derivation of an equation for the minimum perceptible difference as a function of viewing distance and lateral displacement.
[0010] [Figure 3] It is a schematic diagram showing the relationship between the lateral displacement and the viewing distance, horizontal screen resolution, and angular visual acuity in FIG. 2.
[0011] [Figure 4] It is a plot of the minimum perceptible difference as a function of viewing distance in a specific viewing configuration.
[0012] [Figure 5] In this embodiment, Figure 2 shows a plot of multiple depth plane depths recursively determined using the expression of the minimum discernible difference at depth derived from Figure 2.
[0013] [Figure 6] In this embodiment, Figure 5 is a graph showing the normalized depth as a function of the depth plane depth.
[0014] [Figure 7] This flowchart shows a method for representing a 3D scene stored as a 3D dataset in one embodiment.
[0015] [Figure 8] This is a flowchart illustrating a method for representing depth plane data in an embodiment. [Modes for carrying out the invention]
[0016] The apparatus and methods disclosed herein determine depth plane position based on the limits of spatial visual acuity (the ability to perceive detail). This approach differs from methods that rely on binocular visual acuity (the ability to perceive different images with two eyes). By utilizing spatial visual acuity, the embodiments disclosed herein ensure an accurate representation of high-frequency occlusions that exist when one object is obscured by another object from one observation position but visible from another observation position.
[0017] The depth plane positioning methods disclosed herein take motion parallax into account. Motion parallax is when an observer moves and observes a scene from a different viewpoint as they observe it. The change in images from two different viewpoints results in a strong depth cue. Other methods only consider the difference in viewpoint between the two eyes (typically 6.5 cm). Embodiments herein accommodate and are designed to accommodate much longer baselines, such as 28 cm of motion, which results in a more perceptual depth plane.
[0018] Figure 1 is a schematic diagram of a viewer 191 viewing a three-dimensional scene 112 rendered by the display 110 of the device 100. Examples of the device 100 include a head-mounted display, a mobile device, a computer monitor, and a television receiver. The device 100 also includes a processor 102 and a memory 104 communicatively coupled thereto. The memory 104 stores a proxy three-dimensional dataset 170 and software 130. The software 130 includes a decoder 132 in the form of machine-readable instructions and implements one or more functions of the device 100. As used herein, the term “proxy image dataset” means a memory-efficient representation, or proxy, of the original image dataset.
[0019] Figure 1 also includes an encoding device 160, which includes a processor 162 and a memory 164 communicatively coupled thereto. The memory 164 stores a 3D dataset 150, software 166, and a proxy 3D dataset 170. The software 166 includes an encoder 168 in the form of machine-readable instructions and implements one or more functions of the encoding device 160. In an embodiment, the encoder 168 generates a proxy 3D dataset 170 and P depth plane depths 174 from the 3D dataset 150. The device 100 and the encoding device 160 are communicatively connected via a communication network 101.
[0020] Each of the memories 104 and 164 is temporary and / or non-temporary and may include either or both volatile memory (e.g., SRAM, DRAM, compute RAM, other volatile memory, or any combination thereof) and non-volatile memory (e.g., FLASH, ROM, magnetic media, optical media, other non-volatile memory, or any combination thereof). Some or all of the memories 104 and 164 may be integrated into processors 102 and 162, respectively.
[0021] The 3D dataset 150 contains S transverse cross-section images 152. Each transverse cross-section image 152 represents each cross-section of the 3D scene at each scene depth 154 (0, 1, ..., S-1). S is more than P. The proxy 3D dataset 170 contains P proxy images 172 (0, 1, ..., P-1). For each depth plane depth 174(k), the encoder 168 generates a proxy image 172(k) from at least one transverse cross-section image 152. The index k is one of the P integers, for example, an integer between 0 and (P-1), including both ends. One of each scene depth 154 of at least one transverse cross-section image 152 is closest to the depth plane depth 174(k).
[0022] The decoder 132 decodes the proxy 3D dataset 170 and sends the decoded data to the display 110, which displays it as a 3D scene 112. The 3D scene 112 contains P proxy images 172 (0, 1, ..., P-1), each proxy image is located at a depth plane depth 174 (0, 1, ..., P-1) in the direction z parallel to the xy plane of the 3D Cartesian coordinate system 118. On the coordinate system 118, the depth plane depths 174 are z0, z1, ..., z along the z axis. P-1 This is shown as follows. Figure 1 also shows a three-dimensional Cartesian coordinate system 198 defining directions x', y', and z'. When viewed by viewer 191, the directions x, y, and z of coordinate system 118 are parallel to the respective directions x', y', and z' of coordinate system 198.
[0023] Calculation of perceptual depth Figure 2 is a schematic diagram showing the derivation of the formula for minimum discernible difference as a function of viewing distance. In Figure 2, object 221 is located at a distance D from the viewer 191, and object 222 is located at a distance ΔD behind it. From the observation position 211, object 222 is hidden by object 221. When viewer 191 moves to a new position 212 by a distance Δx, viewer 191 can observe object 222. The geometry can be described in terms of the difference Δφ between angles 231 and 232 shown in Figure 2, as shown in formula (1), where Δφ is the observer's angular visual acuity. For television and film production, the International Telecommunication Union Recommendation ITU-R BT.1845 specifies an observer with "normal" visual acuity of 20 / 20, or angular resolution Δφ = 1 / 60 degrees.
number
number
number
[0024] To use Equation 3, you need to specify the range of the depth plane. Recommendation ITU-R BT.1845 states that the closest distance at which the human eye can comfortably focus is D. min = 0.25m is specified. D max For this, we select a value such that the denominator reaches 0 and ΔD becomes infinite, which occurs in the following equation:
number
[0025] The value of Δx must also be specified. This is the minimum movement an observer must make to perceive the change in depth between object 221 and object 222. For images intended to be viewed on a display, this can be calculated from the "ideal viewing distance" defined in ITU-RB T.1845, where the width Δw of each pixel coincides with the visual acuity Δφ, as shown in Figure 3. When the horizontal resolution of the screen is Nx = 3840 pixels, the minimum viewing distance D min In this view, the distance from one edge of the screen to the other is given by equation 4:
number
[0026] The closest viewing distance D = D min When we calculate Δx for this, we get Δx = 0.28m, and therefore D max This equals 960m. Large movements can exceed the just-noticeable difference (JND), but since it is impossible for one observer to see from both positions simultaneously, working memory must be relied upon to compare views from both perspectives.
[0027] Figure 4 shows a plot of ΔD in meters for equation (3), with ΔD / D as a function of viewing distance D when Δφ = 1 / 60 degrees and Δx = 0.28 meters. At close range, very small changes in depth are visible (0.15 mm at D = 25 cm). Depth JND is given by the depth D max It increases with increasing distance as it approaches the target.
[0028] D min Start with D max Using equation 3, which increases by ΔD until it reaches , we can create a table of P depth plane depths 174 where each depth plane depth 174 differs from the last depth by a perceptual amount. The last depth plane is D = D maxThis is how it is set. Therefore, the proxy 3D dataset 170 becomes a memory-efficient representation, or proxy, of the 3D dataset 150. When the viewer 191 moves along the x' axis, the computational resources required for the device 100 to display and refresh the view 3D scene 100 are less for dataset 170 than for dataset 150.
[0029] Under the above conditions, the number of unique depth planes is P = 2890. To show a smooth, continuous gradient spanning half of the screen (for example, a railway disappearing from the bottom to the top of the screen, as shown in 3D scene 112), while allowing for an observer movement Δx = 0.28m, approximately 3000 unique depth planes are required.
[0030] Figure 5 shows the depth plane depth D for each of the 2890 depth planes mentioned above, with depth plane indices k=0 to k=2889. k This plot shows the mapping to 510, where D k is the depth of the k-th depth plane.
[0031] Functional fitting Multiple actual depths D are assigned to each depth plane depth D PQ It is possible to achieve a function fit (invertible) to mapping 510 which maps to . The function form of equation (5) is one such mapping, where depth plane depth D PQ This is the optimal mapping 510 for appropriately selected values of the exponent n and coefficients c1, c2, and c3. The right-hand side of equation (5) may have other forms without departing from the scope of this specification.
number
[0032] A more accurate function fitting can be obtained using the function form defined in equation (6), which is obtained by adding an exponent m to the right-hand side of equation (5). That is, equation (5) is a specific instance of equation (6) where m is equal to 1. In the embodiment, the exponent n = 1.
number
[0033] Depth plane depth D in equation (6) PQ This is an example of a depth plane depth of 174. D PQ Unless otherwise explicitly stated, each depth plane depth D PQ This is a normalized depth in the range of 0 to 1. In other embodiments, each depth plane depth D PQ It has a unit of length, D min From D max It is within the range.
[0034] Equation (7) is the inverted form of equation (6), and therefore the explicit formula for the normalization depth is D' = D / D max This is the depth plane depth D PQThese are functions of the coefficients c1, c2, and c3, as well as the exponents m and n.
number
[0035] Equation (8) is the indexed version of equation (7), and k / P d is D PQ Replace D' with D', and index k goes from 0 to P d The range is up to P d =(P-1). Equation (8) also includes the coefficient μ and the offset β.
number
[0036] In this embodiment, when the device 100 software 130 is executed by the processor, (i) each proxy image 172(0-P d For each of the normalized scene depths D') according to equation (8), k(ii) Determine each proxy image 172(0-P d ) normalized scene depth D' k Includes machine-readable instructions that control the processor to display on display 110 at the scene depth determined therefrom.
[0037] Figure 7 is a flowchart of method 700 for representing a three-dimensional scene stored as a three-dimensional dataset. In embodiments, method 700 is carried out in one or more embodiments of encoding device 160 and / or device 100. For example, method 700 may be carried out by at least one of (i) a processor 162 that executes computer-readable instructions of software 166, and (ii) a processor 102 that executes computer-readable instructions of software 130. Method 700 includes steps 720 and 730. In embodiments, method 700 also includes at least one of steps 710, 740 and 750.
[0038] Step 720 includes determining P depth plane depths along the first field of view direction relative to the first viewpoint. The separation ΔD between each proximal depth D of the depth plane depths and an adjacent distal depth (D+ΔD) is the minimum known difference determined by (i) the proximal depth D, (ii) a lateral offset Δx perpendicular to the first field of view direction and between the first and second viewpoints, and (iii) a field of view angle Δφ tilted by separation ΔD when viewed from the second viewpoint. In the example of step 720, encoder 168 determines depth plane depth 174.
[0039] In the embodiment, the field of view Δφ is 1 arcminute. In the embodiment, each of the P depth plane depths exceeds the minimum depth D0, and D k The steps for determining the depths of P depth planes, where k=1, 2, ..., (P-1), are given by depth D k +1=D k +ΔD k The process includes the step of repeatedly determining the separation ΔD k This is equal to the following equation, which is an example of equation (3):
number
[0040] In the embodiment, method 700 includes step 710, which includes determining a lateral offset Δx from the field of view angle Δφ and a predetermined minimum depth plane depth among P depth plane depths. In an example of step 710, software 166 determines the lateral offset Δx using equation (4), where D is equal to the depth plane depth 174(0).
[0041] Step 730 extracts P proxy images I from the 3D dataset. k This includes generating a proxy 3D dataset containing the following: Generating a proxy 3D dataset includes, for each of the P depth plane depths, (i) constructing a 3D dataset, and (ii) generating a proxy image from the P proxy images from at least one cross-sectional image of a plurality of cross-sectional images, each representing a cross-section of a 3D scene at each of the plurality of scene depths. In the embodiment, one of the scene depths of each of the at least one cross-sectional image is closest to the depth plane depth. In the example of step 730, encoder 168 generates a proxy 3D dataset 170 from 3D dataset 150. As shown in Figure 1, datasets 150 and 170 are cross-sectional image 152 and proxy image 172, respectively.
[0042] If at least one cross-sectional image in step 730 includes multiple cross-sectional images, step 730 may include step 732. Step 732 includes generating a proxy image which involves averaging the multiple cross-sectional images. The final depth plane is D max It can be constructed by averaging all depth values exceeding a certain value. The first depth plane is D minIt can be constructed by averaging all depth values below. In the example of step 732, the encoder 168 generates each proxy image 172 as the average of two or more cross-sectional images 152.
[0043] Step 740 involves determining, for each proxy image I of the P proxy images, where k = 0, 1, 2,..., (P - 1), each scene depth D' k of the proxy image I k as a linear function of the following equation: k including:
Equation
Equation
[0044] In an embodiment, step 740 includes reading the quantities D min , D max and P from the metadata of the 3D data set. For example, the quantities D min , D max and P can be stored as metadata of the 3D data set 150 read by the software 166. In an embodiment, each of D min and D max is a fixed-point value of 10 bits. When the fixed-point value is 0, each value is 0.25 meters and 960 meters. In an embodiment, P is a fixed-point value of 12 bits.
[0045] Step 750 includes displaying a proxy image I at each depth plane depth. In an example of step 750, the apparatus 100 displays at least one proxy image 172(k) at a depth plane depth 174(k) shown as z in the three-dimensional scene 112. If method 700 includes step 740, each depth plane depth of step 750 is equal to the depth D' of each scene of step 740. For example, the depth plane depth 174(k) is equal to the depth D' of the scene.
[0046] In an embodiment, steps 720 and 730 are performed by a first apparatus such as the encoding apparatus 160 of FIG. 1, and method 700 includes step 740. In such an embodiment, step 750 includes a step 752 of transmitting proxy three-dimensional data from the first apparatus to the second apparatus. The second apparatus performs a determination of each scene depth D and displays a proxy image. In an example of step 752, the encoding apparatus 160 transmits a proxy three-dimensional data set 170 to the apparatus 100 and does not generate or store the depth plane depth 174. In this example, the apparatus 100 performs step 740 to determine the depth plane depth 174.
[0047] FIG. 8 is a flowchart showing a method 800 for representing depth plane data. In an embodiment, method 700 is implemented in one or more aspects of the apparatus 100. For example, method 800 can be implemented by a processor 102 that executes computer-readable instructions of software 130.
[0048] Method 800 includes steps 810, 820, and 830, and each step is performed for each of a plurality of two-dimensional images corresponding to each of a plurality of depths D in a three-dimensional scene. In an embodiment, the cross-sectional image 152 constitutes a plurality of two-dimensional images, and the scene depth 154 constitutes a plurality of scene depths D.
[0049] Step 810 involves determining the normalized depth D' from the depth D. In the example of step 810, the software 130 determines each normalized depth from each scene depth 154.
[0050] Step 820 normalizes the perceived depth D according to equation (6). PQ This includes calculating the depth of each scene. In the example of step 820, software 130 calculates the depth of each scene to 154. max Divide by 174 to determine each depth plane depth. In this example, the depth plane depth is the normalized depth.
[0051] Step 830 is normalized perceptual depth D PQ The binary code value D B This includes representing them as follows. In the example of step 830, the software 130 represents each depth plane depth 174 as its respective binary code value. In the embodiment, the binary code value D B The bit depth is either 8, 10, or 12. Step 830 may also include storing each binary code value on a non-temporary storage medium, which may be part of memory 104.
[0052] combination of features The features described above and those claimed below can be combined in various ways without departing from the scope of this specification. The following listed examples illustrate some possible, non-limiting combinations.
[0053] (A1) A method for representing a three-dimensional scene stored as a three-dimensional dataset is disclosed. The method includes determining P depth plane depths along a first field of view direction relative to a first viewpoint. The separation ΔD between each proximal depth D of the depth plane depths and an adjacent distal depth (D+ΔD) is the minimum discernible difference determined by (i) the proximal depth D, (ii) a lateral offset Δx perpendicular to the first field of view direction and between the first and second viewpoints, and (iii) a field of view angle Δφ tilted by separation ΔD when viewed from the second viewpoint. The method also extracts P proxy images I from the three-dimensional dataset. kThis includes generating a proxy 3D dataset containing the following: Generating a proxy 3D dataset includes, for each of the P depth plane depths, (i) constructing a 3D dataset, and (ii) generating a proxy image from at least one cross-sectional image of a plurality of cross-sectional images, each representing a cross-section of a 3D scene at each of the plurality of scene depths, thereby generating a proxy 3D dataset containing P proxy images.
[0054] (A2) In the embodiment of method A1, the field of view Δφ is 1 arcminute.
[0055] (A3) Embodiments of methods A1 and A2 include determining a lateral offset Δx from the field of view angle Δφ and a predetermined minimum depth plane depth among P depth plane depths.
[0056] (A4) In any embodiment of method A1 to A3, each of the P depth plane depths exceeds the minimum depth D0, and D k The steps for determining the depths of P depth planes, where k=1, 2, ..., (P-1), are given by depth D k +1=D k +ΔD k This includes a step of repeatedly making decisions.
[0057] (A5) In the embodiment of method A4, separation ΔD k This is equal to the following equation:
number
[0058] (A6) In an embodiment of any one of methods A1 to A5, when generating a proxy image, at least one cross-sectional image includes multiple cross-sectional images from a plurality of cross-sectional images, and generating a proxy image includes averaging the plurality of cross-sectional images.
[0059] (A7) An embodiment of any one of the methods A1 to A6 is one in which each proxy image I of the P proxy images k For k=0, 1, 2, ..., (P-1), proxy image I k Each scene depth D' k This includes determining it as a linear function of the following equation:
number
[0060] (A8) Determining P depth plane depths and generating proxy 3D datasets are performed by the first device. The first device transmits proxy 3D data to the second device, the second device then transmits the scene depth D' of each scene. k This further includes executing the decision and displaying the proxy image.
[0061] (A9) In either embodiment of method A7 and A8, each scene depth D' k This is equal to:
number
[0062] (A10) The embodiment of A9 uses the metadata of a 3D dataset to obtain quantity D min , D max This includes reading P.
[0063] (A11) Any embodiment of method A9 and A10, D min and D max These are equal to 0.25 meters and 960 meters, respectively.
[0064] (A12) In an embodiment of any one of the methods A7 to A11, c1, m, and n are equal to 2,620,000, 5 / 4, and 3,845 / 4096, respectively.
[0065] (A13) In an embodiment of any one of methods A1 to A12, in the step of generating a proxy image, one of the scene depths of at least one cross-sectional image is closest to the depth plane depth.
[0066] (B1) Encoder including processor and memory. The memory stores machine-readable instructions. When a machine-readable instruction is executed by the processor, it controls the processor to perform the action described in any one of the A1-A13 sections.
[0067] (C1) The display device includes an electronic visual display, a processor, and memory. The memory stores machine-readable instructions, which, when executed by the processor, result in each of the P proxy images I k For k=0, 1, ..., (P-1), (i) The proxy image I is a linear function of the following equation. k Depth of each scene D k Decide:
number
[0068] (D1) The method for representing depth plane data is to represent each of the multiple 2D images corresponding to each of the multiple depths D in the 3D scene, (i) Determine the normalized depth D' from the depth D, (ii) Normalized perceptual depth D equal to the following equation PQ Calculating:
number
[0069] (D2) In an embodiment of method D1, multiple depths D are D PQ The minimum D that is equal to 0 min From D PQ The maximum D is equal to 1. max The range is up to c2, and c2 is -c1(D min / D max ) n It is equal to (c1 + c2 - 1), and c3 is equal to (c1 + c2 - 1).
[0070] (D3) In an embodiment of any one of methods D1 to D2, c1 is equal to 2,620,000, n is equal to 3,872 / 4, and m is equal to 5 / 4.
[0071] (D4) In an embodiment of any one of the methods D1 to D3, the binary code value D B The bit depth is either 8, 10, or 12.
[0072] (D5) In an embodiment of any one of the methods D1 to D4, the binary code value D B The further step includes storing the data in a non-temporary storage medium.
[0073] (E1) The device includes a non-temporary storage medium and a bitstream stored in the non-temporary storage medium. The bitstream includes depth distance data, which is given by the following formula:
number
[0074] (F1) The decoding method is that each of the P proxy images I k For k=0, 1, 2, ..., (P-1), (i) Proxy image I k Each scene depth D' k This involves determining it as a linear function of the following equation:
number
[0075] (F2) In an embodiment of method F1, each scene depth D' k This is equal to:
number
[0076] (F3)Either embodiment of F1 or F2 obtains quantity D from the metadata of the 3D dataset. min , D max This includes reading P.
[0077] (F4) In any embodiment of method F1 to F3, D min and D max These are equal to 0.25 meters and 960 meters, respectively.
[0078] (F5) In an embodiment of any one of the methods F1 to F4, c1, m, and n are equal to 2620000, 5 / 4, and 3845 / 4096, respectively.
[0079] (G1) An encoder including a processor and memory. The memory stores machine-readable instructions, which, when executed by the processor, control the processor to perform the actions described in any one of F1 to F5.
[0080] The above methods and systems can be modified without departing from the scope of these embodiments. Accordingly, it should be noted that matters included in the above description or shown in the accompanying drawings should be interpreted as illustrative, not restrictive. In this specification, unless otherwise indicated, the phrase "in an embodiment" is equivalent to the phrase "in a particular embodiment" and does not refer to all embodiments. The following claims are intended to cover, and for linguistic reasons, all general and specific features described herein, as well as all descriptions of the scope of the methods and systems, and can be said to include them.
Claims
1. A method for reducing the number of depth planes in a 3D scene stored as a 3D dataset, A step of receiving a lateral offset Δx perpendicular to the first field of view and between a first viewpoint and a second viewpoint, wherein the lateral offset Δx is the minimum distance that an observer must take to perceive a change in depth between a first object at a proximal depth D along the first field of view and a second object at an adjacent distal depth (D+ΔD) along the first field of view. The steps include receiving the field of view angle Δφ, which represents the angular visual acuity of the observer, A step of receiving the three-dimensional dataset containing S cross-sectional images, wherein each cross-sectional image corresponds to a depth plane depth and represents each cross-section of the three-dimensional scene at each scene depth along the first field of view direction with respect to the first viewpoint, A step of determining P depth plane depths along the first field of view direction with respect to the first viewpoint, wherein the separation ΔD between each proximal depth D among the P depth plane depths and the adjacent distal depth (D+ΔD) is the minimum known difference determined by (i) the proximal depth D, (ii) the lateral offset Δx, and (iii) the field of view angle Δφ tilted by the separation ΔD as viewed from the second viewpoint, and P is less than S. A step of generating a proxy 3D dataset containing P proxy images from the received 3D dataset, comprising the step of generating a proxy image from the P proxy images from at least one cross-sectional image of the S cross-sectional images for each of the P depth plane depths of the depth plane. A method that includes this.
2. The step of receiving the aforementioned lateral offset Δx is Δx = Nx・D min The process includes the step of determining the horizontal offset Δx by calculating tan(Δφ), where Nx is the horizontal screen resolution and D min The method according to claim 1, wherein is a predetermined minimum depth plane depth of the P depth plane depths.
3. The method according to claim 2, wherein the step of generating a proxy image from at least one cross-sectional image among the S cross-sectional images includes the step of generating the proxy image from a plurality of cross-sectional images among the S cross-sectional images, and the step of generating the proxy image includes the step of averaging the plurality of cross-sectional images.
4. The method according to claim 3, wherein the step of generating a proxy image from at least one cross-sectional image among the S cross-sectional images includes the step of generating the proxy image from at least one cross-sectional image closest to each depth plane depth.
5. Each of the P depth plane depths is a predetermined minimum depth plane depth D min That's all. k The steps are given by k = 0, 1, 2, ..., (P-1), and the steps to determine the P depth plane depths are: depth D k+1 =D k +ΔD k The method according to claim 4, comprising the step of repeatedly determining.
6. The separation ΔD k teeth, [Math 1] The method according to claim 5, which is equivalent to the method according to claim 5.
7. For each proxy image I of the P proxy images k where k = 0, 1, 2,..., (P - 1), Linear functions: [Math 2] Proxy image I k Each of the approximated normalized depth plane depths D' k A step in which m, n, c 1 , c 2 , and c 3 This is the approximated normalized depth plane depth D'. k However, the corresponding depth plane depth D determined according to the method of claim 6 k Selected to be an approximation of the normalized value of P d = (P-1), and k / P d This is normalized perceptual depth D PQ Steps to represent the discrete representation of, The approximated normalized depth plane depth D' k Proxy image I at depth plane depth determined from k The steps to display and The method according to claim 6, further comprising:
8. The steps of determining the P depth plane depths and generating the proxy 3D dataset are performed by the first device. The step of transmitting the proxy 3D data from the first device to the second device, wherein the second device receives the approximated normalized depth plane depth D' k The method according to claim 7, further comprising the step of performing a decision and displaying the proxy image.
9. The P equally spaced normalized depth plane depths are in the range of 0 to 1, and c 3 = c 1 +c 2 -1 and c 2 = -c 1 (D min / D max ) n D min and D max These are the minimum and maximum scene depths of the aforementioned three-dimensional scene, respectively. The method according to claim 8.
10. It is a device, Processor and Memory that stores machine-readable instructions, It has, A device wherein, when the machine-readable instruction is executed by the processor, the processor controls the processor to perform the method according to any one of claims 1 to 9.
11. A display device, Electronic visual displays and Processor and Memory that stores machine-readable instructions, It has, A display device in which, when the machine-readable instruction is executed by the processor, the processor controls the processor to perform the method according to any one of claims 1 to 9 and to display the generated proxy image on the electronic visual display.
12. A method for mapping scene depth to normalized perceived depth associated with depth plane data of a three-dimensional scene, wherein the method is: Minimum scene depth D min The step of receiving, Maximum scene depth D max The step of receiving, For each of the multiple two-dimensional images corresponding to each of the multiple scene depths D within the aforementioned three-dimensional scene, D / D max The steps include calculating and determining the normalized depth D' from the aforementioned scene depth D, The normalized perceptual depth D is equal to the following equation. PQ : [Math 3] The steps to calculate, The normalized perceptual depth D PQ The binary code value D B A step represented as m, n, c 1 , c 2 , and c 3 The steps are determined according to the method of claim 9, A method that includes this.
13. The aforementioned binary code value D B The method according to claim 12, further comprising the step of storing in a non-temporary storage medium.
14. A method for mapping normalized perceived depth associated with depth plane data of a 3D scene to normalized depth distance values, wherein the method is: Multiple normalized perceptual depths D within the aforementioned 3D scene PQ For each of the multiple 2D images corresponding to each of the following, The step of calculating the normalized depth distance value D' as a linear function of the following equation is: [Math 4] D PQ This is a normalized value, and 0 ≤ D PQ The following conditions satisfy ≤ 1, m, n, c 1 , c 2 , and c 3 The steps are determined according to the method of claim 9. A method that includes this.
15. It is a device, Non-temporary storage media and The bitstream stored in the aforementioned non-temporary storage medium, The bitstream includes depth distance data, the depth distance data is a binary code value D representing a normalized depth distance value D' determined according to the method of claim 14. B A device that is encoded in [this format].