Three-dimensional model encoding device, three-dimensional model decoding device, three-dimensional model encoding method, and three-dimensional model decoding method
By shifting and correcting invalid areas in two-dimensional images, the three-dimensional model encoding and decoding technologies improve data compression efficiency and reduce data distribution.
Patent Information
- Application Number
- JP2024111435
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-10-27
- Filing Date
- 2024-07-11
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2038-10-25
AI Technical Summary
Existing three-dimensional model encoding and decoding technologies require improvements to reduce the amount of data distributed and enhance compression efficiency.
A three-dimensional model encoding device and method that generate a corrected two-dimensional image by shifting and correcting invalid areas, followed by two-dimensional encoding, and a decoding device and method that reconstruct the three-dimensional model from encoded data, reducing data distribution and improving coding efficiency.
The proposed solution reduces the amount of data distributed and enhances coding efficiency by correcting invalid areas in the two-dimensional image, allowing for efficient reconstruction of three-dimensional models.
Smart Images

Figure 0007781217000001 
Figure 0007781217000002 
Figure 0007781217000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a three-dimensional model encoding device, a three-dimensional model decoding device, a three-dimensional model encoding method, and a three-dimensional model decoding method. [Background technology]
[0002] Patent Document 1 discloses a method for transferring three-dimensional shape data. In Patent Document 1, three-dimensional shape data is sent over a network for each element, such as a polygon or voxel. The receiving side then imports the three-dimensional shape data and displays an image of each received element. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 9-237354 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the technology disclosed in the above patent document requires further improvement. [Means for solving the problem]
[0005] In order to achieve the above object, a three-dimensional model encoding device according to one embodiment of the present disclosure includes a projection unit that generates a two-dimensional image including one or more effective areas by projecting a three-dimensional model onto at least one or more two-dimensional planes, a correction unit that generates a corrected image by correcting the two-dimensional image, and an encoding unit that generates encoded data by two-dimensionally encoding the corrected image, wherein the correction of the two-dimensional image includes a shift process that shifts the position of the effective area on the two-dimensional image.
[0006] In addition, a three-dimensional model decoding device according to one embodiment of the present disclosure includes a decoding unit that acquires coded data in which a corrected image obtained by correcting a two-dimensional image including one or more effective areas in which a three-dimensional model is projected onto at least one two-dimensional plane is coded, and outputs a three-dimensional model obtained by decoding the acquired coded data, wherein the correction of the two-dimensional image includes a shift process that shifts the position of the effective areas on the two-dimensional image.
[0007] In order to achieve the above object, a three-dimensional model encoding device according to one embodiment of the present disclosure includes a projection unit that generates a two-dimensional image including one or more valid areas onto which a three-dimensional model is projected onto at least one or more two-dimensional planes; a correction unit that generates a corrected image by correcting one or more pixels that constitute invalid areas included in the two-dimensional image onto which the three-dimensional model is not projected, using the two-dimensional image after a shift process has been performed to shift the position of the valid areas on the two-dimensional image; and an encoding unit that generates encoded data by two-dimensionally encoding the corrected image.
[0008] In addition, a three-dimensional model decoding device according to one embodiment of the present disclosure includes a decoding unit that acquires encoded data in which the corrected image is encoded, the corrected image being an image in which a two-dimensional image including one or more valid areas onto which a three-dimensional model is projected onto at least one two-dimensional plane has been corrected, the corrected image being an image in which one or more pixels in invalid areas onto which the three-dimensional model is not projected, included in the two-dimensional image after a shift process has been performed to shift the position of the valid areas on the two-dimensional image, have been corrected, and that outputs the three-dimensional model obtained by decoding the acquired encoded data.
[0009] In order to achieve the above object, a three-dimensional model encoding device according to one embodiment of the present disclosure includes a projection unit that generates a two-dimensional image by projecting a three-dimensional model onto at least one two-dimensional plane, a correction unit that uses the two-dimensional image to generate a corrected image by correcting one or more pixels that constitute an invalid area included in the two-dimensional image onto which the three-dimensional model is not projected, and an encoding unit that generates encoded data by two-dimensionally encoding the corrected image.
[0010] In addition, a three-dimensional model decoding device according to one embodiment of the present disclosure includes a decoding unit that encodes encoded data into a corrected image, which is a corrected two-dimensional image generated by projecting a three-dimensional model onto at least one two-dimensional plane, and in which one or more pixels in an invalid area included in the two-dimensional image onto which the three-dimensional model was not projected are corrected, and outputs a three-dimensional model obtained by decoding the acquired encoded data.
[0011] These general or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium. [Effects of the Invention]
[0012] The present disclosure can provide a three-dimensional model encoding device, a three-dimensional model decoding device, a three-dimensional model encoding method, and a three-dimensional model decoding method that can reduce the amount of data distributed. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a diagram illustrating an overview of a free viewpoint video generation system according to the first embodiment. [Figure 2] FIG. 2 is a block diagram showing the configuration of the three-dimensional space recognition system according to the first embodiment. [Figure 3] FIG. 3 is a diagram illustrating an outline of the operation of the three-dimensional space recognition system according to the first embodiment. [Figure 4] FIG. 4 is a block diagram showing a configuration of a free viewpoint video generation system according to the first embodiment. [Figure 5] FIG. 5 is a diagram illustrating an outline of the operation of the free viewpoint video generation system according to the first embodiment. [Figure 6] FIG. 6 is a flowchart showing the operation of the free viewpoint video generation system according to the first embodiment. [Figure 7] FIG. 7 is a diagram illustrating a method for generating a foreground model according to the first embodiment. [Figure 8] FIG. 8 is a block diagram showing the configuration of a next-generation surveillance system according to the second embodiment. [Figure 9] FIG. 9 is a diagram illustrating an outline of the operation of the next-generation surveillance system according to the second embodiment. [Figure 10] FIG. 10 is a flowchart showing the operation of the next-generation surveillance system according to the second embodiment. [Figure 11] FIG. 11 is a block diagram showing a configuration of a free viewpoint video generation system according to the third embodiment. [Figure 12] FIG. 12 is a flowchart showing the operation of the free viewpoint video generation system according to the third embodiment. [Figure 13] FIG. 13 is a diagram illustrating an example of delivery of a foreground model and a background model according to the third embodiment. [Figure 14] FIG. 14 is a diagram illustrating an example of delivery of a foreground model and a background model according to the third embodiment. [Figure 15] FIG. 15 is a block diagram showing the configuration of a next-generation surveillance system according to the fourth embodiment. [Figure 16] FIG. 16 is a flowchart showing the operation of the next-generation surveillance system according to the fourth embodiment. [Figure 17] FIG. 17 is a block diagram showing a configuration of a free viewpoint video generation system according to the fifth embodiment. [Figure 18] FIG. 18 is a block diagram showing the configuration of a next-generation surveillance system according to the fifth embodiment. [Figure 19] FIG. 19 is a block diagram showing a configuration of a free viewpoint video generation system according to the sixth embodiment. [Figure 20] FIG. 20 is a flowchart showing the operation of the free viewpoint video generation system according to the sixth embodiment. [Figure 21] FIG. 21 is a diagram for explaining the generation and restoration process of a three-dimensional model according to the sixth embodiment. [Figure 22]FIG. 22 is a diagram showing an example of a depth image according to the sixth embodiment. [Figure 23A] FIG. 23A is a diagram showing an example of image value allocation in a depth image according to Embodiment 6. In FIG. [Figure 23B] FIG. 23B is a diagram showing an example of image value allocation in a depth image according to Embodiment 6. In FIG. [Figure 23C] FIG. 23C is a diagram showing an example of image value allocation in a depth image according to Embodiment 6. In FIG. [Figure 24] FIG. 24 is a diagram showing an outline of a three-dimensional data encoding method for encoding three-dimensional data. [Figure 25A] FIG. 25A is a diagram showing an example of a two-dimensional image including a hole region. [Figure 25B] FIG. 25B is a diagram showing an example of a corrected image in which the hole region has been corrected. [Figure 26A] FIG. 26A is a diagram illustrating an example of correction of a hole region by linear interpolation. [Figure 26B] FIG. 26B is a diagram illustrating an example of correction of a hole region by linear interpolation. [Figure 27A] FIG. 27A is a diagram illustrating an example of correction of a hole region by nonlinear interpolation. [Figure 27B] FIG. 27B is a diagram illustrating an example of correction of a hole region by nonlinear interpolation. [Figure 28A] FIG. 28A is a diagram showing another example of correction. [Figure 28B] FIG. 28B is a diagram showing another example of correction. [Figure 28C] FIG. 28C is a diagram showing an example in which the correction in FIG. 28B is expressed in a two-dimensional image. [Figure 28D] FIG. 28D is a diagram showing another example of correction. [Figure 28E] FIG. 28E is a diagram showing another example of correction. [Figure 28F] FIG. 28F is a diagram showing another example of correction. [Figure 29] FIG. 29 is a block diagram illustrating an example of a functional configuration of a 3D model encoding device according to an embodiment. [Figure 30] FIG. 30 is a block diagram illustrating an example of a functional configuration of a 3D model decoding device according to the embodiment. [Figure 31] FIG. 31 is a flowchart showing an example of a 3D model encoding method performed by the 3D model encoding device according to the embodiment. [Figure 32] FIG. 32 is a flowchart showing an example of a 3D model decoding method performed by the 3D model decoding device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] In a three-dimensional encoding device and a three-dimensional encoding method that encode three-dimensional data, and in a three-dimensional decoding device and a three-dimensional decoding method that decodes encoded data into three-dimensional data, it is desirable to be able to reduce the amount of data to be distributed. Therefore, in a three-dimensional encoding device and a three-dimensional encoding method that encode three-dimensional data, it is necessary to improve the compression efficiency of the three-dimensional data.
[0015] The present disclosure aims to provide a three-dimensional encoding device and a three-dimensional encoding method that can improve the compression efficiency of three-dimensional data in a three-dimensional encoding device and a three-dimensional encoding method that encode three-dimensional data, as well as a three-dimensional decoding device and a three-dimensional decoding method that decodes encoded encoded data into three-dimensional data.
[0016] A three-dimensional model encoding device according to one embodiment of the present disclosure includes a projection unit that generates a two-dimensional image by projecting a three-dimensional model onto at least one or more two-dimensional planes, a correction unit that uses the two-dimensional image to generate a corrected image by correcting one or more pixels that constitute an invalid area included in the two-dimensional image onto which the three-dimensional model is not projected, and an encoding unit that generates encoded data by two-dimensionally encoding the corrected image.
[0017] According to this, the corrected image generated by correcting the invalid area is two-dimensionally coded, so that coding efficiency can be improved.
[0018] The correction unit may also correct the invalid area using a first pixel value of a first pixel in a first effective area, which is an effective area adjacent to the invalid area and onto which the three-dimensional model is projected.
[0019] This makes it possible to reduce the difference in pixel values between the first valid area and the invalid area, thereby effectively improving the coding efficiency.
[0020] The correction unit may correct the invalid area by setting all pixel values of the one or more pixels that make up the invalid area to the first pixel value.
[0021] This makes it possible to easily reduce the difference in pixel values between the first valid area and the invalid area.
[0022] In addition, the correction unit may correct the invalid area by further using a second pixel value of a second pixel in a second effective area, which is an effective area on the two-dimensional image on the opposite side of the invalid area from the first effective area.
[0023] This makes it possible to reduce the difference in pixel values between the first and second effective areas and the invalid area, thereby effectively improving the coding efficiency.
[0024] Furthermore, the correction unit may correct the invalid area by using the first pixel value and the second pixel value to change the pixel values of each of a plurality of pixels between the first pixel and the second pixel across the invalid area to pixel values that satisfy a relationship in which the positions and pixel values of each of the plurality of pixels change linearly from the first pixel value to the second pixel value.
[0025] According to this, since the invalid area is linearly interpolated, the processing load related to determining pixel values for interpolation can be reduced.
[0026] Furthermore, the two-dimensional encoding is a process of encoding the corrected image in units of multiple blocks, and when a boundary of the multiple blocks in the two-dimensional encoding is on the invalid area, the correction unit may correct the invalid area by changing a plurality of first invalid pixels between the first pixel and the boundary to the first pixel value and changing a plurality of second invalid pixels between the second pixel and the boundary to the second pixel value.
[0027] According to this, since the invalid area is corrected taking into consideration the boundaries of a plurality of blocks in two-dimensional encoding, the processing load can be effectively reduced and the encoding efficiency can be effectively improved.
[0028] Furthermore, the correction unit may correct the invalid area by using a first pixel value of the first pixel and a second pixel value of the second pixel to change the pixel values of each of a plurality of pixels between the first pixel and the second pixel across the invalid area to pixel values that satisfy a relationship in which the positions and pixel values of each of the plurality of pixels change in a smooth curve from the first pixel value to the second pixel value.
[0029] This makes it possible to effectively reduce the difference in pixel values between the first and second effective areas and the invalid area, thereby improving coding efficiency.
[0030] The first pixel may be a pixel adjacent to the invalid area in the first effective area, and the second pixel may be a pixel adjacent to the invalid area in the second effective area.
[0031] Furthermore, the present invention may further include a generation unit that generates a two-dimensional binary map that indicates whether each of a plurality of areas that constitute a two-dimensional area corresponding to the two-dimensional image is an invalid area or a valid area, and the encoding unit may generate the encoded data by encoding the corrected image and the two-dimensional binary map.
[0032] Therefore, during decoding, it is possible to use the two-dimensional binary map to decode only the valid area out of the valid area and invalid area, thereby reducing the amount of processing required during decoding.
[0033] A three-dimensional model decoding device according to one embodiment of the present disclosure includes a decoding unit that encodes encoded data into a corrected image, which is a corrected two-dimensional image generated by projecting a three-dimensional model onto at least one two-dimensional plane, and in which one or more pixels in an invalid area of the two-dimensional image onto which the three-dimensional model was not projected are corrected, and outputs a three-dimensional model obtained by decoding the acquired encoded data.
[0034] Therefore, the 3D model decoding device 310 can reconstruct a 3D model by acquiring a small amount of coded data.
[0035] A three-dimensional model distribution method according to one embodiment of the present disclosure distributes a first model, which is a three-dimensional model of a target space during a target time period, using a first distribution method, and distributes a second model, which is a three-dimensional model of the target space during the target time period and has smaller changes per hour than the first model, using a second distribution method different from the first distribution method.
[0036] This allows the 3D model distribution method to distribute the first and second models, which have different changes over time, using a distribution method that is appropriate for each model. This allows the 3D model distribution method to achieve appropriate distribution according to requests.
[0037] For example, the delivery cycle of the first delivery method may be shorter than the delivery cycle of the second delivery method.
[0038] According to this, the three-dimensional model distribution method can distribute the first model and the second model, which change over time differently, in a distribution method suited to each of them.
[0039] For example, the first distribution method may use a first encoding method, and the second distribution method may use a second encoding method that has a larger processing delay than the first encoding method.
[0040] According to this, the three-dimensional model distribution method can reduce the processing delay of the first model.
[0041] For example, a first encoding method may be used in the first distribution method, and a second encoding method having a different encoding efficiency from that of the first encoding method may be used in the second distribution method.
[0042] According to this, the three-dimensional model distribution method can use encoding methods suited to the first model and the second model, which change over time differently.
[0043] For example, the first delivery method may have a lower delay than the second delivery method.
[0044] According to this, the three-dimensional model distribution method can reduce the delay of the first model.
[0045] For example, the three-dimensional model distribution method may further include generating the first model by a first generation method, and generating the second model by a second generation method having a different accuracy from that of the first generation method.
[0046] According to this, the three-dimensional model distribution method can generate a first model and a second model that change differently over time using methods suited to each model.
[0047] For example, in generating the first model, the first model may be generated from a third model that is a three-dimensional model of a plurality of objects contained in the target space during the target time period, and the second model that is a three-dimensional model of some of the plurality of objects contained in the target space during the target time period, and the first model that is the difference between the third model and the second model.
[0048] According to this, the three-dimensional model distribution method can easily generate the first model.
[0049] For example, in generating the first model, a third multi-perspective image may be generated which is the difference between a first multi-perspective image in which a plurality of objects contained in the target space during the target time period is photographed and a second multi-perspective image in which some of the plurality of objects are photographed, and the first model may be generated using the third multi-perspective image.
[0050] For example, a terminal to which the first model and the second model are delivered may use the first model and the second model to generate a free viewpoint video, which is a video viewed from a selected viewpoint, and the three-dimensional model delivery method may prioritize delivery of the first model that is necessary for generating the free viewpoint video.
[0051] According to this, the 3D model distribution method can efficiently distribute information required for generating free viewpoint video.
[0052] A three-dimensional model distribution method according to one embodiment of the present disclosure generates a third model that is the difference between a first model, which is a three-dimensional model of a plurality of objects contained in a target space during a target time period, and a second model, which is a three-dimensional model of some of the plurality of objects contained in the target space during the target time period, from the first model and the second model, distributes the second model using a first distribution method, and distributes the third model using a second distribution method that is different from the first distribution method.
[0053] This allows the 3D model distribution method to distribute the second and third models using distribution methods that are appropriate for each model, thereby enabling the 3D model distribution method to achieve appropriate distribution according to requests.
[0054] A three-dimensional model distribution device according to one embodiment of the present disclosure includes a first distribution unit that distributes a first model, which is a three-dimensional model of a target space in a target time period, by a first distribution method, and a second distribution unit that distributes a second model, which is a three-dimensional model of the target space in the target time period and has smaller changes per hour than the first model, by a second distribution method different from the first distribution method.
[0055] This allows the 3D model distribution device to distribute the first model and the second model, which have different changes per unit time, using a distribution method that is appropriate for each model. This allows the 3D model distribution device to realize appropriate distribution according to requests.
[0056] A three-dimensional model distribution device according to one embodiment of the present disclosure includes a three-dimensional model generation unit that generates a third model that is a difference between a first model, which is a three-dimensional model of a plurality of objects included in a target space in a target time period, and a second model, which is a three-dimensional model of some of the plurality of objects included in the target space in the target time period, from the first model and the second model, and a distribution unit that distributes the second model using a first distribution method and distributes the third model using a second distribution method different from the first distribution method.
[0057] This allows the three-dimensional model distribution device to distribute the second model and the third model in a distribution method that is appropriate for each model, thereby enabling the three-dimensional model distribution device to realize appropriate distribution in response to requests.
[0058] A three-dimensional model distribution method according to one aspect of the present disclosure generates a depth image from a three-dimensional model, and distributes the depth image and information for reconstructing the three-dimensional model from the depth image.
[0059] According to this, instead of distributing the 3D model as is, a depth image generated from the 3D model is distributed, thereby reducing the amount of data to be distributed.
[0060] For example, the three-dimensional model distribution method may further include compressing the depth image using a two-dimensional image compression method, and distributing the compressed depth image.
[0061] This allows data to be compressed using a 2D image compression method when distributing 3D models, eliminating the need to create a new compression method specifically for 3D models, and making it easy to reduce the amount of data.
[0062] For example, in generating the depth images, multiple depth images from different viewpoints may be generated from the three-dimensional model, and in compressing the multiple depth images, the multiple depth images may be compressed using the relationship between the multiple depth images.
[0063] This allows the amount of data for multiple depth images to be further reduced.
[0064] For example, the three-dimensional model distribution method may further generate the three-dimensional model using multiple images captured by multiple imaging devices, distribute the multiple images, and the viewpoint of the depth image may be the viewpoint of any of the multiple images.
[0065] According to this, by matching the viewpoint of the depth image with the viewpoint of the captured image, for example, when the captured image is compressed by multi-view coding, it is possible to calculate disparity information between the captured images using the depth image and generate a predicted image between the viewpoints using the disparity information, thereby reducing the amount of code for the captured image.
[0066] For example, in generating the depth image, the depth image may be generated by projecting the three-dimensional model onto an imaging surface at a predetermined viewpoint, and the information may include parameters for projecting the three-dimensional model onto the imaging surface at the predetermined viewpoint.
[0067] For example, the three-dimensional model distribution method may further determine a bit length of each pixel included in the depth image, and distribute information indicating the bit length.
[0068] This allows the bit length to be switched depending on the subject or purpose of use, thereby making it possible to appropriately reduce the amount of data.
[0069] For example, the bit length may be determined in accordance with the distance to the subject.
[0070] For example, the three-dimensional model distribution method may further determine a relationship between a pixel value shown in the depth image and a distance, and distribute information indicating the determined relationship.
[0071] This allows the relationship between pixel values and distance to be switched depending on the subject or purpose of use, thereby improving the accuracy of the restored three-dimensional model.
[0072] For example, the three-dimensional model may include a first model and a second model that changes less per time than the first model, the depth image may include a first depth image and a second depth image, generating the depth image may include generating the first depth image from the first model and generating the second depth image from the second model, determining the relationship may include determining a first relationship between pixel values and distances shown in the first depth image and a second relationship between pixel values and distances shown in the second depth image, and in the first relationship, the distance resolution in a first distance range may be higher than the distance resolution in a second distance range that is farther than the first distance range, and in the second relationship, the distance resolution in the first distance range may be lower than the distance resolution in the second distance range.
[0073] For example, the three-dimensional model may have color information added thereto, and the three-dimensional model distribution method may further generate a texture image from the three-dimensional model, compress the texture image using a two-dimensional image compression method, and distribute the compressed texture image during the distribution.
[0074] A three-dimensional model receiving method according to one aspect of the present disclosure receives a depth image generated from a three-dimensional model and information for reconstructing the three-dimensional model from the depth image, and uses the information to reconstruct the three-dimensional model from the depth image.
[0075] According to this, instead of distributing the 3D model as is, a depth image generated from the 3D model is distributed, thereby reducing the amount of data to be distributed.
[0076] For example, the depth image may be compressed using a two-dimensional image compression method, and the three-dimensional model receiving method may further include decoding the compressed depth image.
[0077] This allows data to be compressed using a 2D image compression method when distributing 3D models, eliminating the need to create a new compression method specifically for 3D models, and making it easy to reduce the amount of data.
[0078] For example, the receiving may include receiving a plurality of depth images, and the decoding may include decoding the plurality of depth images using a relationship between the plurality of depth images.
[0079] This allows the amount of data for multiple depth images to be further reduced.
[0080] For example, the three-dimensional model receiving method may further generate a rendering image using the three-dimensional model and a plurality of images, and the viewpoint of the depth image may be the viewpoint of any one of the plurality of images.
[0081] According to this, by matching the viewpoint of the depth image with the viewpoint of the captured image, for example, when the captured image is compressed by multi-view coding, it is possible to calculate disparity information between the captured images using the depth image and generate a predicted image between the viewpoints using the disparity information, thereby reducing the amount of code for the captured image.
[0082] For example, the information may include parameters for projecting the three-dimensional model onto an imaging plane of the depth image, and the restoration may restore the three-dimensional model from the depth image using the parameters.
[0083] For example, the three-dimensional model receiving method may further include receiving information indicating a bit length of each pixel included in the depth image.
[0084] This allows the bit length to be switched depending on the subject or purpose of use, thereby making it possible to appropriately reduce the amount of data.
[0085] For example, the three-dimensional model receiving method may further include receiving information indicating a relationship between pixel values shown in the depth image and distances.
[0086] This allows the relationship between pixel values and distance to be switched depending on the subject or purpose of use, thereby improving the accuracy of the restored three-dimensional model.
[0087] For example, the three-dimensional model receiving method may further include receiving a texture image compressed using a two-dimensional image compression method, decoding the compressed texture image, and restoring the three-dimensional model with added color information using the decoded depth image and the decoded texture image.
[0088] A three-dimensional model distribution device according to one aspect of the present disclosure includes a depth image generation unit that generates a depth image from a three-dimensional model, and a distribution unit that distributes the depth image and information for restoring the three-dimensional model from the depth image.
[0089] According to this, instead of distributing the 3D model as is, a depth image generated from the 3D model is distributed, thereby reducing the amount of data to be distributed.
[0090] A three-dimensional model receiving device according to one aspect of the present disclosure includes a receiving unit that receives a depth image generated from a three-dimensional model and information for restoring the three-dimensional model from the depth image, and a restoration unit that uses the information to restore the three-dimensional model from the depth image.
[0091] According to this, instead of distributing the 3D model as is, a depth image generated from the 3D model is distributed, thereby reducing the amount of data to be distributed.
[0092] These comprehensive or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0093] Hereinafter, the embodiments will be described in detail with reference to the drawings. Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in an independent claim that represents a superordinate concept will be described as optional components.
[0094] (Embodiment 1) First, an outline of the present embodiment will be described. In this embodiment, a method for generating and distributing a 3D model in a 3D space recognition system such as a next-generation wide-area monitoring system or a free viewpoint video generation system will be described.
[0095] Figure 1 shows an overview of a free-viewpoint video generation system. For example, by using a calibrated camera (e.g., a fixed camera) to capture images of the same space from multiple viewpoints, it is possible to reconstruct the captured space in three dimensions (three-dimensional space reconstruction). This three-dimensional reconstructed data can be used for tracking, scene analysis, and video rendering to generate images viewed from any viewpoint (free-viewpoint camera). This will enable the realization of next-generation wide-area surveillance systems and free-viewpoint video generation systems.
[0096] In such a system, a 3D model generated by 3D reconstruction is distributed via a network, etc., and tracking, scene analysis, video rendering, etc. are performed on the receiving terminal side. However, the huge amount of data for the 3D model poses the problem of insufficient network bandwidth and the long reception time.
[0097] In contrast, in this embodiment, the foreground model and background model that make up the three-dimensional model are distributed separately using different distribution methods. For example, by limiting the number of times that the background model, which is updated less frequently, is distributed, the network bandwidth during distribution can be reduced. This shortens the reception time on the terminal side.
[0098] Next, a configuration of a three-dimensional space recognition system 100 according to this embodiment will be described. Fig. 2 is a block diagram showing the configuration of the three-dimensional space recognition system 100. The three-dimensional space recognition system 100 includes a multi-viewpoint video imaging device 111, a control device 112, an event detection device 113, a calibration instruction device 114, a three-dimensional space reconstruction device 115, and a three-dimensional space recognition device 116.
[0099] FIG. 3 is a diagram showing an outline of the operation of the three-dimensional space recognition system 100.
[0100] The multi-viewpoint video imaging device 111 generates multi-viewpoint video by capturing images of the same space (S101).
[0101] Furthermore, camera calibration is performed to estimate the orientation (camera parameters) of each camera by manually or automatically detecting the correspondence between points in the shooting environment and points on the image, and between points in the images (S102).
[0102] The three-dimensional space reconstruction device 115 generates a three-dimensional model by performing three-dimensional space reconstruction that reconstructs the shooting space three-dimensionally using the multi-view video and the camera parameters (S103). For example, a foreground model and a background model are generated as the three-dimensional model.
[0103] Finally, the 3D space recognition device 116 performs 3D space recognition using the 3D model (S104). Specifically, the 3D space recognition device 116 performs tracking, scene analysis, and video rendering using the 3D model.
[0104] Next, a free viewpoint video generation system 101 including the three-dimensional space recognition system 100 will be described. FIG. 4 is a block diagram showing the configuration of the free viewpoint video generation system 101 according to this embodiment. In addition to the configuration of the three-dimensional space recognition system 100, the free viewpoint video generation system 101 includes a plurality of video display terminals 117, which are user terminals. Furthermore, the three-dimensional space reconstruction device 115 includes a foreground model generation unit 131 and a background model generation unit 132. The three-dimensional space recognition device 116 includes a viewpoint determination unit 141, a rendering unit 142, and a data transfer unit 143.
[0105] 5 is a diagram showing an outline of the operation of the free viewpoint video generation system 101. FIG.
[0106] First, the multi-viewpoint video imaging device 111 generates a multi-viewpoint video by performing multi-viewpoint shooting (S101). The multi-viewpoint video imaging device 111 includes a plurality of imaging devices 121. Each imaging device 121 includes a camera 122, a pan head 123, a memory 124, and a sensor 125.
[0107] The multi-viewpoint video imaging device 111 receives a signal to start or stop imaging from the control device 112, and starts or stops imaging in synchronization with the imaging devices 121 in response to the signal.
[0108] Each imaging device 121 captures an image with a camera 122 and simultaneously records a timestamp of when the image was captured. In addition, the imaging device 121 senses the imaging environment using a sensor 125 (a vibration sensor, an acceleration sensor, a geomagnetic sensor, or a microphone) simultaneously with capturing an image, and outputs the image, the timestamp, and the sensing data to the event detection device 113.
[0109] Furthermore, when the multi-viewpoint video imaging device 111 receives calibration instruction information from the calibration instruction device 114, it adjusts the imaging device 121 in accordance with the calibration instruction information, calibrates the camera 122, and outputs the camera parameters obtained by the calibration to the event detection device 113.
[0110] The memory 124 in each image capture device 121 temporarily stores images, time stamps, sensing data, and camera parameters, and stores shooting settings (frame rate, resolution, etc.).
[0111] Furthermore, camera calibration is performed at any timing (S102). Specifically, the event detection device 113 detects a calibration event from at least one of the video, timestamp, and sensing information obtained from the multi-viewpoint video imaging device 111, the 3D model obtained from the 3D space reconstruction device 115, the free-viewpoint video obtained from the rendering unit 142, terminal information obtained from the video display terminal 117, and control information obtained from the control device 112, and outputs calibration event information including the calibration event to the calibration instruction device 114. The calibration event information includes the calibration event, the importance of the calibration event, and information indicating the imaging device 121 to be calibrated.
[0112] A calibration event is a trigger for calibrating the image capture device 121. For example, the event detection device 113 outputs calibration event information when a shift in the camera 122 is detected, when a predetermined time has passed, when the accuracy of the camera calibration increases, when the accuracy of the model or the free viewpoint video decreases, when the free viewpoint video is not needed, when the video of a certain image capture device 121 cannot be used to generate the free viewpoint video, or when an instruction is received from a system administrator or a user.
[0113] Specifically, the event detection device 113 detects that the camera 122 has shifted when the sensing information exceeds a threshold, when the background area in the video changes by more than a threshold, or when cheers are heard. The predetermined time may be when play is interrupted, such as during halftime or the bottom of the fifth inning, when a certain amount of time has passed since the previous calibration, or when the system is started up. The accuracy of camera calibration increases when a certain number of feature points are extracted from the video, for example. The event detection device 113 also determines a deterioration in the accuracy of the model or free viewpoint video from distortions of the walls or ground in the model or free viewpoint video, etc.
[0114] A free viewpoint video is not needed when any of the video display terminals 117 are not in use, or when a scene is recognized from sound or video and identified as not being important. A video from a certain imaging device 121 cannot be used to generate a free viewpoint video when a sufficient communication bandwidth is not available and the video resolution or frame rate is low, when a synchronization error occurs, or when the area being photographed by the imaging device 121 is not attracting attention for reasons such as the absence of players.
[0115] The importance of a calibration event is calculated based on the calibration event itself or the data observed when the calibration event is detected. For example, a camera misalignment event is considered to be more important than other events. For example, the greater the camera misalignment, the higher the importance of the event.
[0116] The event detection device 113 may also send calibration event information to the video display terminal 117 to notify the user of the imaging device 121 that is currently being calibrated.
[0117] When the calibration instruction device 114 receives the calibration event information from the event detection device 113 , it generates calibration instruction information based on the calibration event information and outputs the generated calibration instruction information to the multi-viewpoint video imaging device 111 .
[0118] The calibration instruction information includes the cameras 122 to be calibrated, the order of the cameras 122 to be calibrated, control information for the pan head 123, zoom magnification change information for the cameras 122, and a calibration method. The control information for the pan head 123 indicates, for example, the amount of rotation of the pan head 123 to return the camera attitude that has shifted due to vibration or the like to its original attitude. The camera zoom magnification change information indicates, for example, the amount of zoom out required to cover the shooting area of the camera 122 that has shifted due to vibration or the like.
[0119] Calibration methods include a method of associating the three-dimensional coordinates of a specific point, line, or surface with the two-dimensional coordinates in the image, and a method of associating the two-dimensional coordinates of a specific point, line, or surface in the image between two or more images. These correspondences can be performed manually, automatically, or both. Furthermore, the accuracy of camera calibration can be improved by using two or more points, lines, or surfaces with known distances, or one or more stereo cameras.
[0120] Next, three-dimensional space reconstruction device 115 performs three-dimensional space reconstruction using the multi-view video (S103). Specifically, event detection device 113 detects a model generation event from at least one of the video, timestamp, and sensing information obtained from multi-view video imaging device 111, terminal information obtained from video display terminal 117, and control information obtained from the control device, and outputs model generation information including the model generation event to three-dimensional space reconstruction device 115.
[0121] The model generation information includes a model generation event and imaging device information. The imaging device information includes video, background image, camera parameters, reliability of the camera parameters, and camera calibration status. A model generation event is a trigger for generating a 3D model of the imaging environment. Specifically, the event detection device 113 outputs the model generation information when a certain number of cameras or more have been calibrated, when a predetermined time has arrived, when free viewpoint video is needed, etc.
[0122] The predetermined time may be when a game is being played or when a certain time has elapsed since the previous model generation. Free viewpoint video is required when the video display terminal 117 is in use, when a scene is recognized from sound or video and identified as an important scene, or when an instruction from a system administrator or a viewing request from a user is received. The reliability of camera parameters is determined from the results of camera calibration, the time when camera calibration was performed, the video, or sensing information. For example, the lower the reprojection error during camera calibration, the higher the reliability is set. Furthermore, the more recently a camera has been calibrated, the higher the reliability is set. Furthermore, the more feature points used in camera calibration, the higher the reliability is set.
[0123] The three-dimensional space reconstruction device 115 generates a three-dimensional model of the shooting environment using the model generation information obtained from the event detection device 113 and stores the generated three-dimensional model. When generating the model, the three-dimensional space reconstruction device 115 preferentially uses images captured by a calibrated camera with high reliability, based on the camera's calibration status and the reliability of the camera parameters. Furthermore, when the three-dimensional space reconstruction device 115 completes generation of the three-dimensional model of the shooting environment, it outputs model generation completion information to the event detection device 113.
[0124] The three-dimensional space reconstruction device 115 outputs a three-dimensional model of the shooting environment to the rendering unit 142 when the three-dimensional space recognition device 116, which is a free viewpoint video generation device, generates a free viewpoint video.
[0125] The foreground model generation unit 131 generates a foreground model, which is a model of a foreground object such as a person or a ball, whose movement changes over time (changes significantly). The background model generation unit 132 generates a background model, which is a model of a background object such as a venue or a goal, whose movement does not change over time (changes little). Hereinafter, a three-dimensional model refers to a model including a foreground model and a background model.
[0126] Foreground model generation unit 131 generates a foreground model in accordance with the frame rate at which imaging device 121 records. For example, if the recording frame rate is 30 frames per second, foreground model generation unit 131 generates a foreground model every 1 / 30 seconds.
[0127] The background model generation unit 132 generates a background model using a background image that does not include a foreground whose movement changes over time, such as a person or a ball. The background model generation unit 132 may reuse a background model that has been generated once within a certain period of time. Alternatively, the background model generation unit 132 may generate a new background model after a certain period of time has elapsed, thereby updating the background model. This reduces the amount of processing required to generate a background model with little movement, thereby reducing CPU usage and memory usage.
[0128] A method for generating a background model and a foreground model will be described below with reference to Figure 7.
[0129] First, the background model generation unit 132 generates a background model (S111). For example, the multiple imaging devices 121 included in the multi-view video imaging device 111 capture images of the background to generate a background image and record the background image. The background model generation unit 132 generates a background model using this background image. As a background model generation method, for example, a method such as a multi-view stereo method can be used, in which the depth of each pixel of an object included in a background image is calculated from multiple stereo camera pairs to identify the three-dimensional position of the object included in the background. Alternatively, the background model generation unit 132 may use a method of extracting features of the background image and identifying the three-dimensional position of the feature of the background image based on the principle of triangulation from the matching results of the features between the cameras. Besides this, any method for calculating a three-dimensional model of an object included in the background may be used.
[0130] Alternatively, some or all of the background model may be created manually. For example, a three-dimensional model of a goal or other object whose shape is determined by the event may be generated in advance using CG or the like. In other words, the background model generation unit 132 may acquire a predetermined and generated background model.
[0131] Furthermore, the background model generation unit 132 may generate a background image using a plurality of captured images including a foreground and a background. For example, the background model generation unit 132 may calculate a background image using an average image of a plurality of captured images. This allows a background image to be generated even in a situation where a background image not including a foreground cannot be captured in advance, thereby enabling generation of a background model.
[0132] Next, the multiple imaging devices 121 included in the multi-viewpoint video imaging device 111 capture images of people (foreground) and the background to generate captured images, and record the captured images (S112).
[0133] Next, foreground model generation unit 131 generates a foreground model (S113). Specifically, foreground model generation unit 131 generates a background difference image by subtracting a background image from images captured from the same viewpoint by the same imaging device 121. Foreground model generation unit 131 generates a foreground model using background difference images from multiple viewpoints. For example, a method of identifying a three-dimensional model of a foreground object existing in space using multiple background difference images, such as a volume intersection method, can be used to generate the foreground model. Alternatively, foreground model generation unit 131 may use a method of extracting feature amounts of the foreground image (background difference image) and identifying the three-dimensional position of the feature amounts of the foreground image based on the principle of triangulation from the matching results of the feature amounts between the cameras. Any other method may be used as long as it calculates a three-dimensional model of an object included in the foreground.
[0134] In this way, a foreground model and a background model are generated.
[0135] Next, three-dimensional space recognition is performed using the three-dimensional model (S104A).First, viewpoint determination unit 141 determines a virtual viewpoint (S105).
[0136] Specifically, the event detection device 113 detects a free viewpoint generation event from model generation completion information obtained from the three-dimensional space reconstruction device 115, terminal information obtained from the video display terminal 117, and control information obtained from the control device 112, and outputs free viewpoint information including the free viewpoint generation event to the viewpoint determination unit 141.
[0137] The free viewpoint generation information includes a free viewpoint generation event, a requested viewpoint, and imaging device information. The requested viewpoint is a viewpoint desired by the user obtained from the video display terminal 117, or a viewpoint specified by a system administrator obtained from a control device, etc. The viewpoint may be a point in three-dimensional space or a line segment. The free viewpoint generation event is a trigger for generating a free viewpoint video of the imaging environment. Specifically, the event detection device 113 outputs the free viewpoint information when a three-dimensional model of the imaging environment is generated, or when a user requests or a system administrator instructs that a free viewpoint video of a time when a generated three-dimensional model exists is viewed or distributed.
[0138] The viewpoint determination unit 141 determines a viewpoint to use when generating a free viewpoint video based on the free viewpoint information obtained from the event detection device 113, and outputs this viewpoint information together with the free viewpoint information to the rendering unit 142. The viewpoint determination unit 141 determines the viewpoint based on a requested viewpoint. If there is no requested viewpoint, the viewpoint determination unit 141 may automatically detect a viewpoint from the video that allows a player to be viewed from the front, or may automatically detect a viewpoint where a calibrated and highly reliable imaging device 121 is nearby based on the reliability of the camera parameters or the calibration status of the camera.
[0139] Once the virtual viewpoint is set, the structure and distance information of the shooting environment seen from the virtual viewpoint are determined from the three-dimensional model (including the foreground model and the background model). The rendering unit 142 performs rendering using the three-dimensional model to generate a free viewpoint video, which is a video seen from the virtual viewpoint (S106).
[0140] Specifically, the rendering unit 142 uses the viewpoint information and free viewpoint information obtained from the viewpoint determination unit 141 and the three-dimensional model of the shooting environment obtained from the three-dimensional space reconstruction device 115 to generate an image of a viewpoint that conforms to the viewpoint information, and outputs the generated image to the data transfer unit 143 as a free viewpoint image.
[0141] That is, the rendering unit 142 generates the free viewpoint video by projecting a three-dimensional model onto the virtual viewpoint position indicated by the viewpoint information. In this case, the rendering unit 142, for example, preferentially acquires color and texture information on the video from a video captured by an imaging device 121 that is close to the virtual viewpoint position. However, if the imaging device 121 that is close is undergoing calibration or the reliability of its camera parameters is low, the rendering unit 142 may preferentially acquire color information from a video captured by an imaging device 121 other than the imaging device 121 that is close to the virtual viewpoint position. Furthermore, if the imaging device 121 that is close to the virtual viewpoint position is undergoing calibration or the reliability of its camera parameters is low, the rendering unit 142 may blur the video or increase the playback speed to make the degradation of image quality less noticeable to the user. In this way, the rendering unit 142 does not necessarily need to preferentially acquire color and texture information from a video captured by an imaging device 121 that is close to the virtual viewpoint position, and any method may be used to acquire color and texture on the video. Furthermore, color information may be added to the three-dimensional model itself in advance.
[0142] Next, the data transfer unit 143 distributes the free viewpoint video obtained from the rendering unit 142 to the video display terminal 117 (S107). The data transfer unit 143 may distribute different free viewpoint video to each video display terminal 117 based on the viewpoint requested by each user, or may distribute the same free viewpoint video generated based on the viewpoint specified by the system administrator or the viewpoint automatically determined by the viewpoint determination unit 141 to multiple video display terminals 117. The data transfer unit 143 may also compress the free viewpoint video and distribute the compressed free viewpoint video.
[0143] Next, each video display terminal 117 displays the distributed free viewpoint video (S108). Here, the video display terminal 117 is equipped with a display, a wireless communication device, and a user input interface. The user uses the video display terminal 117 to send a viewing request to the event detection device 113, requesting to view an arbitrary area of the shooting environment at an arbitrary time from an arbitrary viewpoint. The video display terminal 117 receives the free viewpoint video based on the viewing request from the data transfer unit 143 and displays it to the user.
[0144] Furthermore, the video display terminal 117 receives calibration event information obtained from the event detection device 113 and highlights the camera being calibrated on the display, thereby informing the user that free viewpoint video cannot be generated from a viewpoint close to this imaging device, or that the image quality will be poor.
[0145] Furthermore, the system administrator sends a shooting start or stop signal from the control device 112 to the multi-viewpoint video imaging device 111, causing the multi-viewpoint video imaging device 111 to start or stop synchronous shooting.
[0146] Furthermore, if the system administrator determines that the camera needs to be calibrated, the system administrator can send control information from the control device 112 to the event detection device 113 and calibrate any camera.
[0147] Furthermore, if the system administrator determines that a three-dimensional model of the imaging environment is necessary, the control device 112 can send control information to the event detection device 113, and a three-dimensional model of the imaging environment at any time can be generated using any imaging device 121.
[0148] Furthermore, if the system administrator determines that free viewpoint video is necessary, the control device 112 can send control information to the event detection device 113, which can then generate free viewpoint video at any time and distribute it to the video display terminal 117.
[0149] (Embodiment 2) The above-mentioned free viewpoint video generation function may be used in a surveillance system, in which case an estimated appearance of a suspicious person seen from a viewpoint not captured by a real camera can be presented to security guards to keep them on guard.
[0150] Fig. 8 is a block diagram showing the configuration of a next-generation surveillance system 102 according to this embodiment. The next-generation surveillance system 102 shown in Fig. 8 differs from the free viewpoint video generation system 101 shown in Fig. 4 in that the configuration of a three-dimensional space recognition device 116A is different from that of the three-dimensional space recognition device 116. Furthermore, the next-generation surveillance system 102 includes a monitor 118A, a security guard 118B, and a video imaging device 118C instead of the video display terminal 117.
[0151] The three-dimensional space recognition device 116A includes a tracking unit 144, a scene analysis unit 145, and a data transfer unit 146.
[0152] Fig. 9 is a diagram showing an overview of the operation of the next-generation surveillance system 102. Fig. 10 is a flowchart showing the operation of the next-generation surveillance system 102. Note that the multi-viewpoint shooting (S101), camera calibration (S102), and three-dimensional space reconstruction (S103) are the same as those in Figs. 5 and 6.
[0153] Next, the three-dimensional space recognition device 116A performs three-dimensional space recognition using the three-dimensional model (S104B). Specifically, the tracking unit 144 tracks the person in the three-dimensional space (S105B). The tracking unit 144 also automatically extracts an image in which the person appears.
[0154] Furthermore, the scene analysis unit 145 performs scene analysis (S106B). Specifically, the scene analysis unit 145 performs situation recognition and abnormality detection for people or scenes from three-dimensional space or multi-viewpoint video.
[0155] Next, data transfer unit 146 transfers the result of the three-dimensional space recognition to a terminal or the like carried by monitor 118A or security guard 118B, or to video imaging device 118C (S107B).The result of the three-dimensional space recognition is then displayed on a terminal or the like carried by monitor 118A or security guard 118B, or on a display unit or the like of video imaging device 118C (S108B).
[0156] The above operation will be described in detail below. As with the generation of a free viewpoint video, the scene analysis unit 145 and tracking unit 144 calculate the structure of each subject in the shooting area as seen from a virtual viewpoint and the distance from the virtual viewpoint based on the three-dimensional model generated by the three-dimensional space reconstruction device 115. The scene analysis unit 145 and tracking unit 144 can also preferentially obtain the color and texture of each subject from the video of the imaging device 121 that is close to the virtual viewpoint and use the obtained information.
[0157] Scene analysis using two-dimensional video is performed by analyzing video showing the state of each subject in the shooting area, for example, a person or object, at a certain moment, either by software or by a person viewing the video on a screen. By performing this scene analysis based on three-dimensional model data, the scene analysis unit 145 can observe the three-dimensional posture of a person or the three-dimensional shape of an object in the shooting area, enabling more accurate situation recognition and prediction than when using two-dimensional video.
[0158] In tracking using two-dimensional video, for example, first, a subject within a captured area is identified by scene analysis of the video captured by the imaging device 121. Then, the same subject identified in the video captured by the imaging device 121 at different moments is associated by software or manually. Then, tracking is performed by identifying and associating the subject along a time axis. However, for example, in two-dimensional video captured by the imaging device 121, a subject of interest may be temporarily hidden by another subject, making it impossible to continue identifying that subject. Even in such a case, by using a three-dimensional model, it is possible to continue identifying the subject using three-dimensional position information or three-dimensional shape information of each subject.
[0159] The next-generation surveillance system 102 utilizes scene analysis and tracking functions using such 3D models. This enables early detection of suspicious scenes and improved detection accuracy. Furthermore, even in places where the number of cameras that can be installed is limited, security can be strengthened compared to when 2D images are used.
[0160] The scene analysis unit 145 analyzes the data of the three-dimensional model to, for example, identify the subject. The results of the analysis may be passed to the tracking unit 144 or may be displayed on a display of a terminal or the like together with the free viewpoint video. Data on the analysis results of the free viewpoint video may be stored in a storage device included in the terminal or an external storage device. Depending on the analysis results, the scene analysis unit 145 may request the user via the terminal to determine a virtual viewpoint at another time or another position.
[0161] The tracking unit 144 tracks a specific subject based on the three-dimensional model data. The tracking results may be displayed on a display of a terminal or the like together with the free viewpoint video. Furthermore, for example, if it is not possible to track a specific subject, the tracking unit 144 may request the user via the terminal to determine a virtual viewpoint at another time or another position.
[0162] (Embodiment 3) In this embodiment, a modified example of free viewpoint video generation system 101 according to Embodiment 1 will be described. Fig. 11 is a block diagram showing the configuration of free viewpoint video generation system 103 according to this embodiment. Free viewpoint video generation system 103 shown in Fig. 11 differs from free viewpoint video generation system 101 shown in Fig. 4 in that viewpoint determination unit 151 and rendering unit 152 are provided in video display terminal 117A.
[0163] Data transfer device 119 delivers the three-dimensional models (foreground model and background model) generated by three-dimensional space reconstruction device 115 to video display terminal 117A. Note that data transfer device 119 may also transmit the captured video and camera parameters obtained by multi-viewpoint video imaging device 111 to video display terminal 117A. Furthermore, when generating the three-dimensional model, three-dimensional space reconstruction device 115 may add color information to the three-dimensional model using the captured video or the like, and data transfer device 119 may deliver the three-dimensional model with the added color information to video display terminal 117A. Furthermore, in this case, data transfer device 119 does not need to deliver the captured video to video display terminal 117A.
[0164] Video display terminal 117A is equipped with a display, a wireless communication device, and a user input interface. Using video display terminal 117A, a user sends a viewing request to view an arbitrary area of the shooting environment at an arbitrary time to event detection device 113, and receives a 3D model, shot video, and camera parameters based on the viewing request from data transfer device 119. Then, video display terminal 117A uses the viewpoint information specified by the user and the received 3D model to generate a video from a viewpoint in accordance with the viewpoint information, and outputs the generated video to the display as a free viewpoint video.
[0165] 12 is a flowchart showing the operation of the free viewpoint video generation system 103. Note that steps S101 and S103 are the same as the processes in the first embodiment shown in FIG.
[0166] Next, data transfer device 119 distributes the three-dimensional models (foreground model and background model) generated by three-dimensional space reconstruction device 115 to video display terminal 117A (S107C). At this time, data transfer device 119 distributes the foreground model and the background model using different distribution methods.
[0167] For example, when distributing a three-dimensional model to the video display terminal 117A, the data transfer device 119 distributes the foreground model and the background model separately. At that time, the data transfer device 119 adds, for example, a flag or an identifier to distinguish whether each model is a foreground model or a background model to header information or the like included in the distribution data.
[0168] For example, the delivery cycle of the foreground model and the background model may be different. Furthermore, the delivery cycle of the foreground model may be shorter than the delivery cycle of the background model. For example, if the recording frame rate of the imaging device 121 is 30 frames / second, the data transfer device 119 delivers the foreground model at 30 models / second in accordance with the recording frame rate of the imaging device 121. Furthermore, the data transfer device 119 delivers, for example, one model as the background model.
[0169] Furthermore, when distributing a foreground model, data transfer device 119 may generate a differential model that is the difference between the foreground model at the current time and the foreground model at the previous time, and distribute the generated differential model. Furthermore, data transfer device 119 may predict the movement of the foreground model to generate a predicted model from the foreground model at the previous time, generate a differential model that is the difference between the foreground model at the current time and the predicted model, and distribute the generated differential model and motion information indicating the result of the motion prediction. This reduces the amount of information about the foreground model, thereby suppressing network bandwidth. Furthermore, data transfer device 119 may compress the amount of information in the transmission data by performing variable-length coding or arithmetic coding on the differential model and the motion information.
[0170] Furthermore, when distributing background models, the data transfer device 119 may distribute one background model when the user starts viewing. Alternatively, the data transfer device 119 may transmit background models at predetermined intervals. In this case, the data transfer device 119 may generate a differential model that is the difference between the current background model and the previously distributed background model, and transmit the generated differential model. This makes it possible to reduce the amount of information of the background model to be distributed, thereby suppressing network bandwidth.
[0171] Furthermore, data transfer device 119 may transmit both the foreground model and the background model at the random access point, which allows video display terminal 117A to generate free viewpoint video using appropriate foreground and background models whenever the user switches the time they wish to view.
[0172] 13 is a diagram showing an example of distribution of foreground and background models when one background model is distributed when a user starts viewing. As shown in FIG. 13, data transfer device 119 distributes one background model when a user starts viewing. Video display terminal 117A generates a free viewpoint video using the background model and a foreground model received at each time.
[0173] 14 is a diagram showing an example of distribution of foreground and background models when background models are distributed at regular intervals. As shown in FIG. 14, data transfer device 119 distributes background models at predetermined regular intervals. Here, the regular intervals are longer than the distribution intervals of the foreground models. Video display terminal 117A generates a free viewpoint video using the background model received immediately before and the foreground model received at each time.
[0174] Furthermore, when encoding and distributing a foreground model and a background model, data transfer device 119 may switch the encoding method for each model. In other words, data transfer device 119 may use different encoding methods for the foreground model and the background model. For example, data transfer device 119 applies an encoding method to the foreground model that prioritizes low latency, with the aim of real-time playback on video display terminal 117A. Data transfer device 119 also applies an encoding method to the background model that prioritizes high efficiency in order to reduce the amount of information as much as possible. This allows for the selection of an appropriate encoding method according to the intended use of each model, thereby reducing the amount of data and improving the functionality of the system.
[0175] The data transfer device 119 may use a highly efficient encoding method for the foreground model and a less efficient encoding method for the background model. For example, since the background model is distributed less frequently, using a less efficient encoding method reduces the network load even if the data volume increases. On the other hand, using a less efficient encoding method with lighter processing can reduce the processing load on the background model on the server or terminal. Furthermore, the foreground model is updated more frequently. Therefore, even if the processing load on the server or terminal is high, the network load can be reduced by encoding the foreground model as efficiently as possible. Instead of using a less efficient encoding method, the data transfer device 119 may send the model as is without encoding it.
[0176] Furthermore, data transfer device 119 may distribute the foreground model and the background model using networks or protocols with different characteristics. For example, for the foreground model, data transfer device 119 uses a reliable, high-speed network with little packet loss and a low-latency distribution protocol such as UDP (User Datagram Protocol) for the purpose of real-time playback on video display terminal 117A. For the background model, data transfer device 119 uses a low-speed network and a highly error-resistant protocol such as TCP (Transmission Control Protocol) to ensure the transmission bandwidth for the foreground model and to reliably distribute the background model. Alternatively, low latency for the foreground model may be achieved by applying download distribution using HTTP (Hypertext Transfer Protocol) or the like to the background model and streaming distribution using RTP (Real-time Transport Protocol) or the like to the foreground model.
[0177] Furthermore, the data transfer device 119 may acquire information on the viewpoint position of the user viewing from the video display terminal 117A, and use that information to switch the three-dimensional model to be distributed. For example, the data transfer device 119 may distribute, with priority, foreground and background models necessary for generating the video seen from the viewpoint the user is viewing. Furthermore, the data transfer device 119 may distribute the foreground model necessary for generating the video seen from the viewpoint the user is viewing with high accuracy (high density), and distribute the other models with reduced model accuracy (density) by performing a thinning process or the like. This reduces the amount of data to be distributed. However, such switching does not need to be performed for the background model.
[0178] Furthermore, the data transfer device 119 may change the density or distribution cycle of the three-dimensional models to be distributed depending on the available network bandwidth. For example, the data transfer device 119 may decrease the density of the three-dimensional models or increase the distribution cycle as the network bandwidth becomes narrower. Furthermore, the video display terminal 117A may switch the rendering resolution depending on the density of the three-dimensional models distributed by the data transfer device 119. For example, when the network bandwidth is narrow, the data transfer device 119 distributes the three-dimensional models by decreasing the density of the models through thinning processing or the like. Furthermore, the video display terminal 117A displays the video by reducing the rendering resolution.
[0179] In addition, methods for reducing the density of three-dimensional models can include uniform thinning, or switching between thinning methods or whether or not to thin depending on the target object. For example, the data transfer device 119 distributes important subjects as dense three-dimensional models and other subjects as sparse three-dimensional models. This reduces the amount of data distributed while maintaining the image quality of important subjects. Furthermore, when the network bandwidth becomes narrow, the data transfer device 119 may reduce the temporal resolution of the distributed three-dimensional models, for example, by lengthening the distribution cycle of foreground models.
[0180] Referring again to FIG. 12, next, video display terminal 117A performs three-dimensional space recognition using the distributed three-dimensional model. First, viewpoint determination unit 151 determines a virtual viewpoint (S105C). Next, rendering unit 152 performs rendering using the three-dimensional model to generate a free viewpoint video, which is a video viewed from the virtual viewpoint (S106C). Note that these processes are the same as the processes in steps S105 and S106 in the first embodiment. Next, video display unit 153 displays the generated free viewpoint video (S108C).
[0181] When receiving the three-dimensional model from the data transfer device 119, the video display terminal 117A may receive the foreground model and the background model separately. In this case, the video display terminal 117A may analyze the header information or the like to acquire a flag or identifier for distinguishing whether each model is a foreground model or a background model.
[0182] The reception cycles of the foreground model and the background model may be different. Furthermore, the reception cycle of the foreground model may be shorter than the reception cycle of the background model. For example, if the recording frame rate of the imaging device 121 is 30 frames / second, the video display terminal 117A receives the foreground model at 30 frames / second in accordance with the recording frame rate of the imaging device 121. Furthermore, the video display terminal 117A receives one model as the background model.
[0183] When receiving the foreground model, video display terminal 117A may receive a differential model that is the difference between the foreground model at the current time and the foreground model at the previous time, and generate the foreground model at the current time by adding the foreground model at the previous time and the differential model. Furthermore, video display terminal 117A may receive the differential model and motion information indicating the result of motion prediction, generate a prediction model from the received motion information and the foreground model at the previous time, and generate the foreground model at the current time by adding the differential model and the prediction model. This reduces the amount of information on the foreground model received, thereby suppressing network bandwidth. Furthermore, if the differential model and motion information are compressed by variable-length coding or arithmetic coding, video display terminal 117A may decode the differential model and motion information by variable-length decoding or arithmetic decoding the received data.
[0184] Furthermore, when receiving a background model, the video display terminal 117A may receive one background model when the user starts viewing, and reuse that one background model at all times. Alternatively, the video display terminal 117A may receive a background model at predetermined intervals. In this case, the video display terminal 117 may receive a differential model that is the difference between the previously received background model and the current background model, and generate the current background model by adding the previous background model and the differential model. This reduces the amount of information on the background model received, thereby suppressing network bandwidth.
[0185] Furthermore, video display terminal 117A may receive both the foreground model and the background model at the random access point, which allows video display terminal 117A to generate free viewpoint video using appropriate foreground and background models whenever the user switches the time they wish to view.
[0186] Furthermore, if the video display terminal 117A is unable to receive a three-dimensional model due to a network error or the like, it may perform rendering processing using a three-dimensional model that has already been received. For example, if the video display terminal 117A is unable to receive a foreground model, it may generate a predicted model by predicting movement from the foreground model that has already been received, and use the generated predicted model as the foreground model at the current time. Furthermore, if the video display terminal 117A is unable to receive a background model, it may use a background model that has already been received, or it may use a CG model. Furthermore, if the video display terminal 117A is unable to receive a background model or a foreground model, it may use a model prepared in advance, such as a CG image, or a rendered image. In this way, even if the video display terminal 117A is unable to receive a three-dimensional model, the video display terminal 117A can provide a rendered image to the user.
[0187] In addition, the data transfer device 119 may distribute at least one of the following to the video display terminal 117A: camera parameters, captured images obtained by the multi-viewpoint video imaging device 111, background images, background difference images, time information when each captured image or three-dimensional model is generated, viewpoint position information at the start of rendering, and time information for rendering.
[0188] Furthermore, if imaging device 121 is a fixed camera, data transfer device 119 may deliver the camera parameters to video display terminal 117A only when viewing starts. Furthermore, data transfer device 119 may deliver the camera parameters to video display terminal 117A at the timing when calibration is performed by calibration instruction device 114. Furthermore, if imaging device 121 is not fixed, data transfer device 119 may deliver the camera parameters to video display terminal 117A every time the camera parameters are updated.
[0189] Furthermore, the data transfer device 119 may encode and distribute the captured video, background image, or background difference image obtained by the multi-view video imaging device 111. This reduces the amount of data to be transmitted. For example, the data transfer device 119 may use a multi-view codec (MVC) such as H.264 or H.265 that utilizes the correlation between multi-view images. Furthermore, the data transfer device 119 may encode the video from each imaging device 121 independently using H.264 or H.265 before distributing it. This reduces the amount of data to be distributed to the video display terminal 117A.
[0190] The viewpoint position information at the start of rendering may be specified by the user via the video display terminal 117A at the start. Furthermore, the viewpoint determination unit 151 may switch the viewpoint position depending on the viewing style using the video display terminal 117A or the type of the video display terminal 117A. For example, in the case of viewing on a television, the viewpoint determination unit 151 determines, as the start viewpoint, a recommended viewpoint specified by the system, a viewpoint from the imaging device 121 close to the ball, a viewpoint from the imaging device 121 capturing the center of the field, or a viewpoint with a high viewership rating. In the case of viewing on a personal terminal such as a user's tablet or smartphone, the viewpoint determination unit 151 determines, as the start viewpoint, a viewpoint in which the user's favorite player is captured. In the case of viewing on a head-mounted display, the viewpoint determination unit 151 determines, as the start viewpoint, a recommended viewpoint for VR (Virtual Reality), for example, a viewpoint of a player on the field or a viewpoint from the bench.
[0191] (Fourth embodiment) In this embodiment, a modified example of next-generation surveillance system 102 according to embodiment 2 will be described. Fig. 15 is a block diagram showing the configuration of next-generation surveillance system 104 according to this embodiment. Next-generation surveillance system 104 shown in Fig. 15 differs from next-generation surveillance system 102 shown in Fig. 8 in that tracking unit 154 and scene analysis unit 155 are provided within video display terminal 117B.
[0192] 16 is a flowchart showing the operation of the next-generation monitoring system 104. Note that steps S101, S103, and S107C are the same as the processes in the third embodiment shown in FIG.
[0193] Next, video display terminal 117B performs three-dimensional space recognition using the three-dimensional model. Specifically, tracking unit 154 tracks the person in the three-dimensional space (S105D). Scene analysis unit 155 performs scene analysis (S106D). Then, video display terminal 117B displays the result of the three-dimensional space recognition (S108D). Note that these processes are similar to the processes of steps S105B, S106B, and S108B in the second embodiment.
[0194] (Embodiment 5) In the above embodiment, an example was described in which a three-dimensional model includes a foreground model and a background model, but the models included in a three-dimensional model do not have to be limited to two models, a foreground model and a background model.
[0195] Fig. 17 is a block diagram showing the configuration of free viewpoint video generation system 105 according to this embodiment. Free viewpoint video generation system 105 shown in Fig. 17 differs from free viewpoint video generation system 103 shown in Fig. 11 in the configuration of three-dimensional space reconstruction device 115A. This three-dimensional space reconstruction device 115A includes first model generation unit 133 that generates a first model, second model generation unit 134 that generates a second model, and third model generation unit 135 that generates a third model.
[0196] The three-dimensional space reconstruction device 115A generates three-dimensional models including a first model, a second model, and a third model. The data transfer device 119 delivers the first to third models to the video display terminal 117A separately using different delivery methods. The three-dimensional space reconstruction device 115A updates each model at different frequencies. The data transfer device 119 delivers each model to the video display terminal 117A at different intervals. For example, the first model is a foreground model, the second model is part of a background model, and the third model is a background model other than the second model. In this case, if the recording frame rate of the imaging device 121 is 30 frames per second, the data transfer device 119 delivers the first model at 30 models per second in accordance with the recording frame rate of the imaging device 121. Furthermore, the data transfer device 119 delivers the second model at one model per second and delivers one model as the third model at the start of viewing. This allows areas of the background model that have different update frequencies to be distributed as separate models at separate intervals, thereby reducing network bandwidth.
[0197] Furthermore, the data transfer device 119 may add an identifier to the 3D model to identify two or more models, allowing the video display terminal 117A to analyze the identifier and determine which model the received 3D model corresponds to.
[0198] Although an example in which three models are used has been described here, four or more models may be used.
[0199] Furthermore, when two models are used, the two models may be other than a foreground model and a background model. For example, the three-dimensional data may include a first model that is updated frequently and has a large amount of data, and a second model that is updated infrequently and has a small amount of data. Furthermore, the data transfer device 119 may deliver each model to the video display terminal 117A separately using different delivery methods. In this case, since each model has a different update frequency, the data transfer device 119 delivers each model to the video display terminal 117A at different intervals. For example, if the recording frame rate of the imaging device 121 is 30 frames per second, the data transfer device 119 delivers the first model at 30 frames per second in accordance with the recording frame rate of the imaging device 121. Furthermore, the data transfer device 119 delivers one model as the second model at the start of viewing. This allows three-dimensional models with different data amounts to be delivered at different intervals, thereby reducing network bandwidth.
[0200] The first model and the second model may be models with different levels of importance. Furthermore, the data transfer device 119 may deliver each model to the video display terminal 117A using different delivery methods. Since the importance of each model differs, the data transfer device 119 delivers each model to the video display terminal 117A at different intervals. For example, assume that the first model is a model with high importance and the second model is a model with low importance. In this case, if the recording frame rate of the imaging device 121 is 30 frames per second, the data transfer device 119 delivers the first model at 30 models per second and the second model at 15 models per second, in accordance with the recording frame rate of the imaging device 121. This allows for preferential delivery of three-dimensional models with high importance, thereby enabling appropriate rendering images to be provided to the user of the video display terminal 117A while suppressing network bandwidth.
[0201] The data transfer device 119 may also switch factors other than the distribution cycle depending on the importance. For example, the data transfer device 119 may switch the density of models depending on the priority. For example, when distributing three-dimensional models of a soccer game, the data transfer device 119 may determine that the three-dimensional models of players competing in front of one goal are of high importance, and that the three-dimensional model of a goalkeeper near the other goal is of low importance. The data transfer device 119 then distributes the three-dimensional models of the goalkeeper with a lower density than three-dimensional models with a higher importance. Note that the data transfer device 119 does not need to distribute three-dimensional models with a lower importance. The data transfer device 119 may also determine the level of importance based on, for example, whether the model being determined is close to a specific feature point or object, such as the ball, or whether it is close to a viewpoint from which most viewers view the model. For example, the importance of a model close to a specific feature point or object is set high, and the importance of a model close to a viewpoint from which most viewers view the model is set high.
[0202] Each model may be, for example, a collection of one or more objects (e.g., people, balls, cars, etc.) identified by object recognition, or may be a collection of regions or objects identified based on movement, such as background and foreground.
[0203] A similar modification can also be applied to next-generation monitoring system 104 described in embodiment 4. Fig. 18 is a block diagram showing the configuration of next-generation monitoring system 106 according to this embodiment. Next-generation monitoring system 106 shown in Fig. 18 differs from next-generation monitoring system 104 shown in Fig. 15 in the configuration of three-dimensional space reconstruction device 115A. The functions and the like of three-dimensional space reconstruction device 115A are the same as those in Fig. 17.
[0204] As described above in the first to fourth embodiments, a three-dimensional model distribution device (for example, data transfer device 119) distributes a first model (for example, a foreground model) that is a three-dimensional model of a target space in a target time period by a first distribution method, and distributes a second model (for example, a background model) that is a three-dimensional model of the target space in a target time period and changes less per hour than the first model by a second distribution method different from the first distribution method. In other words, the three-dimensional model distribution device transmits the foreground model and the background model separately.
[0205] For example, the transmission cycles of the first model and the second model are different. For example, the transmission cycle of the first distribution method is shorter than the transmission cycle of the second distribution method. Furthermore, the three-dimensional model distribution device transmits the first model at predetermined regular intervals. At this time, the three-dimensional model distribution device may transmit a differential model that is the difference between the first model at the current time and the first model at the previous time. Furthermore, the three-dimensional model distribution device may transmit movement information of the first model at the current time from the first model at the previous time.
[0206] For example, the three-dimensional model distribution device may transmit the second model at the start of viewing. Alternatively, the three-dimensional model distribution device may transmit the second model at predetermined intervals. Alternatively, the three-dimensional model distribution device may transmit a differential model that is the difference between the current second model and the second model transmitted previously. Alternatively, the three-dimensional model distribution device may transmit the second model at each random access point.
[0207] Furthermore, the three-dimensional model distribution device may transmit information such as a flag for distinguishing whether each model is a first model or a second model.
[0208] Furthermore, the three-dimensional model distribution device may transmit both the first model and the second model at the random access point.
[0209] Furthermore, the three-dimensional model distribution device may generate the first model and the second model using different methods. Specifically, the three-dimensional model distribution device generates the first model using a first generation method and generates the second model using a second generation method that has a different level of accuracy from the first generation method. For example, the three-dimensional model distribution device generates the first model using the first generation method and generates the second model using a second generation method that has higher accuracy than the first generation method. Alternatively, the three-dimensional model distribution device generates the first model using the first generation method and generates the second model using a second generation method that has lower accuracy than the first generation method. For example, when it is necessary to render a first model (foreground model) such as a player or a criminal with the highest possible image quality, the three-dimensional model distribution device generates the first model with high accuracy even if it increases the amount of data. On the other hand, the three-dimensional model distribution device reduces the amount of data by lowering the accuracy of second models of areas that are less important than the foreground, such as spectators or background images.
[0210] For example, the three-dimensional model distribution device generates a first model (foreground model) which is the difference between a third model, which is a three-dimensional model of multiple objects contained in a target space during a target time period, and a second model (background model), which is a three-dimensional model of some of the multiple objects contained in the target space during a target time period.
[0211] For example, the three-dimensional model distribution device generates a third multi-viewpoint image (background difference image) which is the difference between a first multi-viewpoint image (captured image) in which multiple objects included in a target space during a target time period are captured, and a second multi-viewpoint image (background image) in which some of the multiple objects are captured, and generates a first model (foreground model) using the third multi-viewpoint image (background difference image).
[0212] The three-dimensional model distribution device may also generate the first model by a volume intersection method using the second multi-view image (captured image) or the third multi-view image (background subtraction image), and generate the second model using the matching results of feature points between the cameras. This reduces the amount of processing required to generate the first model and improves the accuracy of the second model. The three-dimensional model distribution device may also create the second model manually.
[0213] The 3D model distribution device may distribute data other than the 3D model, including, for example, at least one of camera parameters, multi-viewpoint images, background difference images, time information, and starting viewpoint position.
[0214] Furthermore, the three-dimensional model distribution device may distribute the camera parameters of the fixed camera at the start of viewing, and distribute the camera parameters of the non-fixed camera each time the camera parameters change.
[0215] The viewpoint position at the start of viewing may be specified by the user when viewing begins. Alternatively, the viewpoint position at the start of viewing may be switched depending on the viewing style or type of device. For example, when viewing on a television, a recommended viewpoint, one of the fixed cameras (e.g., close to the ball or in the center of the field), or a viewpoint with a high viewership rating is selected. When viewing on a personal tablet device or smartphone, a viewpoint showing a favorite player is selected. When viewing on a head-mounted display, a recommended viewpoint for VR (e.g., a viewpoint on the field) is selected.
[0216] Furthermore, the first model and the second model are not limited to two models, a foreground model and a background model. Two or more models may be generated and distributed separately using different distribution methods. In this case, the update frequency differs for each model (even in the background, the update frequency differs depending on the area), so the three-dimensional model distribution device distributes each model at different intervals. Furthermore, the three-dimensional model distribution device adds identifiers to identify the two or more models.
[0217] Furthermore, the three-dimensional model distribution device switches the encoding method for each model.
[0218] For example, a first distribution method used for a first model uses a first encoding method. A second distribution method used for a second model uses a second encoding method. The first encoding method and the second encoding method differ in at least one of processing delay and encoding efficiency. For example, the second encoding method has a larger processing delay than the first encoding method. Alternatively, the second encoding method has a higher encoding efficiency than the first encoding method. Alternatively, the second encoding method has a lower encoding efficiency than the first encoding method.
[0219] The first distribution method may have lower latency than the second distribution method. For example, the three-dimensional model distribution device distributes the first model with low latency using a highly reliable line (e.g., using UDP). The three-dimensional model distribution device distributes the second model over a low-speed line (e.g., using TCP). Alternatively, the three-dimensional model distribution device may distribute the second model by download (e.g., HTTP) and the first model by streaming (e.g., RTP).
[0220] Furthermore, the three-dimensional model receiving device (e.g., video display terminal 117A) may use a three-dimensional model that has already been received if it is unable to receive the three-dimensional model due to a network error, etc. For example, if the three-dimensional model receiving device is unable to receive the first model, it generates a prediction model by predicting the movement from the first model that has already been received, and uses the generated prediction model as the first model at the current time.
[0221] Furthermore, if the 3D model receiving device is unable to receive the second model, it uses the second model that it has already received. Alternatively, the 3D model receiving device uses a model or rendering image prepared in advance, such as a CG model or CG image. In other words, the 3D model receiving device may perform different error concealment processes for the first model and the second model.
[0222] Furthermore, the three-dimensional model distribution device may preferentially distribute the first model and the second model required to generate the video from the viewpoint viewed by the user. For example, the three-dimensional model distribution device may distribute the first model required to generate the video from the viewpoint viewed by the user with high accuracy, and thin out the other first models. In other words, the terminal to which the first and second models are distributed (e.g., video display terminal 117A) uses the first and second models to generate a free viewpoint video, which is a video viewed from the selected viewpoint. The three-dimensional model distribution device preferentially distributes the model required to generate the free viewpoint video from among the first models.
[0223] The three-dimensional model distribution device may also change the quality of the three-dimensional model to be distributed depending on the available network bandwidth. For example, the three-dimensional model distribution device switches the density or rendering resolution of the three-dimensional model depending on the network bandwidth. When the bandwidth is tight, the three-dimensional model distribution device reduces the density of the three-dimensional model and the rendering resolution. The density of the three-dimensional model can be switched by uniformly thinning it out or by switching it depending on the target object. When the bandwidth is tight, the three-dimensional model distribution device performs processing to lower the temporal resolution of the three-dimensional model to be distributed, such as by lengthening the distribution cycle of the first model.
[0224] In the above description, an example has been described in which a three-dimensional model is generated using multi-viewpoint video captured by multi-viewpoint video imaging device 111, but the method of generating the three-dimensional model (foreground model and background model) is not limited to the above. For example, the three-dimensional model may be generated using information obtained by means other than a camera, such as LIDAR (Light Detection and Ranging) or TOF (Time of Flight). Furthermore, the multi-viewpoint video used to generate the three-dimensional model may be generated using this information.
[0225] The three-dimensional model may be in any form as long as it represents the three-dimensional position of the target object, for example, a point cloud, voxels, a mesh, polygons, or depth information.
[0226] (Sixth embodiment) In this embodiment, the three-dimensional space reconstruction device 115C generates one or more depth images from a three-dimensional model, compresses the generated depth images, and distributes them to the video display terminal 117C. The video display terminal 117C reconstructs the three-dimensional model from the received depth images. In this way, by efficiently compressing and distributing the depth images, it is possible to reduce the network bandwidth during distribution.
[0227] Fig. 19 is a block diagram showing the configuration of a free viewpoint video generation system 107 according to this embodiment. The free viewpoint video generation system 107 shown in Fig. 19 differs from the free viewpoint video generation system 105 shown in Fig. 17 in the configurations of a three-dimensional space reconstruction device 115C, a data transfer device 119C, and a video display terminal 117C. The three-dimensional space reconstruction device 115C includes a first depth image generation unit 136, a second depth image generation unit 137, and a third depth image generation unit 138 in addition to the configuration of the three-dimensional space reconstruction device 115A. The video display terminal 117C includes a model restoration unit 156 in addition to the configuration of the video display terminal 117A.
[0228] Instead of distributing the three-dimensional model, the three-dimensional space reconstruction device 115C generates one or more depth images (distance images) from the created three-dimensional model. The data transfer device 119C distributes the generated one or more depth images to the video display terminal 117C. In this case, the video display terminal 117C receives the one or more depth images, restores (generates) a three-dimensional model, and generates a rendering image using the restored three-dimensional model and the received captured image.
[0229] 20 is a flowchart showing the operation of the free viewpoint video generation system 107. The processing shown in FIG. 20 differs from the processing shown in FIG. 12 in that steps S121 to S123 are included instead of step S107C.
[0230] Steps S101 and S103 are the same as in the fifth embodiment, and the first model generation unit 133, the second model generation unit 134, and the third model generation unit 135 generate the first model, the second model, and the third model.
[0231] Next, the first depth image generation unit 136 generates one or more first depth images from the first model, the second depth image generation unit 137 generates one or more second depth images from the second model, and the third depth image generation unit 138 generates one or more third depth images from the third model (S121).
[0232] Next, the data transfer device 119C reduces the data amount of the generated first depth image, second depth image, and third depth image by performing two-dimensional image compression processing, etc. Then, the data transfer device 119C delivers the compressed first depth image, second depth image, and third depth image to the video display terminal 117C (S122).
[0233] Next, the model restoration unit 156 of the video display terminal 117C decodes the received first depth image, second depth image, and third depth image, and restores (generates) a first model using the first depth image, restores (generates) a second model using the second depth image, and restores (generates) a third model using the third depth image (S123).
[0234] Then, similar to the fifth embodiment, viewpoint determination unit 151 determines a viewpoint that the user wants to see (S105C). Rendering unit 152 uses the restored first to third models and the received captured image to generate a rendering image that is an image seen from the determined viewpoint (S106C). Video display unit 153 displays the rendering image (S108C).
[0235] In this way, the data transfer device 119C delivers depth images, which are two-dimensional images, instead of delivering three-dimensional models. This allows the data transfer device 119C to compress the depth images using a standard image compression method such as H.264 or H.265 before transmitting them, thereby reducing the amount of data transfer.
[0236] The first to third models may be configured from a point cloud, a mesh, or a polygon.
[0237] Also, here, as in embodiment 5, the case where first to third models are generated has been described as an example, but the same method can be applied when a foreground model and a background model are generated as in embodiments 1 to 4. The same method can also be applied when generating one three-dimensional model.
[0238] Furthermore, although the free viewpoint video generation system has been described as an example here, the same technique can also be applied to next-generation surveillance systems.
[0239] Furthermore, the three-dimensional space reconstruction device 115C may deliver camera parameters corresponding to the depth image in addition to the depth image. For example, the camera parameters are camera parameters at the viewpoint of the depth image. The camera parameters include internal parameters indicating the focal length of the camera, the center of the image, etc., and external parameters indicating the attitude (three-dimensional position and orientation) of the camera, etc. The three-dimensional space reconstruction device 115C generates a depth image from the three-dimensional model using the camera parameters.
[0240] The transmitted information is not limited to camera parameters, but may be any parameters used when generating a depth image from a three-dimensional model. In other words, the parameters may be parameters for projecting the three-dimensional model onto an imaging plane from a predetermined viewpoint (the viewpoint of the depth image). For example, the parameters may be a projection matrix calculated using the camera parameters.
[0241] Furthermore, the video display terminal 117C generates a three-dimensional model by projecting each pixel of one or more depth images into a three-dimensional space using the received camera parameters.
[0242] Furthermore, the three-dimensional space reconstruction device 115C may generate multiple depth images by projecting the three-dimensional model onto the same plane as the imaging surface of each imaging device 121. This makes the viewpoint positions of the captured images and the depth images the same. Therefore, for example, when the data transfer device 119C compresses images captured from multiple viewpoints by the multi-view video imaging device 111 using multi-view coding, which is an extended standard of H.264 or H.265, it can calculate disparity information between the captured images using the depth images and generate a predicted image between the viewpoints using the disparity information. This reduces the amount of code for the captured images.
[0243] Furthermore, the three-dimensional space reconstruction device 115C may generate a depth image by projecting a three-dimensional model onto the same plane as the imaging surface at a viewpoint different from the viewpoint of the imaging device 121. In other words, the viewpoint of the depth image may be different from the viewpoint of the captured image. For example, the three-dimensional space reconstruction device 115C generates a depth image by projecting a three-dimensional model at a viewpoint position from which the video display terminal 117C can easily restore the three-dimensional model. This enables the video display terminal 117C to generate a three-dimensional model with fewer errors. Furthermore, a viewpoint from which the video display terminal 117C can easily restore a three-dimensional model is, for example, a viewpoint from which more objects are captured.
[0244] The data transfer device 119C may also compress and transmit depth images. For example, the data transfer device 119C may compress (encode) depth images using a two-dimensional image compression method such as H.264 or H.265. The data transfer device 119C may also perform compression using a dependency between depth images of different viewpoints, such as a multi-view encoding method. For example, the data transfer device 119C may generate a predicted image between viewpoints using disparity information calculated from camera parameters.
[0245] Furthermore, the three-dimensional space reconstruction device 115C may determine the bit length representing the value of each pixel of the depth image so that the error between the three-dimensional model generated by the three-dimensional space reconstruction device 115C and the three-dimensional model restored by the video display terminal 117C is equal to or less than a certain value. For example, the three-dimensional space reconstruction device 115C may set the bit length of the depth image to a first bit length (e.g., 8 bits) when the distance to the subject is short, and to a second bit length (e.g., 16 bits) longer than the first bit length when the distance to the subject is long. Alternatively, the three-dimensional space reconstruction device 115C may adaptively switch the bit length depending on the distance to the subject. For example, the three-dimensional space reconstruction device 115C may shorten the bit length as the distance to the subject increases.
[0246] In this way, the 3D space reconstruction device 115C controls the bit length of the depth image to be distributed according to the error of the 3D model reconstructed by the video display terminal 117C. This makes it possible to reduce the network load by reducing the amount of information in the depth image to be distributed while keeping the error of the 3D model reconstructed by the video display terminal 117C within an acceptable range. For example, if the bit length of the depth image is set to 8 bits, the error of the 3D model reconstructed by the video display terminal 117C increases compared to when the bit length is set to 16 bits, but the network load for distribution can be reduced.
[0247] Furthermore, if color information is attached to each point cloud constituting the three-dimensional model, the three-dimensional space reconstruction device 115C may generate a depth image and a texture image including the color information by projecting each point cloud and the color information onto the same plane as the imaging plane of one or more viewpoints. In this case, the data transfer device 119C may compress and distribute the depth image and texture image. Furthermore, the video display terminal 117C decodes the compressed depth image and texture image, and generates a three-dimensional model and color information of the point cloud included in the three-dimensional model using the obtained one or more depth images and texture images. Then, the video display terminal 117C generates a rendering image using the generated three-dimensional model and color information.
[0248] The compression of the depth image and the texture image may be performed by the data transfer device 119C or by the three-dimensional space reconstruction device 115C.
[0249] The three-dimensional space reconstruction device 115C or the data transfer device 119C may distribute the above-mentioned background difference image generated by subtracting the background image from the captured image. In this case, the video display terminal 117C may generate a three-dimensional model using the background difference image, and generate a rendering image using the generated three-dimensional model.
[0250] The three-dimensional space reconstruction device 115C or the data transfer device 119C may distribute position information indicating the position of each model in three-dimensional space. This allows the video display terminal 117C to easily integrate each model after generating it using the received position information. For example, the three-dimensional space reconstruction device 115C calculates the position information of each model by detecting a point cloud or the like in three-dimensional space when generating the model. Alternatively, the three-dimensional space reconstruction device 115C may detect a specific subject, such as a player, in advance in two-dimensional captured images and identify the three-dimensional position of the subject (model) using multiple captured images and subject detection information.
[0251] A depth image is two-dimensional image information that represents the distance from a certain viewpoint to a subject, and each pixel of the depth image stores a value that represents distance information to the point cloud of a three-dimensional model projected onto that pixel. Note that the information representing depth does not necessarily have to be an image, and can be anything that represents distance information to each point cloud that makes up the three-dimensional model.
[0252] In the above description, the three-dimensional space reconstruction device 115C generates a three-dimensional model from a background subtraction image or the like, and then generates a depth image by projecting the three-dimensional model to each viewpoint. However, this is not necessarily limited to this. For example, the three-dimensional space reconstruction device 115C may generate a three-dimensional model from something other than an image using LIDAR or the like, and generate a depth image from the three-dimensional model. Furthermore, the three-dimensional space reconstruction device 115C may acquire a previously generated three-dimensional model from the outside, and generate a depth image from the acquired three-dimensional model.
[0253] The three-dimensional space reconstruction device 115C may also set the bit length of the depth image to a different value for each model. For example, the three-dimensional space reconstruction device 115C may set the bit length of the first depth image and the second depth image to different values. The data transfer device 119C may also distribute information indicating the bit lengths of the first depth image and the second depth image to the video display terminal 117C. For example, if the first model is a foreground model and the second model is a background model, the three-dimensional space reconstruction device 115C may set the bit length of the first depth image of the foreground model, which requires higher model accuracy, to 16 bits, and set the bit length of the second depth image of the background model, which can withstand lower model accuracy, to 8 bits. This reduces the amount of information in the distributed depth images, while preferentially allocating bit lengths to depth images of portions, such as the foreground model, that require high-precision model reconstruction on the video display terminal 117C.
[0254] Furthermore, data transfer device 119C may distribute depth images of models requiring high accuracy to video display terminal 117C, but may not distribute depth images of models not requiring high accuracy to video display terminal 117C. For example, data transfer device 119C distributes a first depth image of a foreground model to video display terminal 117C, but does not distribute a second depth image of a background model to video display terminal 117C. In this case, video display terminal 117C uses a background model prepared in advance. This reduces the amount of information of the distributed depth images and suppresses network load.
[0255] Alternatively, the video display terminal 117C may determine whether to use a 3D model reconstructed from the delivered depth image or a 3D model prepared in advance. For example, if the video display terminal 117C is a terminal with high processing capabilities, the video display terminal 117C reconstructs 3D models from the delivered foreground model depth image and background model depth image, respectively, and uses the obtained 3D models for rendering, thereby generating high-quality rendered images for both the foreground and background. On the other hand, if the video display terminal 117C is a terminal with low processing capabilities and needs to reduce power consumption, such as a smartphone terminal, the video display terminal 117C reconstructs the foreground model from the delivered depth image and uses a prepared background model as the background model without using the delivered depth image. This allows for the generation of a high-quality foreground rendered image while reducing the amount of processing. In this way, by switching the 3D model to be used depending on the processing capabilities of the video display terminal 117C, it is possible to balance the quality of the rendered image with the power consumption achieved by reducing the amount of processing.
[0256] A specific example of a method for generating and restoring a three-dimensional model will be described below. Figure 21 is a diagram for explaining the process of generating and restoring a background model as a three-dimensional model.
[0257] First, the three-dimensional space reconstruction device 115C generates a background model from a background image (S101, S103). Note that the details of this process are the same as, for example, step S111 shown in FIG.
[0258] Next, the three-dimensional space reconstruction device 115C generates a depth image of the viewpoint A from the point cloud of the background model (S121). Specifically, the three-dimensional space reconstruction device 115C calculates a projection matrix A using the camera parameters of the viewpoint A. Next, the three-dimensional space reconstruction device 115C creates a depth image (distance image) by projecting the point cloud of the background model onto the projection plane of the viewpoint A using the projection matrix A.
[0259] In this case, multiple point clouds may be projected onto the same pixel in the depth image. In this case, for example, the three-dimensional space reconstruction device 115C uses the value closest to the projection plane of viewpoint A as the pixel value of the depth image. This prevents the depth value of a subject that is hidden in the shadow of the subject from being mixed in with the depth image, thereby enabling a correct depth image to be generated.
[0260] The data transfer device 119C then distributes the generated depth images (S122). At this time, the data transfer device 119C reduces the amount of data by applying standard two-dimensional image compression such as H.264 or H.265 to the depth images. Alternatively, the data transfer device 119C may compress the depth images using a multi-view coding method that utilizes parallax between viewpoints.
[0261] The data transfer device 119C also distributes the camera parameters used in generating the depth image from the three-dimensional model together with the depth image. Note that the data transfer device 119C may distribute the projection matrix A calculated using the camera parameters instead of or in addition to the camera parameters.
[0262] Next, the video display terminal 117C restores a point cloud of the background model by projecting depth images from multiple viewpoints into three-dimensional space (S123). At this time, the video display terminal 117C checks whether there are any problems with the geometric positional relationships between each restored point cloud and each viewpoint, and may readjust the position of the point cloud as necessary. For example, the video display terminal 117C matches feature points using images between viewpoints, and adjusts the position of each point cloud so that the point clouds corresponding to each matched feature point coincide in three-dimensional space. This allows the video display terminal 117C to restore a three-dimensional model with high accuracy.
[0263] Although an example of generating and restoring a background model has been described here, a similar method can be applied to other models such as a foreground model.
[0264] Next, an example of a depth image will be described. FIG. 22 is a diagram showing an example of a depth image. Each pixel of a depth image represents distance information to a subject. For example, a depth image is expressed as an 8-bit monochrome image. In this case, the closer the distance to viewpoint A, the brighter value (value closer to 255) is assigned, and the farther the distance to viewpoint A, the darker value (value closer to 0) is assigned. In the example shown in FIG. 22, subject A is closer to viewpoint A and is therefore assigned a brighter value, and subject B is farther from viewpoint A and is therefore assigned a darker value. The background is further away than subject B, and is therefore assigned a darker value than subject B.
[0265] In addition, in the depth image, the farther the distance from viewpoint A, the brighter a value (a value closer to 255) may be assigned, and the closer the distance from viewpoint A, the darker a value (a value closer to 0) may be assigned. Furthermore, in the example shown in FIG. 22, distance information to the subject is expressed using a depth image, but the transmitted information is not necessarily limited to this and may be in any format as long as it can express the distance to the subject. For example, distance information to subjects A and B may be expressed using text information rather than an image. Furthermore, while the bit length of the depth image is 8 bits here, the bit length is not necessarily limited to this and values greater or smaller than 8 bits may be used. Using a value greater than 8 bits, for example, 16 bits, allows for more detailed reproduction of distance information to the subject, thereby improving the accuracy of restoring the 3D model in the video display terminal 117C. Therefore, the video display terminal 117C can reconstruct a 3D model similar to the 3D model generated by the 3D space reconstruction device 115C. On the other hand, as the amount of information of the depth image to be distributed increases, the network load also increases.
[0266] Conversely, if a value smaller than 8 bits, for example 4 bits, is used, the distance information to the subject becomes coarse, reducing the accuracy of the restoration of the 3D model on the video display terminal 117C. This increases the error between the restored 3D model and the 3D model generated by the 3D space reconstruction device 115C. On the other hand, the amount of information in the depth image to be distributed can be reduced, thereby suppressing the network load.
[0267] The three-dimensional space reconstruction device 115C may determine the bit length of such depth images based on whether a target application requires a highly accurate three-dimensional model on the video display terminal 117C. For example, if the target application does not care about the quality of the rendered image, the three-dimensional space reconstruction device 115C reduces the bit length of the depth images and prioritizes reducing the network load for distribution. On the other hand, if the target application cares about the quality of the image, the three-dimensional space reconstruction device 115C increases the bit length of the depth images and prioritizes increasing the network load for distribution, even if it increases the network load.
[0268] The three-dimensional space reconstruction device 115C may adaptively switch the bit length of the depth image depending on the load of the network over which the depth image is distributed. For example, when the network load is high, the three-dimensional space reconstruction device 115C reduces the network load while reducing the accuracy of the three-dimensional model by setting a small bit length. When the network load is low, the three-dimensional space reconstruction device 115C increases the bit length so that a more detailed three-dimensional model can be generated by the video display terminal 117C. In this case, the three-dimensional space reconstruction device 115C may store information about the bit length of the depth image in header information or the like and distribute it together with the depth image to the video display terminal 117C. This allows the video display terminal 117C to be notified of the bit length of the depth image. The three-dimensional space reconstruction device 115C may add information about the bit length of the depth image to each depth image, add it when the bit length changes, add it periodically, for example, at each random access point, add it only to the first depth image, or distribute it at other times.
[0269] Next, examples of pixel value allocation in a depth image will be described. Figures 23A, 23B, and 23C are diagrams showing first to third examples of pixel value allocation in a depth image.
[0270] In the first allocation method shown in FIG. 23A, values are linearly allocated to pixel values (depth pixel values) of a depth image having a bit length of 8 bits according to distance.
[0271] In the second allocation method shown in FIG. 23B, values are preferentially assigned to pixel values of an 8-bit depth image to objects that are close in distance. This improves the distance resolution of objects that are close in distance. Therefore, by using the second allocation method for a depth image of a foreground model, it is possible to improve the accuracy of the foreground model. The three-dimensional space reconstruction device 115C may distribute information about this second allocation method (i.e., information indicating which pixel value corresponds to which distance) in header information or the like. Alternatively, this information may be predetermined by a standard or the like, and the same information may be used on the transmitting and receiving sides.
[0272] In the third allocation method shown in FIG. 23C, values are preferentially assigned to pixel values of a depth image with a bit length of 8 bits to objects that are far away. This improves the distance resolution of objects that are far away. Therefore, by using the third allocation method for a depth image of a background model, it is possible to improve the accuracy of the background model. The three-dimensional space reconstruction device 115C may distribute information about this third allocation method (i.e., information indicating which pixel value corresponds to which distance) by including it in header information or the like. Alternatively, this information may be predetermined by a standard or the like, and the same information may be used on the transmitting and receiving sides.
[0273] Furthermore, three-dimensional space reconstruction device 115C may switch the allocation method for each model. For example, three-dimensional space reconstruction device 115C may apply the second allocation method to the foreground model and the third allocation method to the background model.
[0274] At this time, the three-dimensional space reconstruction device 115C may add which of the first to third allocation methods to use to header information or the like for each model to be distributed. Alternatively, which allocation method is applied to which model may be determined in advance by a standard or the like.
[0275] Furthermore, the three-dimensional space reconstruction device 115C may add information indicating which of a plurality of allocation methods defined in advance by a standard is to be used to the header information or the like.
[0276] As described above, the three-dimensional space reconstruction device 115C or the data transfer device 119C generates a depth image from a three-dimensional model, and delivers the depth image and information for restoring the three-dimensional model from the depth image to the video display terminal 117C.
[0277] Furthermore, the video display terminal 117C receives a depth image generated from the three-dimensional model and information for restoring the three-dimensional model from the depth image, and restores the three-dimensional model from the depth image using the information.
[0278] In this way, by distributing depth images generated from a three-dimensional model rather than distributing the three-dimensional model as is, the amount of data to be distributed can be reduced.
[0279] Furthermore, in generating the depth image, the three-dimensional space reconstruction device 115C generates the depth image by projecting a three-dimensional model onto an imaging plane of a predetermined viewpoint. For example, the information for restoring the three-dimensional model from the depth image includes parameters for projecting the three-dimensional model onto an imaging plane of a predetermined viewpoint.
[0280] For example, the information for restoring a three-dimensional model from a depth image is a camera parameter. That is, in generating a depth image, the three-dimensional space reconstruction device 115C generates the depth image by projecting a three-dimensional model onto an imaging plane of a predetermined viewpoint using a camera parameter of the viewpoint, and the information includes the camera parameter.
[0281] The information also includes parameters for projecting the three-dimensional model onto the imaging surface of the depth image, and in the restoration, the video display terminal 117C restores the three-dimensional model from the depth image using the parameters.
[0282] For example, the information includes camera parameters of the viewpoint of the depth image, and in the restoration, the video display terminal 117C restores a three-dimensional model from the depth image using the camera parameters.
[0283] Alternatively, the information for restoring the three-dimensional model from the depth image may be a projection matrix. That is, in generating the depth image, the three-dimensional space reconstruction device 115C calculates a projection matrix using camera parameters of a predetermined viewpoint, and generates the depth image by projecting the three-dimensional model onto an imaging plane of the viewpoint using the projection matrix, and the information includes the projection matrix.
[0284] Furthermore, the information includes a projection matrix, and in the restoration, the video display terminal 117C restores a three-dimensional model from the depth image using the projection matrix.
[0285] For example, the three-dimensional space reconstruction device 115C further compresses the depth image using a two-dimensional image compression method, and distributes the compressed depth image in the distribution.
[0286] Furthermore, the depth image is compressed using a two-dimensional image compression method, and the video display terminal 117C further decodes the compressed depth image.
[0287] This allows data to be compressed using a 2D image compression method when distributing 3D models. This eliminates the need to create a new compression method specifically for 3D models, making it easy to reduce the amount of data.
[0288] For example, in generating the depth images, the three-dimensional space reconstruction device 115C generates multiple depth images from different viewpoints from a three-dimensional model, and in compressing the multiple depth images, it compresses the multiple depth images using the relationship between the multiple depth images.
[0289] Furthermore, the video display terminal 117C receives a plurality of depth images in the receiving step, and decodes the plurality of depth images using the relationships between the plurality of depth images in the decoding step.
[0290] This allows the amount of data for multiple depth images to be further reduced by using, for example, a multi-view coding method in a two-dimensional image compression method.
[0291] For example, the three-dimensional space reconstruction device 115C further generates a three-dimensional model using multiple images captured by multiple imaging devices 121, delivers the multiple images to the video display terminal 117C, and the viewpoint of the depth image is the viewpoint of one of the multiple images.
[0292] Furthermore, the video display terminal 117C further receives a plurality of images, and generates a rendering image using the three-dimensional model and the plurality of images, and the viewpoint of the depth image is the viewpoint of one of the plurality of images.
[0293] In this way, by matching the viewpoint of the depth image with the viewpoint of the captured image, the three-dimensional space reconstruction device 115C can calculate disparity information between the captured images using the depth image and generate a predicted image between the viewpoints using the disparity information, for example, when compressing the captured images using multi-view coding. This allows the amount of code for the captured images to be reduced.
[0294] For example, the three-dimensional space reconstruction device 115C further determines the bit length of each pixel included in the depth image and distributes information indicating the bit length.
[0295] Moreover, the video display terminal 117C further receives information indicating the bit length of each pixel included in the depth image.
[0296] This allows the bit length to be switched depending on the subject or purpose of use, thereby making it possible to appropriately reduce the amount of data.
[0297] For example, in determining the bit length, the three-dimensional space reconstruction device 115C determines the bit length according to the distance to the subject.
[0298] For example, three-dimensional space reconstruction device 115C further determines the relationship between the pixel value shown in the depth image and the distance, and distributes information indicating the determined relationship to video display terminal 117C.
[0299] Moreover, the video display terminal 117C further receives information indicating the relationship between the pixel value shown in the depth image and the distance.
[0300] This allows the relationship between pixel values and distance to be switched depending on the subject or purpose of use, thereby improving the accuracy of the restored three-dimensional model.
[0301] For example, the three-dimensional model includes a first model (e.g., a foreground model) and a second model (e.g., a background model) that changes less per time than the first model. The depth image includes a first depth image and a second depth image. In generating the depth image, the three-dimensional space reconstruction device 115C generates a first depth image from the first model and generates a second depth image from the second model. In determining the relationship, the three-dimensional space reconstruction device 115C determines a first relationship between pixel values and distances indicated in the first depth image and a second relationship between pixel values and distances indicated in the second depth image. In the first relationship, the distance resolution in a first distance range (a region where the distance is close) is higher than the distance resolution in a second distance range (a region where the distance is far) that is farther than the first distance range (FIG. 23B). In the second relationship, the distance resolution in the first distance range (a region where the distance is close) is lower than the distance resolution in the second distance range (a region where the distance is far) (FIG. 23C).
[0302] For example, color information is added to the three-dimensional model. The three-dimensional space reconstruction device 115C further generates a texture image from the three-dimensional model, compresses the texture image using a two-dimensional image compression method, and distributes the compressed texture image.
[0303] In addition, the video display terminal 117C further receives a texture image compressed using a two-dimensional image compression method, decodes the compressed texture image, and in the restoration, restores a three-dimensional model with added color information using the decoded depth image and the decoded texture image.
[0304] (Embodiment 7) In this embodiment, a three-dimensional encoding device and a three-dimensional encoding method for encoding three-dimensional data, and a three-dimensional decoding device and a three-dimensional decoding method for decoding the encoded data into three-dimensional data will be described.
[0305] FIG. 24 is a diagram showing an outline of a three-dimensional data encoding method for encoding three-dimensional data.
[0306] In a three-dimensional encoding method for encoding three-dimensional data 200 such as a three-dimensional point group (three-dimensional point cloud or three-dimensional model), two-dimensional compression such as image encoding or video encoding is applied to a two-dimensional image 201 obtained by projecting the three-dimensional data 200 onto a two-dimensional plane. The two-dimensional image 201 obtained by projection contains texture information 202 indicating texture or color, and depth information (distance information) 203 indicating the distance to the three-dimensional point group in the projection direction.
[0307] Such a two-dimensional image obtained by projection may contain hole areas with no texture or depth information due to occlusion areas. A hole region refers to a pixel or a group of pixels to which no three-dimensional data is projected among the multiple pixels that make up a two-dimensional image obtained by projecting three-dimensional data onto a two-dimensional plane. Such hole regions cause discontinuities, sharp edges, and the like in the two-dimensional image obtained by projection. Two-dimensional images containing such discontinuities, sharp edges, and the like have a large amount of high spatial frequency components, resulting in high bit rates for encoding. Therefore, to improve encoding efficiency, it is necessary to minimize the sharp edges around the hole regions.
[0308] For example, it is conceivable to carry out correction to change the pixel values of the hole region so that sharp edges do not occur around the hole region. Next, correction to change the pixel values of the hole region will be described.
[0309] Fig. 25A is a diagram showing an example of a two-dimensional image including a hole region, and Fig. 25B is a diagram showing an example of a corrected image in which the hole region has been corrected.
[0310] 25A is a two-dimensional image obtained by projecting three-dimensional data onto a predetermined two-dimensional plane. Two-dimensional image 210 includes hole regions 214 and 215, which are invalid regions to which no three-dimensional data is projected. Two-dimensional image 210 also includes texture regions 211, 212, and 213, which are valid regions to which three-dimensional data is projected.
[0311] To improve the coding efficiency of such a two-dimensional image 210, it is necessary to appropriately fill the hole regions 214 and 215 with different pixel values, as described above. To improve coding efficiency, for example, it is necessary to minimize discontinuities in texture (or depth) between the hole regions 214 and 215 and the texture regions 211, 212, and 213. In a three-dimensional model coding method according to an embodiment of the present disclosure, the hole regions 214 and 215 are interpolated using pixel values of pixels in the texture regions 211 to 213, thereby reducing the differences between the hole regions 214 and 215 and the texture regions 211, 212, and 213 and performing correction to minimize sharp edges between these multiple regions 211 to 215. For example, at least one of linear interpolation and nonlinear interpolation can be used to correct the hole regions 214 and 215.
[0312] For such correction, a one-dimensional filter may be used for linear interpolation or non-linear interpolation, or a two-dimensional filter may be used.
[0313] In the correction, for example, a pixel value (first pixel value) on the boundary between one of the texture regions 211 to 213 in the two-dimensional image 210 and the hole regions 214 and 215 may be assigned (changed) to the pixel value of a pixel in the hole region 214 or 215, thereby interpolating the hole regions 214 and 215. In this manner, the correction corrects one or more pixels constituting the invalid region. In the correction, the invalid region may be corrected using a first pixel value of a first pixel in a first valid region, which is one of the valid regions adjacent to the invalid region. In addition, in the correction, the invalid region may be corrected using a second pixel value of a second pixel in a second valid region, which is an valid region on the opposite side of the first valid region in the two-dimensional image from the invalid region. For example, the first pixel may be a pixel adjacent to the invalid region in the first valid region. Similarly, the second pixel may be a pixel adjacent to the invalid region in the second valid region.
[0314] 25B, a two-dimensional image 220 is generated that has hole regions 224 and 225 in which pixel values have been changed to the pixel values of pixels (for example, pixel 226) in the texture regions 211 to 213. The two-dimensional image 220 is an example of a corrected image.
[0315] The pixel values to be assigned to the hole regions 214, 215 may be determined to be the pixel values of pixels on the boundary of a texture region among the multiple texture regions 211-213 that has the largest number of pixels on the boundary adjacent to the hole region 214, 215. For example, if a hole region is a region surrounded by multiple texture regions, the pixel values of the multiple pixels that make up the hole region may be replaced with the pixel values of pixels in the texture region that corresponds to the longest boundary line among multiple boundary lines that correspond to the multiple texture regions, respectively. Note that in order to interpolate the hole region, it is not necessary to apply the pixel values of pixels in the texture region adjacent to the hole region directly to the hole region, but it is also possible to apply the average or median value of the pixel values of the multiple pixels on the boundary with the hole region in the texture.
[0316] Furthermore, any method may be used, not limited to the above, as long as it is possible to set the value of the hole region to a value close to that of the texture region. For example, the average or median value of all the pixel values of the multiple pixels that make up the texture region may be determined as the pixel value of the multiple pixels that make up the hole region.
[0317] 26A and 26B are diagrams illustrating an example of hole region correction using linear interpolation. In FIGS. 26A and 26B, the vertical axis indicates pixel value, and the horizontal axis indicates pixel position. While FIGS. 26A and 26B illustrate a one-dimensional example, the method may also be applied to two dimensions. The pixel value may be, for example, a luminance value, a color difference value, an RGB value, or a depth value.
[0318] Linear interpolation is a correction method for correcting an invalid area using the first pixel and the second pixel of each of two texture areas A and B adjacent to the hole area. Here, the hole area is an example of an invalid area, texture area A is an example of a first valid area, and texture area B is an example of a second valid area.
[0319] In correction by linear interpolation, the hole region is corrected by using a first pixel value V1 of a first pixel P1 in texture region A and a second pixel value V2 of a second pixel P2 in texture region B to change the pixel values of each of the multiple pixels between the first pixel P1 and the second pixel P2 across the hole region to pixel values that satisfy a linear relationship between the first pixel value V1 and the second pixel value V2 in terms of the relationship between the positions and pixel values of each of the multiple pixels. In other words, in correction by linear interpolation, the multiple pixel values corresponding to each of the multiple pixels constituting the hole region between texture region A and texture region B are changed to pixel values specified by points on the line that correspond to the positions of each of the multiple pixels when a line is drawn connecting a first point indicated by the position of the first pixel P1 and the first pixel value V1 with a second point indicated by the position of the second pixel P2 and the second pixel value V2 in terms of the relationship between the positions and pixel values of each pixel.
[0320] In addition, in the case of correction by linear interpolation, as shown in Figure 26B, when the difference ΔV2 between the first pixel value V11 of the first pixel P11 and the second pixel value V12 of the second pixel P12 is larger than a predetermined value, even if the hole region is replaced with the first pixel value V11 and the second pixel value V12, discontinuity remains between the texture regions A and B and the hole region, and high spatial frequency components are included around the hole region, so coding efficiency may not be significantly improved. For this reason, correction by nonlinear interpolation, as shown in Figures 27A and 27B, for example, may be performed. This can reduce the discontinuity between the texture regions A and B and the hole region.
[0321] 27A and 27B are diagrams illustrating an example of hole region correction using nonlinear interpolation. In FIGS. 27A and 27B, the vertical axis represents pixel values, and the horizontal axis represents pixel positions. Although FIGS. 27A and 27B illustrate one-dimensional examples, they may also be applied to two dimensions.
[0322] In correction by nonlinear interpolation, the hole region is corrected by using a first pixel value V1 of a first pixel P1 in texture region A and a second pixel value V2 of a second pixel P2 in texture region B to change the pixel values of each of the multiple pixels between the first pixel P1 and the second pixel P2 across the hole region to pixel values that satisfy a relationship in which the pixel values change along a smooth curve from the first pixel value V1 to the second pixel value V2, depending on the relationship between the positions and pixel values of each of the multiple pixels. Here, the smooth curve is a curve that smoothly connects a first line, whose pixel value is the first pixel value V1 in texture region A, to the position of the first pixel P1, and smoothly connects a second line, whose pixel value is the second pixel value V2 in texture region B, to the position of the second pixel P2. For example, the smooth curve is a curve that has two inflection points and whose pixel value changes monotonically from the first pixel value V1 to the second pixel value V2 depending on the pixel position. For example, as shown in FIG. 27A, when the first pixel value V1 is greater than the second pixel value V2, the smooth curve is a curve in which the pixel value monotonically decreases depending on the position from the first pixel value V1 to the second pixel value V2.
[0323] In the case of the texture regions A and B and hole region corresponding to Fig. 26B, even if the difference ΔV2 between the first pixel value V11 of the first pixel P11 and the second pixel value V12 of the second pixel P12 is larger than a predetermined value as shown in Fig. 27B, correction by nonlinear interpolation replaces the pixel values of the multiple pixels constituting the hole region with pixel values corresponding to a smooth curve, thereby effectively reducing discontinuity between the texture regions A and B and the hole region, thereby improving coding efficiency.
[0324] Furthermore, the correction method is not limited to the above-described correction method explained with reference to FIGS. 26A to 27B, and other correction methods may also be used for correction.
[0325] 28A to 28F are diagrams showing other examples of correction.
[0326] As shown in Figure 28A, the hole area may be corrected by gradually changing the pixel values corresponding to the pixels that make up the hole area between texture area A and texture area B, from the first pixel value V21 of the first pixel P21 in texture area A to the second pixel value V22 of the second pixel P22 in texture area B.
[0327] 28B, multiple pixel values corresponding to multiple pixels constituting a hole region between texture region A and texture region B may be corrected by using pixel values on the boundary with the hole region in texture region A or texture region B. As a result, all of the pixel values of the multiple pixels constituting the hole region are unified to the pixel values of the pixels in texture region A or texture region B. In this correction, for example, the hole region is corrected by setting all pixel values of the multiple pixels constituting the hole region to the first pixel value V31 of the first pixel P31 on the boundary with the hole region in texture region A.
[0328] Here, Fig. 28C is a diagram showing an example of the correction of Fig. 28B expressed in a two-dimensional image. As shown in Fig. 28C(a), when first pixels P31a-P31e on the boundary between texture region A and the hole region have pixel values A-E, respectively, each pixel in the hole region is corrected by being assigned the first pixel value of the first pixel located at the same position as the pixel in the vertical direction, as shown in Fig. 28C(b). That is, in the correction, pixel value A of first pixel P31a is assigned to the pixel in the hole region located to the horizontal side of first pixel P31a. Similarly, in the correction, pixel values B-E of first pixels P31b-P31e are assigned to the pixel in the hole region located to the horizontal side of each of first pixels P31b-P31e, respectively.
[0329] 28C, the vertical direction may be interpreted as the horizontal direction, and the horizontal direction may be interpreted as the vertical direction. That is, each pixel in the hole region is corrected by being assigned the first pixel value of the first pixel located at the same position in the first direction, as shown in (b) of FIG. 28C. In other words, in the correction, the pixel value of the first pixel is assigned to a pixel in the hole region located on the second direction side of the first pixel, which is perpendicular to the first direction.
[0330] In addition, in the methods other than FIG. 28B, such as those in FIGS. 26A to 27B, 28A, and 28D to 28F, correction is performed by assigning the pixel value calculated in each method based on the first pixel value of the first pixel located at the same position in the vertical direction.
[0331] As in the correction shown in FIG. 28D, when the boundary between multiple coding blocks in two-dimensional coding is on a hole region, the region of the hole region that is closer to texture region A than the boundary between the coding blocks is corrected with the pixel values of the pixel values on the boundary between the hole region and texture region A. That is, in this correction, multiple first invalid pixels between the first pixel P41 and the boundary are changed to a first pixel value V41. Also, the region of the hole region that is closer to texture region B than the boundary between the coding blocks is corrected with the pixel values of the pixel values on the boundary between the hole region and texture region B. That is, in this correction, multiple second invalid pixels between the second pixel P42 and the boundary are changed to a second pixel value V42. In this way, the hole region may be corrected. Here, the coding block is, for example, a macroblock when the coding method is H.264, or a CTU (Coding Tree Unit) or CU (Coding Unit) when the coding method is H.265. .
[0332] As shown in (a) of Figure 28E, if there is no hole area between texture area A and texture area B and the difference in pixel values between texture area A and texture area B is greater than a predetermined value, it is possible to assume that there is a virtual hole area on the boundary between texture area A and texture area B, and to perform correction on the virtual hole area using the methods of Figures 26A to 28D. For example, (b) of Figure 28E is an example in which the correction by nonlinear interpolation described in Figures 27A and 27B is applied. As a result, even if there is no hole area, if there is a steep edge between texture area A and texture area B, the above correction can be performed to reduce the steep edge between texture areas A and B, and the amount of code can be effectively reduced.
[0333] FIG. 28F shows an example of a case similar to FIG. 28E(a) where edges are corrected using a different method.
[0334] FIG. 28F(a) is a diagram similar to FIG. 28E(a). Thus, when there is no hole region between texture region A and texture region B and the difference in pixel values between texture region A and texture region B is greater than a predetermined value, as shown in FIG. 28F(b), a virtual hole region may be generated by shifting texture region B away from texture region A, and the generated virtual hole region may be corrected using the methods of FIGS. 26A to 28D. For example, FIG. 28F(b) is an example in which the nonlinear interpolation correction described with reference to FIGS. 27A and 27B is applied. As a result, as in the case of FIG. 28E, even when there is no hole region, if there is a steep edge between texture region A and texture region B, the above correction can be performed to reduce the steep edge between texture regions A and B, thereby effectively reducing the amount of code.
[0335] In the correction, a smoothing filter such as a Gaussian filter or a median filter may be applied to the two-dimensional image obtained by projecting the three-dimensional model, regardless of whether it is a texture region or a hole region, and the texture region may be reassigned to the image after the filter application so that the values in the hole region approach the values in the texture region. This eliminates the need to identify the filter application region, and allows the values in the hole region to be corrected with a low amount of processing.
[0336] 28F, after creating a two-dimensional image obtained by projection, the pixels of each of the multiple texture regions may be shifted left and right, up and down within the two-dimensional image to generate hole regions. Also, hole regions may be created between multiple texture regions in the process of generating a two-dimensional image by projecting a three-dimensional point cloud onto a two-dimensional plane.
[0337] Fig. 29 is a block diagram showing an example of the functional configuration of a 3D model encoding device according to an embodiment. Fig. 30 is a block diagram showing an example of the functional configuration of a 3D model decoding device according to an embodiment. Fig. 31 is a flowchart showing an example of a 3D model encoding method performed by a 3D model encoding device according to an embodiment. Fig. 32 is a flowchart showing an example of a 3D model decoding method performed by a 3D model decoding device according to an embodiment.
[0338] The three-dimensional model encoding device 300 and the three-dimensional model encoding method will be described with reference to FIGS.
[0339] The 3D model coding device 300 includes a projection unit 301, a correction unit 302, and a coding unit 304. The 3D model coding device 300 may further include a generation unit 303.
[0340] First, the projection unit 301 generates a two-dimensional image by projecting a three-dimensional model onto at least one two-dimensional plane (S11). The generated two-dimensional image includes texture information and depth information.
[0341] The correction unit 302 uses the two-dimensional image to correct one or more pixels that constitute an invalid area (i.e., a hole area) included in the two-dimensional image where the three-dimensional model is not projected, thereby generating a corrected image (S12). As the correction, the correction unit 302 performs any of the corrections described above with reference to Figs. 25A to 28F.
[0342] Meanwhile, the generation unit 303 generates a two-dimensional binary map indicating whether each of a plurality of regions constituting the two-dimensional region corresponding to the two-dimensional image is an invalid region or a valid region (S13). The two-dimensional binary map enables the three-dimensional model decoding device 310, which has received the encoded data, to easily distinguish between invalid regions and valid regions in the two-dimensional image.
[0343] The encoding unit 304 generates an encoded stream as encoded data by two-dimensionally encoding the corrected image (S14). The encoding unit 304 may generate encoded data by further encoding the two-dimensional binary map together with the corrected image. Here, the encoding unit 304 may also encode projection information and parameters related to the projection when the two-dimensional image was generated, and generate encoded data.
[0344] Each of the projection unit 301, correction unit 302, generation unit 303, and encoding unit 304 may be realized by a processor and memory, or by a dedicated circuit. In other words, these processing units may be realized by software or hardware.
[0345] Next, the 3D model decoding device 310 and the 3D model decoding method will be described with reference to FIGS.
[0346] The three-dimensional model decoding device 310 includes a decoding unit 311 , a map reconstruction unit 312 , and a three-dimensional reconstruction unit 313 .
[0347] First, the decoding unit 311 acquires coded data and decodes the acquired coded data to acquire a corrected image and a two-dimensional binary map (S21). The coded data is coded data output by the three-dimensional model coding device 300. In other words, the coded data is coded data of a corrected image obtained by correcting a two-dimensional image generated by projecting a three-dimensional model onto at least one two-dimensional plane, and in which one or more pixels in an invalid area included in the two-dimensional image and onto which the three-dimensional model was not projected have been corrected.
[0348] The map reconstruction unit 312 reconstructs the decoded two-dimensional binary map to obtain the original map indicating valid pixels and invalid pixels (S22).
[0349] The three-dimensional reconstruction unit 313 reconstructs three-dimensional data from the corrected image using the projection information and the reconstructed two-dimensional binary map (S23). The three-dimensional reconstruction unit 313 obtains three-dimensional points by reprojecting valid pixels in the valid area shown in the two-dimensional binary map into three-dimensional space using the decoded depth information, and obtains the color of the three-dimensional points from the decoded texture information. Therefore, the three-dimensional reconstruction unit 313 does not reproject invalid pixels in the invalid area shown in the two-dimensional binary map. The depth information is a distance image indicating the distance corresponding to each pixel in the two-dimensional image. The texture information is a two-dimensional color image indicating the texture or color corresponding to each pixel in the two-dimensional image. In this way, the three-dimensional reconstruction unit 313 reconstructs three-dimensional points from the valid area in the corrected image, so that in the decoder, pixels in the valid area in the corrected image are not affected by pixels in the invalid area in the corrected image.
[0350] In the 3D model decoding method, the process of step S22 by the map reconstruction unit 312 does not necessarily have to be performed. In other words, the 3D model decoding device 310 does not necessarily have to include the map reconstruction unit 312.
[0351] Each of the decoding unit 311, the map reconstruction unit 312, and the 3D reconstruction unit 313 may be realized by a processor and memory, or by a dedicated circuit. In other words, these processing units may be realized by software or hardware.
[0352] In the 3D model coding device 300 according to this embodiment, the correction unit 302 performs two-dimensional coding on the corrected image generated by correcting the invalid area, thereby improving coding efficiency.
[0353] Furthermore, the correction unit 302 corrects the invalid area using a first pixel value of a first pixel in a first effective area, which is an effective area adjacent to the invalid area and onto which the 3D model is projected. This reduces the difference in pixel values between the first effective area and the invalid area, thereby effectively improving coding efficiency.
[0354] Furthermore, the correction unit 302 corrects the invalid area by further using a second pixel value of a second pixel in a second effective area, which is an effective area on the opposite side of the invalid area from the first effective area in the two-dimensional image. This reduces the difference in pixel values between the first effective area and the second effective area and the invalid area, thereby effectively improving coding efficiency.
[0355] Furthermore, the correction unit 302 performs linear interpolation on the invalid area, thereby reducing the processing load associated with determining pixel values for interpolation.
[0356] Furthermore, the correction unit 302 corrects invalid areas taking into consideration the boundaries of multiple blocks in two-dimensional coding, so that the processing load can be effectively reduced and coding efficiency can be effectively improved.
[0357] Furthermore, since the correction unit 302 performs correction using nonlinear interpolation, it can effectively reduce the difference in pixel values between the first and second effective areas and the invalid area, thereby improving coding efficiency.
[0358] Furthermore, the 3D model encoding device 300 generates a two-dimensional binary map, and outputs the encoded data obtained by encoding the two-dimensional binary map together with the corrected image. Therefore, during decoding by the 3D model decoding device 310, it is possible to decode only the valid regions out of the valid and invalid regions using the two-dimensional binary map, thereby reducing the amount of processing required during decoding.
[0359] The 3D model decoding device 310 according to this embodiment can reconstruct a 3D model by obtaining a small amount of coded data.
[0360] The 3D model encoding device 300 may add to the encoded data filter information (including filter application on / off information, filter type, filter coefficients, etc.) applied to a 2D image (projected 2D image) generated by projecting a 3D model onto a 2D plane. This allows the 3D model decoding device 310 to know the filter information applied to the decoded 2D image. After the 3D model decoding device 310 decodes the 3D model, this filter information can be reused when the 3D model is again encoded using the method described in this embodiment.
[0361] In this embodiment, the three-dimensional model encoding device 300 adds a two-dimensional binary map to the encoded data and transmits it to the three-dimensional model decoding device 310 to distinguish between valid and invalid areas of the decoded two-dimensional image. However, the two-dimensional binary map does not necessarily have to be added to the encoded data. For example, instead of generating a two-dimensional binary map, the three-dimensional model encoding device 300 may assign a value A, which is not used in texture areas, to hole areas. As a result, if each pixel value of the decoded two-dimensional image is value A, the three-dimensional model decoding device 310 can determine that the pixel is included in a hole area and decide that the pixel is an invalid pixel in an invalid area and not to reproject the pixel into three-dimensional space. In the case of an RGB color space, values such as (0,0,0) or (255,255,255) may be used as value A. This eliminates the need to add a two-dimensional binary map to the encoded data, thereby reducing the amount of code.
[0362] The above describes a three-dimensional encoding device and a three-dimensional encoding method for encoding three-dimensional data according to an embodiment of the present disclosure, as well as a three-dimensional decoding device and a three-dimensional decoding method for decoding the encoded data into three-dimensional data, but the present disclosure is not limited to these embodiments.
[0363] Furthermore, the processing units included in the three-dimensional encoding device and three-dimensional encoding method for encoding three-dimensional data according to the above-described embodiments, and the three-dimensional decoding device and three-dimensional decoding method for decoding the encoded data into three-dimensional data, are typically realized as LSIs, which are integrated circuits. These units may be individually integrated into single chips, or some or all of them may be integrated into a single chip.
[0364] Furthermore, the integration is not limited to LSI, but may be realized by dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays), which can be programmed after LSI fabrication, or reconfigurable processors, which allow the connections and settings of circuit cells within LSIs to be reconfigured, may also be used.
[0365] In each of the above embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0366] The present disclosure may also be realized as various methods executed by a three-dimensional encoding device and a three-dimensional encoding method that encode three-dimensional data, and a three-dimensional decoding device and a three-dimensional decoding method that decode the encoded data into three-dimensional data.
[0367] The division of functional blocks in the block diagram is an example, and multiple functional blocks may be realized as a single functional block, one functional block may be divided into multiple blocks, or some functions may be moved to another functional block.Furthermore, the functions of multiple functional blocks having similar functions may be processed in parallel or in time-sharing by a single piece of hardware or software.
[0368] The order in which the steps in the flowchart are executed is merely an example for specifically explaining the present disclosure, and other orders may be used. Some of the steps may be executed simultaneously (in parallel) with other steps.
[0369] The three-dimensional encoding device and three-dimensional encoding method for encoding three-dimensional data according to one or more aspects, and the three-dimensional decoding device and three-dimensional decoding method for decoding the encoded data into three-dimensional data have been described above based on the embodiments, but the present disclosure is not limited to these embodiments. As long as they do not deviate from the spirit of the present disclosure, various modifications conceivable by those skilled in the art to the present embodiments, and forms constructed by combining components of different embodiments, may also be included within the scope of one or more aspects. [Industrial Applicability]
[0370] The present disclosure can be applied to a three-dimensional encoding device and a three-dimensional encoding method that encode three-dimensional data, and a three-dimensional decoding device and a three-dimensional decoding method that decodes encoded data into three-dimensional data. [Explanation of symbols]
[0371] 100 Three-dimensional spatial recognition system 101, 103, 105, 107 Free viewpoint video generation system 102, 104, 106 Next-generation surveillance system 111 Multi-viewpoint video imaging device 112 Control device 113 Event detection device 114 Calibration instruction device 115, 115A, 115C Three-dimensional space reconstruction device 116, 116A 3D spatial recognition device 117, 117A, 117B, 117C video display terminals 118A Warden 118B Security guard 118C Video Imaging Device 119, 119C Data transfer device 121 Imaging device 122 Camera 123 Pan head 124 memory 125 sensors 131 Foreground model generation unit 132 Background model generation unit 133 First Model Generation Unit 134 Second Model Generation Unit 135 Third Model Generation Unit 136 First depth image generation unit 137 Second depth image generation unit 138 Third depth image generation unit 141, 151 Viewpoint determination unit 142, 152 Rendering section 143, 146 Data transfer section 144, 154 Tracking section 145, 155 Scene analysis section 153 Video display unit 156 Model Restoration Unit 200 3D data 201, 210, 220 2D images 202 Texture Information 203 Depth Information 211~213 Texture area 214, 215, 224, 225 hole areas 300 Three-dimensional model encoding device 301 Projection section 302 Correction Unit 303 Generation part 304 Encoding section 310 Three-dimensional model decoding device 311 Decoding Unit 312 Map Reconstruction Unit 313 Three-dimensional reconstruction unit
Claims
1. a projection unit that generates a two-dimensional image including one or more effective areas by projecting the three-dimensional model onto at least one or more two-dimensional planes; a correction unit that corrects the two-dimensional image to generate a corrected image; an encoding unit that generates encoded data by two-dimensionally encoding the corrected image, The correction of the two-dimensional image includes a shift process for shifting the position of the effective area on the two-dimensional image. Three-dimensional model coding device.
2. The shifting process includes generating a virtual area between adjacent effective areas by shifting the positions of the effective areas in a direction separating the effective areas.
2. The three-dimensional model encoding device according to claim 1.
3. The correction unit performs the shift process when there is no invalid area between the adjacent valid areas.
2. The three-dimensional model encoding device according to claim 1.
4. The correction unit performs the shift process when a difference in pixel values between the adjacent valid areas is greater than a predetermined value.
2. The three-dimensional model encoding device according to claim 1.
5. The correction unit corrects an invalid area adjacent to the valid area using pixel values of the valid area.
2. The three-dimensional model encoding device according to claim 1.
6. a decoding unit that acquires coded data in which a corrected image obtained by correcting a two-dimensional image including one or more effective areas in which a three-dimensional model is projected onto at least one two-dimensional plane is coded, and outputs the three-dimensional model obtained by decoding the acquired coded data, The correction of the two-dimensional image includes a shift process for shifting the position of the effective area on the two-dimensional image. 3D model decoding device.
7. generating a two-dimensional image including one or more effective areas projected from the three-dimensional model onto at least one or more two-dimensional planes; generating a corrected image by correcting the two-dimensional image; generating encoded data by two-dimensionally encoding the corrected image; The correction of the two-dimensional image includes a shift process for shifting the position of the effective area on the two-dimensional image. Three-dimensional model coding method.
8. Obtaining encoded data in which a corrected two-dimensional image including one or more effective areas in which the three-dimensional model is projected onto at least one two-dimensional plane is corrected, and a corrected image is encoded; outputting a three-dimensional model obtained by decoding the acquired encoded data; The correction of the two-dimensional image includes a shift process for shifting the position of the effective area on the two-dimensional image. Three-dimensional model decoding method.
Citation Information
Patent Citations
Method for transferring and displaying three-dimensional shape data
JP1997237354A
Digital image covering method, image processor and data recording medium
JP1998276432A
High efficiency encoder of video signal
JP1998336646A
Enhanced virtual environment
JP2006503379A
Method and device for estimating positional orientation
JP2010267232A