Information processor, information processing method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2023-07-25
- Publication Date
- 2026-07-29
AI Technical Summary
Existing technologies for generating and transmitting three-dimensional model data face challenges in efficiently reducing the amount of data to fit within transmission bands, leading to unintended image losses.
An information processing device that acquires transmission band information and movement information to determine an optimal data generation method, reducing data transmission by adjusting temporal and spatial resolutions based on the available band and movement speed of subjects and virtual viewpoints.
The method effectively reduces the amount of transmitted data while maintaining image quality, ensuring it fits within the available transmission band and minimizing data loss.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing technique for generating transmission data. [Background technology]
[0002] There is known a technology that generates a three-dimensional model from subject images captured from multiple viewpoints by a plurality of different imaging devices (hereinafter referred to as cameras) and generates an image seen from a virtual viewpoint (hereinafter referred to as a virtual viewpoint image) using the three-dimensional model. One example of a system form in which a user can view the virtual viewpoint image is a form in which a server generates transmission data such as a three-dimensional model or a virtual viewpoint image and sends the transmission data to a user terminal. However, in this system form, the amount of transmission data may not fit within the transmission band, resulting in unintended loss of images.
[0003] In response to this, Patent Document 1 discloses a technology that can efficiently reduce the amount of three-dimensional model data for reproducing a virtual viewpoint image. In the technology of Patent Document 1, model data of a plurality of hierarchical levels, each of which has a different amount of data, is generated for each subject, and transmission data is composed of model data of a hierarchical level determined for each subject based on attribute data that associates the attributes of the content with the required hierarchical level. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2019-54488 A Summary of the Invention [Problem to be solved by the invention]
[0005] According to the technology of Patent Document 1 mentioned above, it is possible to reduce the amount of transmission data so that it fits within the transmission band, but a technology that can reduce the amount of transmission data more efficiently is desired.
[0006] Therefore, an object of the present invention is to make it possible to reduce the amount of transmission data more efficiently. [Means for solving the problem]
[0007] The information processing device of the present invention is characterized by having a bandwidth information acquisition means for acquiring information on a transmission bandwidth available for data transmission, a model acquisition means for acquiring a three-dimensional model of a subject in three-dimensional space, a movement information acquisition means for acquiring movement information of at least one of the subject and a virtual viewpoint in the three-dimensional space, a determination means for determining a transmission data generation method based on the transmission bandwidth information and the movement information, and a data generation means for generating transmission data of the three-dimensional model or transmission data of a two-dimensional image generated from the three-dimensional model based on the transmission data generation method. Effect of the Invention
[0008] According to the present invention, the amount of transmission data can be further efficiently reduced. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 illustrates an example of the overall configuration of a system according to an embodiment. [Diagram 2] FIG. 2 is an explanatory diagram of a use case in the system of the embodiment. [Diagram 3] 4 is a diagram illustrating an example of a functional configuration of a transmission data generating unit according to the first embodiment. FIG. [Figure 4] FIG. 2 illustrates an example of a hardware configuration of a transmission data generating unit. [Diagram 5] 4 is a flowchart of a transmission data generation process according to the first embodiment. [Figure 6] 11 is a flowchart of a transmission data reduction method decision process according to the first embodiment. [Figure 7] FIG. 4 is an explanatory diagram of a method for reducing a time direction resolution in the first embodiment. [Figure 8]FIG. 4 is an explanatory diagram of a method for reducing spatial resolution in the first embodiment. [Figure 9] FIG. 4 is a diagram illustrating a specific example of determining a data amount reduction method according to the first embodiment. [Figure 10] FIG. 11 is a diagram illustrating a functional configuration of a transmission data generating unit according to the second embodiment. [Figure 11] 10 is a flowchart of a transmission data generation process according to the second embodiment. [Figure 12] 13 is a flowchart of a transmission data reduction method decision process according to the second embodiment. [Figure 13] FIG. 11 is a diagram illustrating a specific example of determining a data amount reduction method according to the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Hereinafter, an embodiment according to the present invention will be described with reference to the drawings. The following embodiments do not limit the present invention, and not all of the combinations of features described in the present embodiment are necessarily essential to the solution of the present invention. The configuration of the embodiment may be appropriately modified or changed depending on the specifications of the device to which the present invention is applied and various conditions (conditions of use, environment of use, etc.). In each of the following embodiments, the same reference symbols are given to the same or similar configurations and processing steps, and duplicated descriptions are omitted. In addition, in the reference symbols given to the same or similar configurations, when only the alphabets added after the numbers are different, it is assumed that it indicates different instances of the configurations having the same function. For example, the cameras 102A and 102B in FIG. 1 described later indicate different instances having the same function. In addition, having the same function refers to having at least a specific function (such as an imaging function in the case of a camera), and for example, the functions and performance of the cameras 102A and 102B may differ from each other.
[0011] <First embodiment> 1 is a diagram showing an example of the overall configuration of a system 100 to which an information processing device according to this embodiment is applied. The system 100 according to this embodiment is configured to include a plurality of imaging devices (cameras 102A to 102P), a three-dimensional model generating unit 103, and a transmission data generating unit 104. The cameras 102A to 102P and the three-dimensional model generating unit 103, and further the three-dimensional model generating unit 103 and the transmission data generating unit 104 are connected by a LAN (Local Area Network) (not shown). Note that the connection form is not limited to a LAN, and may be either a wired connection or a wireless connection.
[0012] The cameras 102A-102P are disposed at different positions around a field 101, which is an example of a subject to be photographed. Within the field 101, there is a subject to be three-dimensionally modeled. In the case of a sport (soccer in the example of FIG. 1) being taken as an example of a subject to be photographed, the subject is players playing the sport, equipment, tools, and the like. In the example of FIG. 1, the subjects, soccer players and a ball, are illustrated as three-dimensional modeling subjects 105A-105D in an approximate manner. Hereinafter, unless otherwise required, the cameras 102A-102P will be simply referred to as cameras 102, and the three-dimensional modeling subjects 105A-105D will be simply referred to as three-dimensional modeling subjects 105. These cameras 102 transmit the whole or a part of the images they have captured to the three-dimensional model generating unit 103.
[0013] The three-dimensional model generating unit 103 is a server that performs a process of generating a three-dimensional model using images acquired from each camera 102. The three-dimensional model generating unit 103 performs a shape estimation process for generating a three-dimensional model required for generating a virtual viewpoint image using images captured by each camera 102. For example, a volume intersection method is used for the shape estimation process. As a process of the volume intersection method, the three-dimensional model generating unit 103 first tiled a rectangular parallelepiped of a unit volume called a voxel in a target space for generating a three-dimensional model. Then, the three-dimensional model generating unit 103 projects a foreground silhouette of a target subject of interest extracted from an image captured by each camera 102 onto a group of voxels in the imaging range of each camera 102, and scrapes off voxels that fall outside the foreground silhouette. The three-dimensional model generating unit 103 performs this process for all cameras 102 to generate a three-dimensional model. At the same time, the three-dimensional model generating unit 103 uses image information contained in the foreground silhouette of each camera 102 as a texture image, and pastes it onto the three-dimensional model generated as described above, thereby coloring the three-dimensional model. In the following description of the embodiment, the three-dimensional model will be described as a three-dimensional point cloud that is a collection of colored voxels. Of course, the three-dimensional model is not limited to this, and may be, for example, a three-dimensional model based on a three-dimensional polygon mesh.
[0014] The transmission data generating unit 104 is an application example of the information processing device according to this embodiment, and is a server that generates transmission data to be sent to a user terminal from a three-dimensional model generated by the above-mentioned three-dimensional model generating unit 103. The transmission data generating unit 104 of this embodiment generates transmission data from three-dimensional model data (three-dimensional point cloud data which is a collection of colored voxels) generated by the three-dimensional model generating unit 103 and transmits it to the user terminal. The transmission data generating unit 104 may generate transmission data from two-dimensional image data generated by performing rendering using virtual viewpoint information indicating the position of a virtual viewpoint and the direction of a line of sight, and transmit the data to the user terminal. In this embodiment, an example will be described in which transmission data is generated from three-dimensional model data before rendering (three-dimensional point cloud data which is a collection of colored voxels) and transmitted to a user terminal. In this way, when transmitting three-dimensional model data before rendering to a user terminal, a rendering process is performed in the user terminal, and a virtual viewpoint image which is a two-dimensional image is generated and displayed. The information processing device according to this embodiment may include not only the transmission data generating unit 104 but also the above-mentioned three-dimensional model generating unit 103.
[0015] FIG. 2 is a diagram showing an example of a use case 200 in which the transmission data generated by the transmission data generating unit 104 is sent to the user terminals 210A to 210C, and the user terminals 210A to 210C generate and display a virtual viewpoint image from the transmission data. As shown in FIG. 2, the transmission data generating unit 104 is connected to the user terminals 210A to 210C via the Internet 201. Hereinafter, unless otherwise required, the user terminals 210A to 210C will be simply referred to as user terminals 210. The user terminal 210 includes the functions of a data receiving unit 213, a display unit 214, an operation unit 211, and an information transmitting unit 212. Note that the transmission path is not limited to the Internet 201, and may be any network in which a server (the transmission data generating unit 104) that generates the transmission data and a client terminal (the user terminal 210) that displays the virtual viewpoint image are connected.
[0016] The operation unit 211 acquires the position of a virtual viewpoint that the user desires to view and the line of sight direction of the virtual viewpoint, in other words, the position of a virtual camera and its shooting direction, as operation information input from the user of the user terminal 210. Examples of the operation unit 211 capable of acquiring the position and direction of the virtual viewpoint include input devices capable of specifying a three-dimensional spatial position and direction, such as a joystick type or a game controller type. Other examples of the operation unit 211 include a head-mounted display equipped with a sensor that can detect the direction in which the user's face is facing (line of sight direction) when worn by the user on the head.
[0017] Virtual viewpoint information including the position and direction of the virtual viewpoint input by the user via the operation unit 211 is sent to the display unit 214 or the information transmission unit 212. For example, in the case where transmission data of three-dimensional model data is sent from the transmission data generation unit 104, the virtual viewpoint information is sent to the display unit 214. In addition, in the case where transmission data of two-dimensional image data is sent from the transmission data generation unit 104, the virtual viewpoint information is sent to the information transmission unit 212. In addition, the virtual viewpoint information is also sent to the information transmission unit 212 in the case of a second embodiment described later. When the information transmission unit 212 receives virtual viewpoint information from the operation unit 211 , the information transmission unit 212 transmits the virtual viewpoint information to the transmission data generation unit 104 via the Internet 201 .
[0018] When the data receiving unit 213 of the user terminal 210 receives the transmission data from the transmission data generating unit 104 , it sends the received transmission data to the display unit 214 . The display unit 214 displays a virtual viewpoint image based on the transmission data received by the data receiving unit 213. For example, in a case where transmission data of a three-dimensional model is sent from the transmission data generating unit 104, the display unit 214 performs rendering processing based on the three-dimensional model data and the virtual viewpoint information acquired by the operation unit 211 to generate and display a virtual viewpoint image (two-dimensional image). Also, in a case where two-dimensional image data (virtual viewpoint image) is sent from the transmission data generating unit 104, the display unit 214 displays the virtual viewpoint image.
[0019] For example, in the case of transmitting transmission data of three-dimensional model data, the transmission data generation unit 104 generates transmission data of the three-dimensional model sent from the three-dimensional model generation unit 103 and transmits it to the user terminal 210. In addition, in the case of transmitting transmission data of two-dimensional image data, for example, the transmission data generation unit 104 receives virtual viewpoint information sent from the user terminal 210 and generates two-dimensional image data by performing rendering using the virtual viewpoint information and the three-dimensional model. Then, the transmission data generation unit 104 transmits the transmission data of the two-dimensional image data to the user terminal 210.
[0020] Here, when generating transmission data, the transmission data generating unit 104 according to the first embodiment acquires transmission bandwidth information indicating the transmission bandwidth available on the transmission path and movement information of the subject in the three-dimensional space. Then, based on the available transmission bandwidth information and the movement information of the subject, the transmission data generating unit 104 generates transmission data with a data amount that fits within the available transmission bandwidth. In other words, when the transmission data amount does not fit within the available transmission bandwidth, the transmission data generating unit 104 reduces the transmission data amount so that it fits within the available transmission bandwidth.
[0021] Hereinafter, the generation process of the transmission data in the transmission data generation unit 104 of this embodiment will be described using an example in which the transmission data is three-dimensional model data. Note that the case in which the transmission data is rendered two-dimensional image data will be described later.
[0022] The three-dimensional model data is a three-dimensional point cloud data of the three-dimensional modeling object 105 as shown in FIG. 1 described above, and is point cloud data for each object existing in a continuous area in a three-dimensional space. Therefore, the transmission data generating unit 104 can obtain position information of the three-dimensional modeling object 105 in the three-dimensional space (position information of the object) by calculating the geometric mean position of the continuous three-dimensional point cloud corresponding to a certain three-dimensional modeling object 105 of interest. The three-dimensional model data is also data in chronological order generated by the three-dimensional model generating unit 103 from images sequentially acquired by the camera 102 for each frame in chronological order. Therefore, the transmission data generating unit 104 can obtain the position information of each three-dimensional modeling object 105 in chronological order, and can obtain movement information indicating the movement of the three-dimensional modeling object 105 by using the position information in chronological order. That is, the transmission data generating unit 104 can obtain movement information of objects such as players and tools in FIG. 1.
[0023] In this embodiment, an example is given in which movement information is acquired from three-dimensional model data, but the present invention is not limited to this, and the three-dimensional model generating unit 103 or the transmission data generating unit 104 may acquire movement information of a subject, etc., from images captured sequentially by the camera 102. There are various existing techniques for acquiring movement (motion) information of a subject, etc., from images captured by the camera 102, and any of these techniques may be used.
[0024] Furthermore, the transmission data generating unit 104 acquires transmission bandwidth information indicating an available transmission bandwidth based on the usage status of the Internet 201 and the transmission bandwidth previously allocated to the transmission data generating unit 104 and the user terminal 210. The available transmission bandwidth information is a value indicating the amount of data that can be sent per unit time, expressed in units such as Byte / second (hereinafter, referred to as B / s).
[0025] As described above, the transmission data generation unit 104 of this embodiment acquires the movement information of the three-dimensional modeling object 105 and the transmission band information indicating the available transmission band, and determines a method for generating transmission data that fits within the available transmission band based on the information. Then, the transmission data generation unit 104 generates transmission data that fits within the available transmission band from the three-dimensional model data using the determined transmission data generation method.
[0026] FIG. 3 is a diagram showing an example of the functional configuration of the transmission data generation unit 104 according to this embodiment, which acquires movement information of the three-dimensional modeling object 105 and transmission bandwidth information indicating an available transmission bandwidth, and, based on that information, generates transmission data that fits within the available transmission bandwidth.
[0027] The model acquisition unit 301 acquires the three-dimensional model data 306 generated by the three-dimensional model generation unit 103, and outputs it to the subject position acquisition unit 302, the generation method determination unit 304, and the data generation unit 305. The three-dimensional model data 306 acquired by the model acquisition unit 301 is stored in a storage device such as a RAM (Random Access Memory) (not shown) provided in the transmission data generation unit 104, and used for subsequent processing. In the case of this embodiment, the storage device sequentially stores at least past three-dimensional model data 306 for a certain period from the present (current time).
[0028] The subject position acquisition unit 302 detects the current position of the three-dimensional modeling target 105 in three-dimensional space based on the three-dimensional model data 306 received from the model acquisition unit 301. As described above, the subject position acquisition unit 302 calculates the geometric mean position of the continuous three-dimensional point group corresponding to the three-dimensional modeling target 105, thereby acquiring position information of the three-dimensional modeling target 105 in three-dimensional space, that is, position information of the subject. The subject position acquisition unit 302 also causes the storage device described above to store the acquired position information. The current position information of the subject acquired by the subject position acquisition unit 302 and past position information for a certain period from the present that has already been stored in the storage device are then sent to the generation method determination unit 304.
[0029] The bandwidth information acquisition unit 303 acquires available transmission bandwidth information 307 and sends it to the generation method determination unit 304. As described above, the available transmission bandwidth information 307 is a value indicating the amount of data that can be transmitted per unit time expressed in units such as B / s, and is acquired based on the usage status of the Internet 201 shown in Fig. 2 and a previously assigned value.
[0030] The generation method determination unit 304 acquires movement information of the subject based on the current position information acquired by the subject position acquisition unit 302 and past position information for a certain period of time stored in the storage device. Furthermore, the generation method determination unit 304 determines a transmission data generation method for the downstream data generation unit 305 to generate transmission data based on the subject movement information, three-dimensional model data 306, and available transmission band information 307. The process of determining the transmission data generation method in the generation method determination unit 304 will be described in detail later. Then, the generation method determination unit 304 sends information indicating the determined transmission data generation method to the data generation unit 305.
[0031] The data generation unit 305 generates transmission data 308 from the three-dimensional model data 306 based on the transmission data generation method determined by the generation method determination unit 304, and transmits the transmission data 308 to the user terminal 210. When the transmission data is generated from two-dimensional image data, the data generation unit 305 generates the transmission data from the two-dimensional image data generated by performing rendering based on the three-dimensional model data 306 and the virtual viewpoint information sent from the user terminal 210. Details of the transmission data generation process in the data generation unit 305 will be described later.
[0032] Next, an example of the hardware configuration of the transmission data generation unit 104 will be described with reference to Fig. 4. As shown in Fig. 4, the hardware configuration of the transmission data generation unit 104 includes a CPU 401, a ROM 402, a RAM 403, an auxiliary storage device 404, a display unit 405, an operation unit 406, a communication I / F 407, and a bus 408. The CPU 401 realizes each functional unit of the transmission data generation unit 104 shown in Fig. 3 by controlling the entire transmission data generation unit 104 using the information processing program and data according to this embodiment stored in the ROM 402 or the auxiliary storage device 404. Note that the transmission data generation unit 104 may have one or more dedicated hardware pieces different from the CPU 401, and at least a part of the processing by the CPU 401 may be executed by the dedicated hardware pieces. Examples of the dedicated hardware pieces include an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), and a DSP (Digital Signal Processor).
[0033] The ROM 402 stores programs that do not require modification. The RAM 403 temporarily stores programs and data supplied from the ROM 402 or the auxiliary storage device 404, and data supplied from the outside via the communication I / F 407. The auxiliary storage device 404 is composed of, for example, a hard disk drive, and stores various data such as image data and audio data. The information processing program according to this embodiment is stored in the ROM 402 or the auxiliary storage device 404, and is expanded in the RAM 403 and executed by the CPU 401. It is assumed that the aforementioned storage device provided inside the transmission data generation unit 104 is the RAM 403.
[0034] The display unit 405 is configured to have, for example, a liquid crystal display or an LED, and displays a GUI (Graphical User Interface) for the user to operate the transmission data generation unit 104. The CPU 401 also operates as a display control unit that controls the display unit 405. The operation unit 406 is composed of, for example, a keyboard, a mouse, a joystick, a touch panel, etc., and receives operations from a user to input various instructions to the CPU 401. The CPU 401 also operates as an operation control unit that controls the operation unit 406.
[0035] The communication I / F 407 is used when the transmission data generation unit 104 communicates with an external device, etc. For example, when the transmission data generation unit 104 is connected to an external device by wire, a communication cable is connected to the communication I / F 407. When the transmission data generation unit 104 has a function of wirelessly communicating with an external device, the communication I / F 407 includes an antenna. A bus 408 connects each section within the transmission data generating unit 104 to transmit information. In addition, while FIG. 4 shows an example in which the display unit 405 and the operation unit 406 exist inside the transmission data generation unit 104, at least one of the display unit 405 and the operation unit 406 may exist as a separate device outside the transmission data generation unit 104.
[0036] Hereinafter, the process of generating transmission data from the three-dimensional model data 306 based on the three-dimensional model data 306, subject movement information, and available transmission bandwidth information 307 in the transmission data generation unit 104 according to the first embodiment will be described with reference to Figs. 5 to 8. Fig. 5 is a flowchart showing the flow of a transmission data generation process executed in the transmission data generation unit 104 shown in Fig. 3. In the following explanation of each flowchart, the symbol S represents a processing step (process).
[0037] In S501, the model acquisition unit 301 acquires the three-dimensional model data 306 from the three-dimensional model generation unit 103 as described above. Then, the model acquisition unit 301 outputs the three-dimensional model data 306 to the subject position acquisition unit 302, the generation method determination unit 304, and the data generation unit 305. Furthermore, the model acquisition unit 301 stores the three-dimensional model data 306 in the storage device (RAM 403). After this S501, the process of the transmission data generation unit 104 proceeds to S502. Also in S501, the subject position acquisition unit 302 acquires position information of the subject, which is the three-dimensional modeling target 105. As described above, the subject position acquisition unit 302 detects the current position of the subject, which is the three-dimensional modeling target 105, from the three-dimensional model data 306. Furthermore, the subject position acquisition unit 302 sequentially stores the subject position information in the storage device described above.
[0038] Next, in S502, the generation method determination unit 304 checks the amount of transmission data before reduction. In this embodiment, the amount of transmission data before reduction is the amount of transmission data when the time direction resolution and the space direction resolution in certain three-dimensional model data 306 are default values when a three-dimensional model is generated by the three-dimensional model generation unit 103. In this embodiment, the default value of the time direction resolution is 60 fps (frames per second), and the default value of the space direction resolution is 1 mm / voxel. That is, in S502, the generation method determination unit 304 checks that the amount of transmission data before reduction is the amount of data at 60 fps and 1 mm / voxel. After this S502, the processing of the transmission data generation unit 104 proceeds to S504.
[0039] In addition, in S503, which is performed in parallel with or prior to the above-mentioned S501, the bandwidth information acquisition unit 303 acquires the available transmission bandwidth information 307. As described above, the unit of the available transmission bandwidth is a unit of data amount per unit time, such as MB / second. In the case of this embodiment, the unit time is the time of one frame, and the bandwidth information acquisition unit 303 converts the available transmission bandwidth into the data amount per frame. Here, for example, if the available transmission bandwidth is 120 MB / second and the time direction resolution is 60 fps, the data amount per frame is 2 MByte. Then, the bandwidth information acquisition unit 303 sends information indicating the transmission data amount obtained by converting the available transmission bandwidth information 307 into the data amount per frame to the generation method determination unit 304. After this S503, the processing of the transmission data generation unit 104 proceeds to S504.
[0040] When the process proceeds to S504, the generation method determination unit 304 determines whether or not the transmission data amount needs to be reduced based on the transmission data amount before reduction confirmed in S502 and the available transmission band information 307 acquired by the band information acquisition unit 303 in S503. Then, the generation method determination unit 304 branches the process according to the determination result. For example, if the transmission data amount before reduction is equal to or less than the amount of data that can be transmitted in the available transmission band, the generation method determination unit 304 determines that the reduction of the transmission data amount is not necessary. That is, in this case, the generation method determination unit 304 does not change the time direction resolution and the spatial direction resolution from the default values. On the other hand, if the transmission data amount before reduction exceeds the amount of data that can be transmitted in the available transmission band, the generation method determination unit 304 advances the process to S505.
[0041] In S505, the generation method determination unit 304 determines a method for reducing the amount of transmission data. FIG. 6 is a flowchart showing a detailed flow of the transmission data reduction method determination process executed by the generation method determination unit 304 in S505 of FIG.
[0042] In S601 of FIG. 6, the generation method determination unit 304 executes a movement information acquisition process of the subject. That is, the generation method determination unit 304 acquires movement information of the subject (the three-dimensional modeling target 105) based on the current position information of the subject acquired by the subject position acquisition unit 302 and position information for a certain period of time stored in the storage device. The generation method determination unit 304 of this embodiment calculates the movement speed as the movement information of the subject. For example, the generation method determination unit 304 acquires position information of a certain target subject at two times, and calculates the movement speed from the time difference between the subject position information of the two times. In addition, for example, the generation method determination unit 304 may calculate the movement speed of the target subject using position information for each of a plurality of consecutive times in the past. However, when calculating the movement speed using subject position information of a plurality of times, there is a possibility that the movement direction of the subject is not linear within the period during which the plurality of position information is acquired. For this reason, the generation method determination unit 304 calculates the movement speed for each of two pieces of position information among the subject position information for multiple times, and calculates the average value of the movement speed for each of the two pieces of position information as the movement speed based on the subject position information for multiple times.
[0043] Next, in the process of S602, the generation method determination unit 304 determines whether the moving speed of the subject calculated in S601 is equal to or greater than a predetermined threshold, and branches the process according to the determination result. If the moving speed of the subject is equal to or greater than the threshold, the generation method determination unit 304 advances the process to S603, whereas if the moving speed is less than the threshold, the process advances to S604.
[0044] When the process proceeds to S603, the generation method determination unit 304 determines to use a method of reducing the spatial resolution as a method of reducing the amount of transmission data. On the other hand, if the process proceeds to S604, the generation method determination unit 304 determines to use a method of reducing the time direction resolution as a method of reducing the amount of transmission data. After the method for reducing the amount of transmission data is determined by the process of S603 or S604, the process of the transmission data generating unit 104 proceeds to S506 in FIG.
[0045] In S506, the generation method determination unit 304 branches the process depending on whether the transmission data amount reduction method determined in S505 is a method of reducing spatial resolution or a method of reducing temporal resolution. That is, if the transmission data amount reduction method determined in S505 is a method of reducing temporal resolution, the generation method determination unit 304 advances the process to S507, whereas if the transmission data amount reduction method determined in S505 is a method of reducing spatial resolution, the generation method determination unit 304 advances the process to S508.
[0046] When the process proceeds to S507, the generation method determination unit 304 determines a method for reducing the time direction resolution. Although details will be described later, when the process proceeds to S507, the generation method determination unit 304 performs a process of determining which of a plurality of different time direction resolutions to use based on the amount of transmission data before reduction and the available transmission band as a process for determining a method for reducing the time direction resolution.
[0047] On the other hand, if the process proceeds to S508, the generation method determination unit 304 determines a method for reducing the spatial resolution. Details will be described later, but if the process proceeds to S508, the generation method determination unit 304 performs a process of determining which of a plurality of different spatial resolutions to use as a method for reducing the spatial resolution based on the amount of transmission data before reduction and the available transmission band. After the process of S 507 or S 508 , the process of the transmission data generating unit 104 proceeds to S 509 which is executed by the data generating unit 305 .
[0048] In S509 , the data generation unit 305 generates the transmission data 308 from the three-dimensional model data 306 . For example, if it is determined in S504 that reduction in the data amount is not necessary and the process proceeds to S509, the data generator 305 generates the transmission data 308 according to the default time direction resolution and spatial direction resolution. For example, if the process proceeds from S504 to S505, and then from S506 to S507 to S509, the data generation unit 305 generates transmission data 308 having a default spatial resolution and a temporal resolution determined in S507. For example, if the process proceeds from S504 to S505, and then from S506 to S508 to S509, the data generation unit 305 generates transmission data 308 in which the time direction resolution is the default value and the spatial direction resolution is the time direction resolution determined in S508. Then, the transmission data 308 generated in the process of S 509 is output from the transmission data generating unit 104 .
[0049] Fig. 7 is a diagram showing an example of a time direction resolution determined by the generation method determination unit 304. Fig. 7 illustrates three time direction resolutions, including a default time direction resolution of 60 fps: large 701, a time direction resolution of 30 fps: medium 702, and a time direction resolution of 15 fps: small 703. Note that these are merely examples, and the frame rate of the time direction resolution may be set in a smaller number of steps.
[0050] 7, in the case of large time resolution 701, from time t0 to time t1, from time t1 to time t2, and from time t3 to time t4 each indicate a time interval of 60 fps. Also, subject 710 at time t0, subject 711 at time t1, subject 712 at time t2, subject 713 at time t3, and subject 714 at time t4 indicate the position and posture of each subject at each time.
[0051] In the case of medium time resolution 702 in Fig. 7, the time intervals from time t0 to time t2 and from time t2 to time t4 each indicate a time interval of 30 fps. Also, subject 710 at time t0, subject 712 at time t2, and subject 714 at time t4 indicate the position and orientation of each subject at each time. That is, in the case of medium time resolution 702, the time interval between frames is 1 / 30 seconds, and the position and orientation of the subject are the position and orientation every 1 / 30 seconds.
[0052] In Fig. 7, in case of small time resolution 703, the time interval from time t0 to time t4 indicates a time interval of 15 fps. Also, subject 710 at time t0 and subject 714 at time t4 indicate the position and orientation of each subject at each time. In other words, in case of small time resolution 703, the time interval between frames is 1 / 15 seconds, and the position and orientation of the subject are the same every 1 / 15 seconds.
[0053] Fig. 8 is a diagram showing an example of the spatial resolution determined by the generation method determination unit 304. Fig. 8 illustrates three spatial resolutions, including a default value of 1 mm / voxel: large spatial resolution 801, a default value of 2 mm / voxel: medium spatial resolution 802, and a default value of 4 mm / voxel: small spatial resolution 803. Note that these are only examples, and the voxel size of the time resolution may be set in a smaller number of parts.
[0054] In the case of the spatial resolution: large 801 in FIG. 8, a three-dimensional model 811 of the subject is constructed by combining voxels 812 of 1 mm / voxel, which is the default value. In the case of medium spatial resolution 802 in FIG. 8, a three-dimensional model 813 of the subject is constructed by combining voxels 814 of 2 mm / voxel. In the case of small spatial resolution 803 in FIG. 8, a three-dimensional model 815 of the subject is constructed by combining voxels 816 of 4 mm / voxel.
[0055] If it is determined in S504 that data amount reduction is not necessary, the generation method determination unit 304 determines the time resolution and the spatial resolution to be the default values shown in Fig. 7 and Fig. 8. That is, the generation method determination unit 304 determines the time resolution to be 60 fps, time resolution: large 701, and the spatial resolution to be 1 mm / voxel, spatial resolution: large 801. As a result, the data generation unit 305 generates transmission data 308 in which the time interval between frames is 1 / 60 seconds, the position and orientation of the subject are represented by the position and orientation every 1 / 60 seconds, and a three-dimensional model of the subject is configured by a combination of voxels of 1 mm / voxel.
[0056] On the other hand, when the process proceeds to S507, the generation method determination unit 304 determines either a temporal resolution of 30 fps: medium 702 or a temporal resolution of 15 fps: small 703 based on the default value of a temporal resolution of 60 fps: large 701. That is, the generation method determination unit 304 determines either a temporal resolution of 30 fps: medium 702 or a temporal resolution of 15 fps: small 703, which are temporal resolutions with a lowered frame rate, based on the amount of transmission data before reduction and the available transmission band. Note that when the process proceeds to S507, the generation method determination unit 304 determines the spatial resolution to be a default value of a spatial resolution of 1 mm / voxel: large 801.
[0057] For example, when the time direction resolution: medium 702 is determined, the time interval between frames is 1 / 30 seconds. In this case, the data generating unit 305 reduces the amount of transmission data by selecting frames at a predetermined interval from each frame of the base 60 fps, that is, by selecting every other frame, for example. In the example of FIG. 7, the data generating unit 305 selects the frames at time t0, time t2, and time t4, and deletes the frames at time t1 and time t3 to generate data of 30 fps with the time direction resolution: medium 702. As a result, the amount of transmission data when the time direction resolution: medium 702 is determined is reduced to 1 / 2 of the amount of transmission data before reduction, that is, the time direction resolution: high 701.
[0058] Furthermore, for example, when the time direction resolution: small 703 is determined, the time interval between frames is 1 / 15 seconds. In this case, the data generating unit 305 reduces the amount of transmission data by selecting frames at a predetermined interval from each frame of the base 60 fps, that is, by selecting every third frame, for example. In the example of FIG. 7, the data generating unit 305 selects the frames at time t0 and time t4, and deletes the frames at time t1, time t2, and time t3, thereby generating data of 15 fps with the time direction resolution: small 703. As a result, the amount of transmission data when the time direction resolution: small 703 is determined is reduced to 1 / 4 of the amount of transmission data before reduction, that is, the time direction resolution: large 701.
[0059] In addition, the method of reducing the time direction resolution is not limited to simply deleting frames as described above, but may be to use data of unused frames to interpolate data of frames that are actually used as transmission data. For example, the interpolation may be performed by performing blurring processing on voxel information with respect to the relative direction of voxels of unused frames. In addition, in this embodiment, the case where the transmission data is three-dimensional model data has been described as an example, but the same method of reducing the time direction resolution can be used even when the transmission data is rendered two-dimensional image data. In other words, even when the transmission data is rendered two-dimensional image data, a method of deleting unused frames can be used according to the required reduction in the amount of transmission data.
[0060] Also, when the process proceeds to S508, the generation method determination unit 304 determines either 2 mm / voxel spatial resolution: medium 802 or 4 mm / voxel spatial resolution: small 803 based on the default value of 1 mm / voxel spatial resolution: large 801. That is, the generation method determination unit 304 determines either 2 mm / voxel spatial resolution: medium 802 or 4 mm / voxel spatial resolution: small 803, which are spatial resolutions with large voxel sizes, based on the amount of transmission data before reduction and the available transmission band. Note that when the process proceeds to S508, the generation method determination unit 304 determines the time direction resolution to be 60 fps temporal resolution: large 701, which is the default value.
[0061] For example, when the spatial resolution is determined to be medium 802, a three-dimensional model 813 of the subject is constructed from voxels 814 of 2 mm / voxel. In this case, the data generating unit 305 selects, for example, one of eight 1 mm voxels included in the 2 mm voxel. Alternatively, the data generating unit 305 may determine the presence or absence of a voxel by selecting one of the eight 1 mm voxels by majority vote, or by comparing the eight 1 mm voxels with a threshold value. In addition, the color to be colored into the voxel may be selected from one representative point, or may be the average color of the eight 1 mm voxels.
[0062] Furthermore, for example, when spatial resolution: small 803 is determined, a three-dimensional model 815 of the subject is constructed from 4 mm / voxel voxels 816. In this case as well, the data generation unit 305 determines the presence or absence of a voxel by, for example, selecting one of the 1 mm voxels contained in the 4 mm voxel, or by a majority vote or threshold value of the 1 mm voxels. Note that the color to be applied to the voxels when spatial resolution: small 803 is determined may be determined by selecting a representative point or using an average value, as described above.
[0063] The reduction in the amount of transmission data due to the spatial resolution as described above, when calculated only in terms of the number of voxels, is 1 / 8 for spatial resolution: medium 802 and 1 / 64 for spatial resolution: small 803 compared to the spatial resolution: large 801 before reduction. However, generally, voxels constituting the interior of a three-dimensional model do not have the same amount of data as the surface. For this reason, when reducing the amount of transmission data due to the spatial resolution, the reduction in the amount of transmission data is estimated based on the change in the surface area of the three-dimensional model. As a result, the amount of transmission data reduced by the spatial resolution is approximately 1 / 4 for spatial resolution: medium 802 and approximately 1 / 16 for spatial resolution: small 803 compared to the spatial resolution: large 801 before reduction. Note that the estimated reduction depends on how the data of the voxels constituting the three-dimensional model is stored as described above, and therefore needs to be calculated according to the data format of the system to be applied.
[0064] In this embodiment, the case where the transmission data is three-dimensional model data has been described as an example, but when the transmission data is rendered two-dimensional image data, the method of reducing the spatial resolution is to reduce the rendering resolution. This process can be performed by determining the amount of transmission data reduction corresponding to medium and small spatial resolutions and determining the rendering resolution. For example, the amount of transmission data can be set to approximately 1 / 2 for medium spatial resolution and approximately 1 / 4 for small spatial resolution, and the rendering resolution can be set to achieve these amounts of data.
[0065] FIG. 9 is a diagram showing a specific example of numerical values used when determining a data amount reduction method in the transmission data generation unit 104 of the first embodiment described above. In Fig. 9, case (a) 911, case (b) 912, and case (c) 913 illustrated in case 901 each show an example of a situation in which the shape and moving speed of the object are different. Fig. 9 shows examples of numerical values of the number of objects of interest 902, the amount of data transmitted per frame 903, the amount of data transmitted per second 904, the available transmission bandwidth 905, the moving speed of the object 906, and the moving speed threshold value 907 for each case 901. Below, the process of determining the data amount reduction method for each of cases (a) 911 to (c) 913 will be described in accordance with the processing steps of the flowcharts in Figs. 5 and 6 described above.
[0066] The example of case (a) 911 shows a case where the three-dimensional model data acquired in S501 of Fig. 5 is data in which the transmission data amount per frame for one person of interest subjects is 5 MB / frame. Here, since the default frame rate is 60 fps, in the transmission data amount confirmation process in S502, it is confirmed that the transmission data amount per second before reduction is 300 MB / sec. In addition, in the example of case (a) 911, the available transmission band acquired in S503 is, for example, 400 MB / sec. Therefore, in the data amount reduction necessity determination process in S504 based on this information, it is determined that the data amount is within the available transmission band and therefore reduction of the transmission data amount is unnecessary, and the process proceeds to the process of S509, and transmission data is generated without reducing the transmission data amount.
[0067] An example of case (b) 912 shows a case where the three-dimensional model data acquired in S501 of Fig. 5 requires 10 MB / frame as the amount of transmission data per frame for one person of interest. In this case, the transmission data amount confirmation process in S502 confirms that the transmission data amount per second before reduction is 600 MB / sec. Also in this case, the available transmission bandwidth acquired in S503 is, for example, 400 MB / sec. In this case, since the data amount exceeds the available transmission bandwidth, the data amount reduction necessity determination process in S504 determines that the transmission data amount needs to be reduced, and the process proceeds to S505.
[0068] When the process proceeds to S505, information on the moving speed of the target subject and the threshold value of the moving speed is acquired as the process of S601 in Fig. 6. In the case (b) 912, since the moving speed of the target subject is 15km / h and the threshold value of the moving speed is 10km / h, the determination process of S602 determines that the moving speed of the target subject is equal to or greater than the threshold value, and the process proceeds to S603. That is, in S603, it is decided to reduce the spatial resolution, and therefore, in Fig. 5, the process proceeds to S508 via the process of S506.
[0069] Then, when the process proceeds to S508, a reduction method is determined taking into consideration how much the spatial resolution needs to be reduced to fit within the available transmission bandwidth. In the case (b) 912, the amount of transmission data before reduction is 600 MB / sec, and the available transmission bandwidth is 400 MB / sec. In other words, if the amount of transmission data before reduction is halved, for example, it can be fit within the available transmission bandwidth. Therefore, in the process of determining the method of reducing the spatial resolution in S508, spatial resolution: medium 802 in FIG. 8 is determined, and the amount of transmission data is 300 MB / sec. When the process proceeds to S508, the default value of 60 fps and temporal resolution: large 701 is used as the time resolution, and therefore transmission data of 60 fps 2 mm / voxel is generated in S509.
[0070] In other words, the result of the example of case (b) 912 means that when the subject's moving speed is equal to or greater than a threshold value, emphasis is placed on reproducing the subject's movement, and transmission data is generated that includes a three-dimensional model that maintains the temporal resolution (frame rate).
[0071] In the case (c) 913 example, like the case (b) 912 example described above, the process proceeds from S501 to S505. In the case (c) 913, in S601 of Fig. 6, the moving speed of the target subject is 5km / h and the threshold value of the moving speed is 10km / h, so in S602, it is determined that the moving speed of the target subject is less than the threshold value, and the process proceeds to S604. That is, in the case (c) 913, it is determined in S603 that the time direction resolution is to be reduced, and therefore in Fig. 5, the process proceeds to S507 via the process of S506.
[0072] Then, when the process proceeds to S507, a reduction method is determined taking into consideration how much the time direction resolution needs to be reduced to fit within the available transmission bandwidth. In the case (c) 913, the amount of transmission data before reduction is 600 MB / sec, and the available transmission bandwidth is 400 MB / sec. In other words, if the amount of transmission data before reduction is, for example, halved, it can be fit within the available transmission bandwidth. Therefore, in the process of determining the reduction direction of the time direction resolution in S507, the time direction resolution: medium 702 in FIG. 7 is determined, and the amount of transmission data is 300 MB / sec. When the process proceeds to S507, the spatial resolution is set to the default value of 1 mm / voxel, spatial direction resolution: large 801, and therefore transmission data of 30 fps·1 mm / voxel is generated in S509.
[0073] In other words, the result of the example case (c) 913 means that when the object's moving speed is below a threshold, emphasis is placed on reproducing the object's shape, and transmission data is generated that includes a three-dimensional model that maintains spatial resolution.
[0074] In each of the above-mentioned cases 901, the case where the number of target objects is one has been described as an example, but the number of target objects may be two or more. When the number of target objects is two or more, the moving speed of the object may be, for example, the moving speed of the object present at the center of the image, or the fastest moving speed among the objects, or the average value of the moving speeds of all the target objects.
[0075] As described above, in the transmission data generating unit 104 of this embodiment, when the amount of data becomes equal to or exceeds the available transmission band, the amount of transmission data can be reduced according to the movement information of the subject. That is, according to this embodiment, the amount of transmission data can be efficiently reduced.
[0076] <Second embodiment> In the first embodiment, an example of efficiently reducing the amount of transmission data using the movement information of the subject has been described, but in the following second embodiment, an example of reducing the amount of transmission data using the movement information of the virtual viewpoint in addition to the movement information of the subject will be described. Note that in the second embodiment, components with the same reference numbers as those in the first embodiment operate in the same manner as those described in the first embodiment, and therefore their description will be omitted. The system configuration of the second embodiment is the same as that in FIG. 1, the use case is the same as that in FIG. 2, and the hardware configuration of the transmission data generation unit 104 is the same as that in FIG. 4, and therefore their illustration and description will be omitted.
[0077] 10 is a diagram showing an example of a functional configuration capable of generating transmission data that fits within an available transmission band based on movement information of a three-dimensional modeling target 105, information on an available transmission band, and movement information of a virtual viewpoint in a transmission data generation unit 104 according to the second embodiment. The transmission data generation unit 104 of the second embodiment generates transmission data 308 that fits within an available transmission band based on three-dimensional model data 306, available transmission band information 307, and virtual viewpoint information 1003.
[0078] The transmission data generation unit 104 of the second embodiment includes a virtual viewpoint position acquisition unit 1001. The virtual viewpoint position acquisition unit 1001 receives virtual viewpoint information 1003, which is input by the user via the above-mentioned operation unit 211 and transmitted from the information transmission unit 212 via the Internet 201. The virtual viewpoint position acquisition unit 1001 acquires position information of the virtual viewpoint from the virtual viewpoint information 1003. This virtual viewpoint position information is sent to a generation method determination unit 1002, and the storage device sequentially stores past virtual viewpoint position information for a certain period from the present (current time). The current virtual viewpoint position information acquired from the virtual viewpoint information 1003 and past virtual viewpoint position information for a certain period held in the storage device are sent to the generation method determination unit 1002.
[0079] The generation method determination unit 1002 of the second embodiment determines a transmission data generation method based on the virtual viewpoint position information, the three-dimensional model data 306, the subject position information (position information of the three-dimensional modeling target 105), and the available transmission band information 307. The transmission data generation method determination process in the generation method determination unit 1002 of the second embodiment will be described in detail later. Then, the generation method determination unit 1002 sends information indicating the determined transmission data generation method to the data generation unit 305. The data generation unit 305 generates and outputs transmission data 308 from the three-dimensional model data 306 or the two-dimensional image data based on the transmission data generation method determined by the generation method determination unit 1002, in the same manner as in the example of the first embodiment described above.
[0080] Hereinafter, with reference to Figures 11 to 13, a process will be described in which the transmission data generation unit 104 according to the second embodiment generates transmission data from the three-dimensional model data 306 based on the position information of the subject, the position information of the virtual viewpoint, and available transmission bandwidth information.
[0081] Fig. 11 is a flowchart showing the flow of the transmission data generation process executed by the transmission data generation unit 104 according to the second embodiment shown in Fig. 10. The processes from S501 to S504 are similar to the processes described in Fig. 5. In the case of the flowchart of Fig. 11, if it is determined in S504 that the amount of data needs to be reduced, the process proceeds to S1101.
[0082] In S1101, the virtual viewpoint position acquisition unit 1001 acquires the position information of the virtual viewpoint from the virtual viewpoint information 1003. In addition, the virtual viewpoint position acquisition unit 1001 stores the position information of the virtual viewpoint in the storage device (RAM 403).
[0083] Next, in the process of S1102, the generation method determination unit 304 determines a method for reducing the amount of transmission data. FIG. 12 is a flowchart showing a detailed flow of the transmission data reduction method determination process executed by the generation method determination unit 304 in S1102 of FIG.
[0084] In S1201 of FIG. 12, the generation method determination unit 1002 acquires a moving speed as moving information of the virtual viewpoint based on the current position information of the virtual viewpoint acquired by the virtual viewpoint position acquisition unit 1001 and the position information of the virtual viewpoint for a certain period of time stored in the storage device. The moving speed of the virtual viewpoint can be calculated from, for example, the position information of two consecutive times and the time difference between those two times, in the same manner as the calculation process of the movement information of the subject. In addition, the moving information of the virtual viewpoint can also be calculated using the position information of multiple virtual viewpoints obtained for multiple consecutive times in the past. However, when calculating the moving speed using the position information of these multiple past virtual viewpoints, there is a possibility that the moving direction of the virtual viewpoint during the period when the multiple position information was acquired is not linear. Therefore, when calculating the moving speed using multiple past position information, the generation method determination unit 1002 calculates the moving information for each of two consecutive position information, and calculates the average value as the moving speed.
[0085] Next, in the process of S1202, the generation method determination unit 1002 determines whether the moving speed of the virtual viewpoint is equal to or greater than a predetermined threshold value, and branches the process according to the determination result. If the moving speed of the virtual viewpoint is equal to or greater than the threshold value, the generation method determination unit 1002 advances the process to S603, whereas if the moving speed of the virtual viewpoint is less than the threshold value, the process advances to S601. The processes from S601 to S604 are similar to those described with reference to FIG. 6, and therefore their description will be omitted.
[0086] In the second embodiment, when it is determined in S1202 that the moving speed of the virtual viewpoint is equal to or greater than the threshold value, it is determined to use a method of reducing the spatial resolution in S603 while maintaining the time resolution (frame rate) regardless of the moving speed of the subject. In other words, if the moving speed of the virtual viewpoint is fast, the movement of the entire virtual viewpoint image becomes fast regardless of the moving speed of the subject, so it is determined to use a method of reducing the spatial resolution.
[0087] FIG. 13 is a diagram showing a specific example of numerical values used when determining the data amount reduction process in the transmission data generation unit 104 of the second embodiment. In Fig. 13, case (a) 1311, case (b) 1312, case (c) 1313, and case (d) 1314 illustrated in case 1301 show examples of situations in which the shape and moving speed of the subject and the moving speed of the virtual viewpoint are different. Fig. 13 shows the number of subjects of interest 1302, the amount of transmission data per frame 1303, the amount of transmission data per second 1304, the available transmission bandwidth 1305, the moving speed of the subject 1308, and a threshold value 1309 for each case 1301. In the case of the second embodiment, a moving speed of the virtual viewpoint 1306 and a threshold value 1307 for the moving speed of the virtual viewpoint are further added.
[0088] The data amount reduction method decision process will be described below for each of the cases (a) 1311 to (d) 1314 in accordance with the processing steps of the flowcharts of FIGS. In the case of case (a) 1311, the same processing results as those in the example of case (a) 911 in Fig. 9 are obtained in S501, S502, and S503 in Fig. 11. That is, since the amount of transmission data falls within the available transmission band, it is determined in the data amount reduction necessity determination process of S504 that reduction in the amount of transmission data is unnecessary, and the process proceeds to S509, where transmission data is generated without reducing the amount of transmission data.
[0089] An example of case (b) 1312 shows a case where the three-dimensional model data acquired in S501 requires 10 MB / frame as the amount of transmission data per frame for one person of interest. In this case, the transmission data amount confirmation process in S502 confirms that the amount of transmission data per second before reduction is 600 MB / sec. Also in this case, the available transmission bandwidth acquired in S503 is, for example, 400 MB / sec. In this case, since the amount of data exceeds the available transmission bandwidth, the data amount reduction necessity determination process in S504 determines that the amount of transmission data needs to be reduced, and the process proceeds to S1101.
[0090] When the process proceeds to S1101, information on the moving speed of the virtual viewpoint and the threshold value of the moving speed is acquired as the process of S1201 in Fig. 12. In the case (b) 1312, since the moving speed of the virtual viewpoint is 20 km / h and the threshold value of the moving speed is 16 km / h, in the determination process of S1202, it is determined that the moving speed of the virtual viewpoint is equal to or greater than the threshold value, and the process proceeds to the process of S603. In other words, in S603, it is determined that the spatial direction resolution is to be reduced, and therefore, in Fig. 11, the process proceeds to S508 via the process of S506.
[0091] Then, when the process proceeds to S508, a reduction method is determined taking into consideration how much the spatial resolution needs to be reduced to fit within the available transmission bandwidth. In the case (b) 1312 example, the amount of transmission data before reduction is 600 MB / sec, and the available transmission bandwidth is 400 MB / sec. In other words, if the amount of transmission data before reduction is halved, for example, it can be fit within the available transmission bandwidth. Therefore, in the process of determining the method of reducing the spatial resolution in S508, the spatial resolution: medium 802 in FIG. 8 is determined, and the amount of transmission data is 300 MB / sec. When the process proceeds to S508, the default value of 60 fps and temporal resolution: large 701 is used as the time resolution, and therefore transmission data of 60 fps 2 mm / voxel is generated in S509.
[0092] In other words, the result of the example of case (b) 1312 means that when the moving speed of the virtual viewpoint is equal to or greater than a threshold value, emphasis is placed on reproducing the movement of the virtual viewpoint, and transmission data is generated that includes a three-dimensional model that maintains the temporal resolution (frame rate).
[0093] In the case of the example of case (c) 1313, the process proceeds from S501 to S505, similarly to the above-mentioned example of case (b) 1312. In the case of case (c) 1313, the moving speed of the virtual viewpoint is 12 km / h and the threshold value of the moving speed is 16 km / h, so in S1201 of Fig. 12, it is determined that the moving speed of the virtual viewpoint is less than the threshold value, and the process proceeds to S601.
[0094] When the process proceeds to S601, the moving speed of the target object and the threshold value for that moving speed are obtained in the same manner as described above. In the case (c) 1313 example, the moving speed of the target object is 15 km / h and the threshold value for the moving speed of the object is 10 km / h, so in S602, it is determined that the moving speed of the target object is equal to or greater than the threshold value, and the process proceeds to S603. In other words, it is decided to reduce the spatial resolution, and the process proceeds to S508 via S506 in FIG. 11.
[0095] When the process proceeds to S508, a reduction method is determined taking into consideration how much the spatial resolution needs to be reduced to fit within the available transmission bandwidth. In the case (c) 1313, the amount of transmission data before reduction is 600 MB / sec, and the available transmission bandwidth is 400 MB / sec. In other words, if the amount of transmission data before reduction is halved, for example, it can be fit within the available transmission bandwidth. Therefore, in the process of determining the method of reducing the spatial resolution in S508, the spatial resolution: medium 802 in FIG. 8 is determined, and the amount of transmission data is 300 MB / sec. When the process proceeds to S508, the default value of 60 fps and temporal resolution: large 701 is used as the time resolution, and therefore transmission data of 60 fps 2 mm / voxel is generated in S509.
[0096] The result of the example of case (c) 1313 means that even if the moving speed of the virtual viewpoint is less than the threshold, if the moving speed of the subject is equal to or greater than the threshold, importance is attached to reproducing the movement of the subject. In other words, in case (c) 1313, transmission data is generated that includes a three-dimensional model that maintains the time direction resolution (frame rate).
[0097] In the case of the example of case (d) 1314, the process proceeds from S501 to S505, similarly to the above-mentioned example of case (b) 1312. In the case of case (d) 1314, the moving speed of the virtual viewpoint is 12 km / h and the threshold value of the moving speed is 16 km / h, so in S1201 of Fig. 12, it is determined that the moving speed of the virtual viewpoint is less than the threshold value, and the process proceeds to S601.
[0098] When the process proceeds to S601, the moving speed of the target object and the threshold value for that moving speed are obtained in the same manner as described above. In the case (d) 1314 example, the moving speed of the target object is 5 km / h and the threshold value for the moving speed of the object is 10 km / h, so in S602, it is determined that the moving speed of the target object is less than the threshold value and the process proceeds to S604. In other words, it is decided to reduce the time direction resolution, and the process proceeds to S507 via S506 in FIG. 11.
[0099] When the process proceeds to S507, a reduction method is determined taking into consideration how much the time direction resolution needs to be reduced to fit within the available transmission bandwidth. In the case (d) 1314, the amount of transmission data before reduction is 600 MB / sec, and the available transmission bandwidth is 400 MB / sec. In other words, if the amount of transmission data before reduction is halved, for example, it can be fit within the available transmission bandwidth. Therefore, in the process of determining the method of reducing the time direction resolution in S507, the time direction resolution: medium 702 in FIG. 7 is determined, and the amount of transmission data is 300 MB / sec. When the process proceeds to S507, the spatial resolution is set to the default value of 1 mm / voxel, spatial direction resolution: large 801, and therefore transmission data of 30 fps·1 mm / voxel is generated in S509.
[0100] The result of the example of case (d) 1314 means that when the moving speed of the virtual viewpoint and the object is less than the threshold, importance is attached to reproducing the shape of the object. That is, in the example of case (d) 1314, transmission data including a three-dimensional model that maintains the spatial resolution is generated.
[0101] As described above, in the transmission data generating unit 104 of the second embodiment, when the amount of data becomes equal to or exceeds the available transmission band, the amount of transmission data can be reduced based on the movement information of the subject and the virtual viewpoint. That is, according to this embodiment, the amount of transmission data can be efficiently reduced.
[0102] In the case of the technology disclosed in the above-mentioned Patent Document 1, the amount of data can be reduced by reducing the spatial resolution, but the amount of data cannot be reduced by reducing the time resolution as in the first and second embodiments. In the technology of Patent Document 1, even in a scene in which the subject of interest remains stationary, the transmission data of the stationary subject continues to be transmitted at a high frame rate, occupying most of the communication bandwidth. In contrast, according to the first and second embodiments, even in a case where the available transmission bandwidth is tight, an appropriate method of reducing the transmission data can be adopted according to the movement of the subject or the virtual viewpoint, and the deterioration of the three-dimensional model can be minimized.
[0103] <Other embodiments> In the above-mentioned first embodiment, an example of reducing the amount of transmission data based on information on available transmission bands and movement information of a subject has been described, and in the second embodiment, an example of reducing the amount of transmission data by separately handling the movement information of a subject and the movement information of a virtual viewpoint has been described. As another embodiment, the amount of transmission data may be reduced based on information on available transmission bands and the movement information of a virtual viewpoint. In this case, if the movement speed of the virtual viewpoint is equal to or greater than a threshold value, it is determined that the spatial resolution is reduced, whereas if the movement speed of the virtual viewpoint is less than the threshold value, it is determined that the time resolution is reduced.
[0104] In the first and second embodiments, only the case where only one of the time resolution and the space resolution is reduced to satisfy the available transmission bandwidth has been described, but it is also possible to combine both the reduction in the time resolution and the reduction in the space resolution. When combining both, it is possible to gradually reduce the amount of transmission data by repeating the process from S504 to S509 in the flowchart of FIG. 5 of the first embodiment or the flowchart of FIG. 11 of the second embodiment multiple times. When repeating the process from S504 to S509 multiple times, the threshold value of the moving speed of the subject or the threshold value of the moving speed of the virtual viewpoint may be appropriately changed according to the number of times of repetition, and the reduction method may be determined for each loop of repetition.
[0105] In the second embodiment, the moving speed of the virtual viewpoint and the moving speed of the subject are acquired independently to determine the transmission data generation method, but it is also possible to determine the transmission data generation method based on the relative moving speed between the virtual viewpoint and the subject. For example, when the virtual viewpoint and the subject are moving in the same direction (their speed vectors are approximately the same), the subject viewed from the virtual viewpoint is almost stationary on the virtual viewpoint image. In this case, maintaining the spatial resolution and reducing the time resolution can maintain higher quality of the image quality of the virtual viewpoint image.
[0106] The present invention can also be realized by supplying a program for implementing one or more of the functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions. The above-mentioned embodiments are merely examples of the implementation of the present invention, and the technical scope of the present invention should not be interpreted as being limited by these. In other words, the present invention can be implemented in various forms without departing from its technical concept or main characteristics.
[0107] The disclosure of this embodiment includes the following configuration, method, and program. (Configuration 1) A bandwidth information acquisition means for acquiring information on a transmission bandwidth available for data transmission; a model acquisition means for acquiring a three-dimensional model of a subject in a three-dimensional space; a movement information acquisition means for acquiring movement information of at least one of the subject and a virtual viewpoint in the three-dimensional space; a determination means for determining a transmission data generation method based on the information on the transmission band and the movement information; a data generating means for generating transmission data of the three-dimensional model or transmission data of a two-dimensional image generated from the three-dimensional model based on the transmission data generating method; 13. An information processing device comprising: (Configuration 2) The transmission data generation method includes a plurality of transmission data generation methods for generating transmission data each having a different amount of data, 2. The information processing device according to configuration 1, wherein the determining means determines one of the plurality of transmission data generation methods based on the information on the transmission band and the movement information. (Configuration 3) the plurality of transmission data generation methods include a first transmission data generation method that is a default, and a second transmission data generation method that generates transmission data with a reduced amount of data compared to the first transmission data generation method; The information processing device according to configuration 2, wherein the decision means determines whether or not to select the second transmission data generation method based on information about the transmission bandwidth and an amount of data generated by the first transmission data generation method. (Configuration 4) The second transmission data generation method includes a third transmission data generation method for generating transmission data with reduced time resolution and a fourth transmission data generation method for generating transmission data with reduced spatial resolution; The information processing device according to configuration 3, characterized in that when it is determined that the second transmission data generation method is to be selected, the decision means decides on either the third transmission data generation method or the fourth transmission data generation method based on the transmission bandwidth information and the mobility information. (Configuration 5) the movement information acquisition means acquires a movement speed as the movement information of either the subject or the virtual viewpoint; 5. The information processing apparatus according to configuration 4, wherein the decision means decides on the fourth transmission data generation method when the moving speed is equal to or greater than a predetermined threshold value. (Configuration 6) the movement information acquisition means acquires a movement speed as the movement information of either the subject or the virtual viewpoint; 6. The information processing apparatus according to configuration 4 or 5, wherein the decision means decides on the third transmission data generation method when the moving speed is less than a predetermined threshold value. (Configuration 7) the movement information acquisition means acquires, as the movement information of the subject and the virtual viewpoint, movement information of the subject and movement information of the virtual viewpoint; 5. The information processing apparatus according to configuration 4, wherein the decision means decides on the fourth transmission data generation method when the moving speed of the virtual viewpoint is equal to or greater than a predetermined threshold value. (Configuration 8) the movement information acquisition means acquires, as the movement information of the subject and the virtual viewpoint, movement information of the subject and movement information of the virtual viewpoint; The information processing device according to configuration 4 or 7, wherein the decision means decides to use the fourth transmission data generation method when the moving speed of the virtual viewpoint is less than a predetermined threshold and the moving speed of the subject is equal to or greater than a predetermined threshold. (Configuration 9) the movement information acquisition means acquires, as the movement information of the subject and the virtual viewpoint, movement information of the subject and movement information of the virtual viewpoint; The information processing device according to any one of configurations 4, 7, and 8, wherein the decision means decides to use the third transmission data generation method when the moving speed of the virtual viewpoint is less than a predetermined threshold and the moving speed of the subject is less than a predetermined threshold. (Configuration 10) the movement information acquisition means acquires a relative movement speed between the subject and the virtual viewpoint as the movement information of the subject and the virtual viewpoint; 5. The information processing apparatus according to configuration 4, wherein the decision means decides on the fourth transmission data generation method when the relative moving speed is equal to or greater than a predetermined threshold value. (Configuration 11) the movement information acquisition means acquires a relative movement speed between the subject and the virtual viewpoint as the movement information of the subject and the virtual viewpoint; 11. The information processing apparatus according to configuration 4 or 10, wherein the decision means decides on the third transmission data generation method when the relative moving speed is less than a predetermined threshold value. (Configuration 12) 12. The information processing device according to any one of configurations 4 to 11, wherein the third transmission data generation method is a method for generating transmission data having a frame rate lower than that of the first transmission data generation method. (Configuration 13) 12. The information processing device according to any one of configurations 4 to 11, wherein the fourth transmission data generation method is a method for generating transmission data in which a voxel size of a three-dimensional model is larger than that of the first transmission data generation method. (Configuration 14) The information processing device according to any one of configurations 4 to 11, wherein the fourth transmission data generation method is a method of generating transmission data having a lower resolution for rendering a two-dimensional image than the first transmission data generation method. (Configuration 15) The information processing device described in any one of configurations 5 to 9, characterized in that the movement information acquisition means acquires an average value of multiple movement speeds calculated using multiple consecutive positions of the subject of interest at each time as the movement speed of the subject of interest, or acquires an average value of multiple movement speeds calculated using multiple consecutive positions of a virtual viewpoint at each time as the movement speed of the virtual viewpoint. (Configuration 16) The information processing device according to any one of configurations 5 to 9, wherein the movement information acquisition means acquires any one of the movement speed of a subject at a center position among a plurality of subjects, the movement speed of a subject moving the fastest, and an average value of the movement speeds of the plurality of subjects. (Configuration 17) The information processing device according to configuration 5 or 6, characterized in that the process of acquiring the movement information by the movement information acquisition means and the process of determining the transmission data generation method by the determination means are repeated a plurality of times, and the threshold value is changed depending on the number of repetitions. (Method 1) a bandwidth information acquisition step of acquiring information on a transmission bandwidth available for data transmission; a model acquisition step of acquiring a three-dimensional model of a subject in a three-dimensional space; a movement information acquisition step of acquiring movement information of at least one of the subject and a virtual viewpoint in the three-dimensional space; a determination step of determining a transmission data generation method based on the information on the transmission band and the movement information; a data generation step of generating transmission data of the three-dimensional model or transmission data of a two-dimensional image generated from the three-dimensional model based on the transmission data generation method; 13. An information processing method comprising: (Program 1) A program for causing a computer to function as the information processing device according to any one of configurations 1 to 17. [Explanation of symbols]
[0108] 104: Transmission data generation unit, 301: Model acquisition unit, 302: Subject position acquisition unit, 303: Bandwidth information acquisition unit, 304: Generation method determination unit, 305: Data generation unit
Claims
1. A model acquisition means for acquiring a three-dimensional model of an object in three-dimensional space, Movement information acquisition means for acquiring movement information of at least one of the subject and the virtual viewpoint in the three-dimensional space, A determination means for determining which transmission data to be transmitted from among the transmission data based on the three-dimensional model, based on the movement information, A transmission means for transmitting the transmission data determined to be transmitted by the determination means, An information processing device characterized by having the following features.
2. The information processing apparatus according to Claim 1, characterized in that the transmission data based on the three-dimensional model is either the transmission data of the three-dimensional model or the transmission data of a two-dimensional image generated from the three-dimensional model.
3. The information processing apparatus according to Claim 1, wherein the determination means determines the transmission data to be transmitted based on information indicating the available transmission bandwidth and the movement information.
4. The information processing apparatus according to claim 3, characterized in that the determination means determines whether to transmit a first transmission data or a second transmission data having a reduced data volume compared to the first transmission data.
5. The information processing apparatus according to claim 4, characterized in that the second transmission data is transmission data with reduced temporal resolution or transmission data with reduced spatial resolution.
6. The movement information acquisition means acquires the movement speed of either the subject or the virtual viewpoint as movement information. The information processing apparatus according to claim 5, characterized in that the determination means determines, when the movement speed is greater than or equal to a predetermined threshold, to transmit data with reduced spatial resolution.
7. The movement information acquisition means acquires the movement speed of either the subject or the virtual viewpoint as movement information. The information processing apparatus according to claim 5, characterized in that the determination means determines, when the moving speed is less than a predetermined threshold, to transmit data with reduced temporal resolution.
8. The movement information acquisition means acquires movement information of the subject and movement information of the virtual viewpoint as movement information. The information processing apparatus according to claim 5, characterized in that the determination means determines to transmit data with reduced spatial resolution when the movement speed of the virtual viewpoint is greater than or equal to a predetermined threshold.
9. The movement information acquisition means acquires movement information of the subject and movement information of the virtual viewpoint as movement information. The information processing apparatus according to claim 5, characterized in that the determination means determines to transmit data with reduced spatial resolution when the movement speed of the virtual viewpoint is less than a predetermined threshold and the movement speed of the subject is greater than or equal to a predetermined threshold.
10. The movement information acquisition means acquires movement information of the subject and movement information of the virtual viewpoint as movement information. The information processing apparatus according to claim 5, characterized in that the determination means determines to transmit data with reduced temporal resolution when the movement speed of the virtual viewpoint is less than a predetermined threshold and the movement speed of the subject is less than a predetermined threshold.
11. The movement information acquisition means acquires the relative movement speed between the subject and the virtual viewpoint as movement information. The information processing apparatus according to claim 5, characterized in that the determination means determines to transmit data with reduced spatial resolution when the relative movement speed is greater than or equal to a predetermined threshold.
12. The movement information acquisition means acquires the relative movement speed between the subject and the virtual viewpoint as the movement information, The information processing apparatus according to claim 5, characterized in that the determination means determines to transmit data with reduced temporal resolution when the relative movement speed is less than a predetermined threshold.
13. The information processing apparatus according to claim 5, characterized in that the transmission data with reduced temporal resolution is transmission data with a lower frame rate than the first transmission data.
14. The information processing apparatus according to claim 5, characterized in that the transmission data with reduced spatial resolution is transmission data with a lower spatial resolution of the three-dimensional model than the first transmission data.
15. The information processing apparatus according to claim 5, characterized in that, when the transmission data with reduced spatial resolution is transmission data of a two-dimensional image, the resolution of the two-dimensional image is lower than the resolution of the two-dimensional image in the first transmission data.
16. The information processing apparatus according to claim 1, wherein the movement information acquisition means acquires the average value of a plurality of movement speeds calculated using the position of the subject of interest at multiple consecutive time intervals as the movement speed of the subject of interest, or acquires the average value of a plurality of movement speeds calculated using the position of a virtual viewpoint at multiple consecutive time intervals as the movement speed of the virtual viewpoint.
17. The information processing apparatus according to claim 1, wherein the movement information acquisition means acquires one of the following from among a plurality of subjects: the movement speed of the subject at the center position, the movement speed of the subject moving the fastest, or the average value of the movement speeds of the plurality of subjects.
18. A model acquisition step of acquiring a three-dimensional model of an object in three-dimensional space, A movement information acquisition step that acquires movement information of at least one of the subject and the virtual viewpoint in the three-dimensional space, A determination step in which, among the transmission data based on the three-dimensional model, the transmission data to be transmitted is determined based on the movement information, A transmission step which transmits the transmission data that was determined to be transmitted in the determination step, An information processing method characterized by having the following features.
19. Computers, A model acquisition means for acquiring a three-dimensional model of an object in three-dimensional space, Movement information acquisition means for acquiring movement information of at least one of the subject and the virtual viewpoint in the three-dimensional space, A determination means for determining which transmission data to be transmitted from among the transmission data based on the three-dimensional model, based on the movement information, A transmission means for transmitting the transmission data determined to be transmitted by the determination means, A program to cause an information processing device to function as such.