Information processing apparatus, information processing method, and program
The information processing device optimizes data reduction by selecting image capture devices with minimal impact on image quality, addressing the challenge of excessive data in virtual viewpoint image generation, thereby extending recording time and maintaining image quality.
Patent Information
- Application Number
- JP2024095704
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-13
- Publication Date
- 2025-12-25
AI Technical Summary
Generating high-quality virtual viewpoint images using multiple image capture devices and high-resolution images results in excessive data that exceeds storage capacity, making it difficult to determine the optimal data reduction while maintaining image quality, especially in varying shooting conditions.
An information processing device that acquires free space on a recording medium and transmits instructions to specific image capture devices to stop data output when capacity thresholds are reached, selecting devices based on their impact on image quality and overlap with other devices to minimize data reduction impact.
Reduces recorded data while preserving image quality by strategically stopping data output from selected capture devices, extending recording time and maintaining image quality based on shooting conditions.
Smart Images

Figure 2025187140000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] A technology that uses multiple image capture devices installed at different positions surrounding an area to be captured and synchronized to capture the image, and then generates a virtual viewpoint image using the captured images, has been attracting attention. To generate a high-quality virtual viewpoint image, it is desirable to use many captured images captured from different positions and to use high-resolution captured images. That is, it is desirable to use many image capture devices and capture the subject at a high resolution and a high frame rate. However, using many image capture devices or capturing images at a high resolution and a high frame rate increases the amount of captured image data, which may exceed the amount of data that can be stored on a recording medium depending on the subject and the shooting environment. As a result, it becomes impossible to record the captured images, and a virtual viewpoint image cannot be generated. Thus, while generating a high-quality virtual viewpoint image is desirable, reducing the amount of captured image data is also desirable, resulting in a trade-off between the two.
[0003] Patent Document 1 describes a method in which, among multiple imaging devices, an imaging device corresponding to an external signal is preset. Then, when the external signal is received, the frame rate of imaging devices other than the preset imaging device is lowered or the captured image is recorded at a high compression rate, thereby reducing the amount of data recorded on a recording medium. In other words, it describes a method in which an imaging device for which data volume is to be reduced is preset. Considering the description in Patent Document 1, it is conceivable that, among multiple imaging devices, an imaging device for which data volume is to be reduced is preset when generating a virtual viewpoint image. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-19450 Summary of the Invention [Problem to be solved by the invention]
[0005] However, when capturing virtual viewpoint images, the capture time is not always clearly defined, making it difficult to determine how much data volume should be reduced while maintaining a suitable image quality. For example, when capturing a baseball game, the capture time varies greatly depending on the game situation. A game with few points requires a shorter capture time, while a game with many points requires a longer capture time.
[0006] An object of the present disclosure is to reduce the amount of data recorded for capturing images for generating a virtual viewpoint image while suppressing the effect on the image quality of the virtual viewpoint image in accordance with the shooting conditions. [Means for solving the problem]
[0007] An information processing device according to one aspect of the present disclosure has the following configuration: an acquisition means for acquiring free space of a recording medium for recording a plurality of pieces of image data output from a plurality of image capture devices used to generate a virtual viewpoint image; a transmitting means for transmitting an instruction not to output image data to a specific image capture device among the plurality of image capture devices when the available capacity is equal to or less than a threshold value; It has. [Effects of the Invention]
[0008] According to the present disclosure, it is possible to reduce the amount of data recorded for captured images used to generate a virtual viewpoint image while suppressing the effect on the image quality of the virtual viewpoint image according to the shooting conditions. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a diagram illustrating an example of the configuration of an image processing system 1. FIG. [Figure 2]FIG. 1 illustrates an example of a hardware configuration of an information processing device 100. [Figure 3] 1 is a diagram illustrating an example of the configuration of an information processing device 100. FIG. [Figure 4] 1 is a flowchart illustrating the operation flow of the image processing system 1. [Figure 5] FIG. 4 is a diagram illustrating a method for selecting an imaging device to be subjected to data reduction according to the first embodiment. [Figure 6] FIG. 10 is a diagram illustrating a method for selecting an imaging device to be subjected to data reduction according to the second embodiment. [Figure 7] FIG. 11 is a diagram illustrating a method for selecting an imaging device to be subjected to data reduction according to the third embodiment. [Figure 8] FIG. 10 is a diagram illustrating an example of a hardware configuration of an information processing device 100 according to a fifth embodiment. [Figure 9] FIG. 10 is a diagram illustrating an example of the configuration of an information processing device 100 according to a fifth embodiment. [Figure 10] 10 is a flowchart illustrating an operation flow of the image processing system 1 according to the fifth embodiment. [Figure 11] 13 is a diagram illustrating a relationship between an imaging device according to a fifth embodiment and the texture of a photographed subject. FIG. [Figure 12] FIG. 13 is a diagram illustrating a method for selecting an imaging device to be subjected to data reduction according to the fifth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] <Embodiment> The information processing device according to the present disclosure includes an acquisition unit that acquires free space on a recording medium that records multiple pieces of image data output from multiple image capture devices used to generate a virtual viewpoint image, and a transmission unit that, when the free space falls below a threshold, transmits instruction information for preventing the image data output from a specific one of the image capture devices from being recorded on the recording medium.
[0011] Furthermore, if the instruction information is information indicating an instruction not to output the image data generated by the specific imaging device, the transmitting means may transmit the instruction information to the specific imaging device. In this way, the image data generated by the specific imaging device is not recorded on a recording medium, thereby reducing the amount of data recorded on the recording medium. Furthermore, since the generation of image data in the specific imaging device continues, the output of the image data can be resumed as necessary.
[0012] Furthermore, if the instruction information is information indicating an instruction to stop generating the image data, the transmission unit may transmit the instruction information to the specific imaging device. In this way, it is possible to reduce the amount of data recorded on a recording medium and also reduce the cost of generating image data in the specific imaging device.
[0013] Furthermore, if the instruction information is information indicating an instruction not to record the image data generated by the specific imaging device on the recording medium, the transmitting means may transmit the instruction information to the recording medium. In this way, it is possible to check whether an abnormality has occurred in the transmission bandwidth of the image data while reducing the amount of data recorded on the recording medium.
[0014] In one aspect of the present disclosure, among the plurality of imaging devices, imaging devices other than the specific imaging device output the image data even when the free space falls below a threshold.
[0015] In this way, even if the free space falls below the threshold, it is possible to continue generating virtual viewpoint images.
[0016] In one aspect of the present disclosure, the specific image capture device is determined based on the positions and orientations of the multiple image capture devices.
[0017] In one aspect of the present disclosure, the specific imaging device is an imaging device among the plurality of imaging devices that has a large area where its imaging range overlaps with other imaging devices.
[0018] In one aspect of the present disclosure, the system includes a first calculation means for calculating a first maximum angle, which is the largest angle among multiple angles formed by a line connecting a certain coordinate on an XY plane with the coordinates of the multiple imaging devices. Furthermore, for each of the multiple imaging devices, excluding one imaging device, the system calculates a second maximum angle, which is the largest angle among multiple angles formed by a line connecting the certain coordinate with the remaining imaging devices. The system also includes a second calculation means for calculating the second maximum angle when only one of the multiple imaging devices is excluded. In this case, the specific imaging device is the imaging device excluded from the multiple imaging devices when the second maximum angle is smaller than the first maximum angle.
[0019] In one aspect of the present disclosure, each of the plurality of image capture devices includes a calculation unit that calculates a similarity score indicating a degree of similarity between the image capture device and the other image capture devices in terms of three-dimensional coordinates and orientation, and the specific image capture device is an image capture device among the plurality of image capture devices that has a large similarity score.
[0020] In this way, when setting an imaging device that generates captured images that are set as targets not to be recorded in order to reduce the amount of data on the recording medium, it is possible to set an imaging device that has little impact on the image quality of the virtual viewpoint image.
[0021] In one aspect of the present disclosure, the specific imaging device is a plurality of imaging devices, and the storage device includes a determination unit that determines the number of the specific imaging devices based on the available capacity.
[0022] In this way, the number of imaging devices to be targeted for data reduction can be determined based on the available capacity without the user having to give instructions at any time. The number of specific imaging devices is set in advance. The available capacity threshold and the number of specific imaging devices are also recorded in association with each other. A plurality of combinations of the available capacity threshold and the number of specific imaging devices may be provided. In this case, the number of specific imaging devices can be set to be larger as the available capacity becomes smaller.
[0023] The present disclosure also relates to a computer program for controlling each of the above-described means by a computer.
[0024] <Example> Hereinafter, examples for carrying out the present disclosure will be described with reference to the drawings. Note that the following examples do not limit the present disclosure, and not all of the combinations of features described in the examples are necessarily essential to the solutions of the present disclosure. Note that the same components will be described with the same reference numerals, and duplicate descriptions will be omitted.
[0025] A virtual viewpoint image is an image generated by a user freely manipulating the position and orientation of a virtual camera, and is also called a free viewpoint image, an arbitrary viewpoint image, etc. Unless otherwise specified, the term "image" will be explained as including the concepts of both moving images and still images.
[0026] The viewpoint information used to generate a virtual viewpoint image is information indicating the position and orientation (line of sight direction) of the virtual viewpoint. Specifically, the viewpoint information is a parameter set including a parameter indicating the three-dimensional position of the virtual viewpoint and a parameter indicating the orientation of the virtual viewpoint in the pan, tilt, and roll directions. Note that the content of the viewpoint information is not limited to the above. For example, the parameter set serving as viewpoint information may include a parameter indicating the size of the field of view (angle of view) of the virtual viewpoint. Furthermore, the viewpoint information may have multiple parameter sets. For example, the viewpoint information may have multiple parameter sets corresponding to multiple frames constituting a moving image of the virtual viewpoint image, and may be information indicating the position and orientation of the virtual viewpoint at each of multiple consecutive time points.
[0027] The image processing system described below has multiple imaging devices that capture images of an imaging area from multiple directions. The imaging area may be, for example, a stadium where sports such as soccer or karate are held, or a stage where a concert or a play is held. The multiple imaging devices are installed at different positions surrounding the imaging area and capture images synchronously. Note that the multiple imaging devices do not need to be installed around the entire periphery of the imaging area; depending on installation space restrictions, they may be installed only around a portion of the periphery of the imaging area. Furthermore, the number of imaging devices is not limited to the example shown in the figure. For example, if the imaging area is a soccer stadium, approximately 30 imaging devices may be installed around the stadium. Furthermore, imaging devices with different functions, such as telephoto cameras and wide-angle cameras, may be installed.
[0028] In this embodiment, the multiple image capturing devices are cameras each having an independent housing and capable of capturing images from a single viewpoint. However, this is not limiting, and two or more image capturing devices may be configured in the same housing. For example, a single camera equipped with multiple lens groups and multiple sensors and capable of capturing images from multiple viewpoints may be installed as the multiple image capturing devices.
[0029] A virtual viewpoint image is generated, for example, by the following method. First, multiple images (multiple captured images) are obtained by capturing images from different directions using multiple imaging devices. Next, a foreground image in which a foreground region corresponding to a predetermined object, such as a person or a ball, is extracted, and a background image in which a background region other than the foreground region is extracted are obtained from the multiple captured images. Furthermore, a foreground model representing the three-dimensional shape of the predetermined object and texture data for coloring the foreground model are generated based on the foreground image, and texture data for coloring a background model representing the three-dimensional shape of a background, such as a stadium, is generated based on the background image. Then, the texture data is mapped to the foreground model and background model, and rendering is performed according to the virtual viewpoint indicated by the viewpoint information, thereby generating a virtual viewpoint image. However, the method for generating a virtual viewpoint image is not limited to this, and various methods can be used, such as a method of generating a virtual viewpoint image by projective transformation of captured images without using a three-dimensional model.
[0030] A foreground image is an image in which an object region (foreground region) is extracted from an image captured by an imaging device. An object extracted as a foreground region is a dynamic object (moving body) that moves (its absolute position and shape can change) when images are captured from the same direction in a time series. Examples of objects include players, referees, and other people on the field where a sport is being played, such as a ball in a ball game, or singers, musicians, performers, and presenters in a concert or entertainment event.
[0031] A background image is an image of at least a region (background region) different from the foreground object. Specifically, a background image is an image in which the foreground object has been removed from the captured image. Furthermore, the background refers to an imaged object that remains stationary or nearly stationary when images are captured from the same direction in chronological order. Examples of such imaged objects include a stage for a concert, a stadium where an event such as a sport is held, a structure such as a goal used in a ball game, or a field. However, the background is at least a region different from the foreground object, and the imaged object may include other objects in addition to the object and background.
[0032] A virtual camera is a virtual camera that is different from the multiple imaging devices actually installed around the imaging area, and is a concept for conveniently explaining a virtual viewpoint related to the generation of a virtual viewpoint image. That is, a virtual viewpoint image can be considered to be an image captured from a virtual viewpoint set in a virtual space associated with the imaging area. The position and orientation of the viewpoint in the virtual image capture can be expressed as the position and orientation of the virtual camera. In other words, a virtual viewpoint image can be said to be an image that simulates an image captured by a camera if it were assumed that the camera were located at the virtual viewpoint set in space.
[0033] Example 1 In this embodiment, an example is described in which the image processing system 1 is used to identify an imaging device that has little impact on the image quality of the virtual viewpoint image even if it is set as a target for data volume reduction, and the recording time of the virtual viewpoint image is extended by stopping the data recording of the identified imaging device.
[0034] FIG. 1 is a diagram illustrating an example of the configuration of an image processing system 1 according to this embodiment. The image processing system 1 includes multiple image capture devices 10, multiple image processing devices 20, an information processing device 100, a modeling device 30, a database 40, a virtual viewpoint input device 50, a rendering device 60, and a display device 70. Each component of the image processing system 1 may function as a single device, or some of the functions may be assigned to an external server or other device. Specifically, the rendering device 60 may be handled by an external server, and the display device 70 may be handled by a tablet terminal. Note that one image capture device 10 and one image processing device 20 may be treated as a pair, or one image processing device 20 may be treated as a pair for multiple image capture devices 10. Alternatively, the image processing system 1 may include one image processing device 20 for all image capture devices 10. In this embodiment, one image capture device 10 and one image processing device 20 are treated as a pair.
[0035] The imaging device 10 photographs the imaging area 2 and transmits image data of the photographed image to the image processing device 20. Multiple imaging devices 10 are installed to surround the imaging area 2. The multiple imaging devices 10 synchronize their photographing timing to photograph the imaging area 2 from different positions at the same time. The multiple photographed images obtained by synchronized photographing are used to estimate the three-dimensional shape of the subject 3.
[0036] The image processing device 20 processes image data acquired from the imaging device 10 into a format suitable for transmission. For example, the image processing device 20 can perform lossless compression to reduce the amount of data required during transmission. Alternatively, the image processing device 20 may separate the acquired image data into the foreground (i.e., the subject) and the background, generating image data of only the foreground, thereby reducing the amount of data. The image processing device 20 may separate the foreground (i.e., the subject 3) from the rest of the background (foreground-background separation processing) to later estimate the three-dimensional shape of the subject 3, generating a mask image that distinguishes the foreground and background in binary. In this case, the image processing device 20 transmits texture data of only the foreground together with the mask image, thereby reducing data during transmission without affecting the three-dimensional shape estimation. Furthermore, the image processing device 20 may perform data reduction by performing lossless or lossy compression on various image data, as described above. The image processing device 20 also acquires imaging device control information, such as hardware configuration information including the sensor size and lens information of the imaging device 10, as well as focal length and aperture, by communicating with the imaging device 10.
[0037] The modeling device 30 estimates the three-dimensional shape of the subject based on the image data transmitted from the image processing device 20. The three-dimensional shape of the subject is estimated using, for example, a volume intersection method, and shape data representing the three-dimensional shape, such as a point cloud or mesh, is generated. The modeling device 30 may also generate texture data representing color information for each component of the generated three-dimensional shape. Depending on the three-dimensional shape estimation method, additional data may be generated. In this embodiment, the modeling device 30 generates shape data and texture data and outputs them to the database 40. The modeling device 30 also acquires three-dimensional coordinates and orientation information of the image capture device 10. For example, markers are placed in the imaging area 2 and captured by multiple image capture devices 10. The modeling device 30 processes the captured image data to acquire the three-dimensional coordinates and orientation information of the image capture device. This process is called calibration, and since it is an existing technology, a detailed description will be omitted. The three-dimensional coordinates of the image capture device here refer to three-dimensional coordinates when the same coordinate system is set for multiple image capture devices. The origin of this coordinate system may be set to a specific object or marker in real space. In generating a virtual viewpoint image, the coordinate system of a virtual space including a background model, which is a 3D model representing the background of a stadium or the like, generated separately, is aligned with the coordinate system used during calibration of the multiple image capture devices. Therefore, calibration of the multiple image capture devices can be considered a process of identifying information about the positions and orientations of the multiple image capture devices in real space. Calibration of the multiple image capture devices can also be considered a process of identifying the positions and orientations in virtual space that correspond to the positions and orientations of the multiple image capture devices in real space. The coordinate system of the virtual space and the coordinate system used during calibration do not need to be strictly aligned; for example, a predetermined amount of misalignment is acceptable. In this case, the position and orientation of the 3D model of the subject generated based on the multiple image capture devices are determined taking into account the predetermined amount of misalignment.
[0038] The information processing device 100 is responsible for monitoring the image processing device 20, the modeling device 30, and the database 40. The information processing device 100 may be realized as a single device, or the control functions of each part may be divided and controlled by a plurality of devices in cooperation with each other.
[0039] The database 40 records the shape data and texture data generated by the modeling device 30 on a recording medium. The database 40 also manages the usage status of the recording medium installed therein, and transmits free space information indicating the free space to the information processing device 100. Note that it is sufficient if the information processing device 100 can provide information for identifying the free space on the recording medium, and the maximum recording capacity and the currently used capacity may be transmitted to the information processing device 100.
[0040] A user viewing a virtual viewpoint image inputs virtual viewpoint information to the virtual viewpoint input device 50. The virtual viewpoint information may include some or all of position and orientation information indicating the three-dimensional coordinates and orientation information of a virtual camera, and information indicating a virtual lens such as a focal length and an aperture.
[0041] The rendering device 60 generates a virtual viewpoint image based on the virtual viewpoint information input from the virtual viewpoint input device 50. The process of generating a virtual viewpoint image is called rendering. The generated virtual viewpoint image is output to the display device 70.
[0042] The display device 70 displays the virtual viewpoint image generated by the rendering device 60. This allows the user to view the virtual viewpoint image.
[0043] 2 is a diagram showing an example of the hardware configuration of a computer that can realize the image processing system 1 according to this embodiment. The information processing device 100 includes a CPU 201, a ROM 202, a RAM 203, an auxiliary storage device 204, a display unit 205, an operation unit 206, a communication I / F 207, and a bus 208.
[0044] The CPU 201 realizes each function of the information processing device 100 using computer programs and data stored in the ROM 202 and RAM 203. Note that the information processing device 100 may have one or more dedicated hardware components different from the CPU 201, and at least a portion of the processing by the CPU 201 may be executed by the dedicated hardware components. Examples of the dedicated hardware components include an ASIC (application-specific integrated circuit), an FPGA (field-programmable gate array), and a DSP (digital signal processor). The ROM 202 stores programs that do not require modification. The RAM 203 temporarily stores programs and data supplied from the auxiliary storage device 204, as well as data supplied from the outside via the communication I / F 207. The auxiliary storage device 204 is formed, for example, by a hard disk drive or the like, and stores various data such as image data and sound data.
[0045] The display unit 205 is configured with, for example, a liquid crystal display, an LED, or the like, and displays the generated virtual viewpoint image and a GUI (Graphical User Interface) for the user to operate the information processing device 100. The operation unit 206 is configured with, for example, a keyboard, a mouse, a joystick, a touch panel, or the like, and receives operations by the user to input various instructions to the CPU 201. The CPU 201 operates as a display control unit that controls the display unit 205 and an operation control unit that controls the operation unit 206. Furthermore, the CPU 201 is also responsible for controlling the ROM 202 and RAM 203, controlling data transfer from the auxiliary storage device 204, and controlling the bus 208 and communication I / F 207.
[0046] The communication I / F 207 is used for communication with devices external to the information processing device 100. For example, when the information processing device 100 is connected to an external device via a wired connection, a communication cable is connected to the communication I / F 207. When the information processing device 100 has a function for wireless communication with an external device, the communication I / F 207 is equipped with an antenna. The bus 208 connects each unit of the information processing device 100 to transmit information.
[0047] In this embodiment, the display unit 205 and the operation unit 206 are assumed to exist inside the information processing device 100. However, at least a part of the display unit 205 and the operation unit 206 may exist outside the information processing device 100 as separate devices.
[0048] FIG. 3 is a diagram illustrating an example of the configuration of the information processing device 100. As shown in FIG.
[0049] The recording setting acquisition unit 301 acquires the current recording settings from the image processing device 20 and the database 40. The recording settings include hardware information such as the sensor size and lens of the multiple image capturing devices 10.
[0050] The reduction device selection unit 302 receives the recording settings acquired by the recording setting acquisition unit 301 and the three-dimensional coordinates and orientation information of the multiple imaging devices 10 acquired by the imaging device position information acquisition unit 303. Then, based on the recording settings and the three-dimensional coordinates and orientation information of the multiple imaging devices 10, it selects reduction devices that indicate imaging devices for which data volume reduction is desired. The reduction device selection unit 302 also determines the number of imaging devices to be selected as reduction devices based on available capacity. For example, it calculates the ratio of available capacity to maximum storage capacity and determines the number of imaging devices to be selected as reduction devices based on whether the ratio of available capacity is equal to or less than a threshold. The maximum available capacity is set to 100%, and when the available capacity falls below 50%, the number of imaging devices to be selected as reduction devices is set to 10. Furthermore, when the available capacity falls below 30%, the number of imaging devices to be selected as reduction devices is set to 20. Multiple thresholds may be set in this manner, and the number of imaging devices to be selected as reduction devices may be increased each time the available capacity falls below a threshold. It should be noted that the number of imaging devices to be selected as reduction devices to be set for the available capacity is not limited to these values, and the user can set any value in advance.
[0051] By controlling the image processing device 20 that controls the selected imaging device 10 to stop transmitting image data, the amount of data recorded in the database is reduced and the remaining recording time is extended. The method for selecting the reduced device will be described later.
[0052] The image capture device position information acquisition unit 303 acquires the three-dimensional coordinate information and orientation information of the multiple image capture devices 10 from the modeling device 30 .
[0053] FIG. 4 is a flowchart showing a process of reducing recorded data by the information processing apparatus 100 according to the present embodiment.
[0054] In S401, the recording setting acquisition unit 301 acquires the free space of the recording medium from the database 40.
[0055] In S402, the reduction device selection unit 302 starts reducing recorded data when the free space of the recording medium obtained from the database 40 reaches a preset threshold. If the free space is equal to or less than the threshold, the process proceeds to S403. If the free space is greater than the threshold, the process proceeds to S401.
[0056] In S403, the imaging device position information acquisition unit 303 acquires the three-dimensional coordinates and orientation information of the multiple imaging devices 10 from the modeling device 30. In addition, the recording setting acquisition unit 301 acquires hardware information of the multiple imaging devices 10 from the image processing device 20.
[0057] In S404, the reduction device selection unit 302 selects the number of image capture devices to be reduced and the image capture devices to be reduced, depending on the free space on the recording medium of the database 40 acquired in S401.
[0058] In this embodiment, the reduction device selection unit 302 selects one imaging device for data reduction using a method of calculating the angle between imaging devices for each unit within the imaging area 2. FIG. 5 shows a top view of an arrangement of four imaging devices. First, the reduction device selection unit 302 acquires the sensor size, lens hardware information, three-dimensional coordinate information, and posture information of each imaging device from the recording setting acquisition unit 301 and the imaging device position information acquisition unit 303, and calculates the angle of view (imaging range). Next, the reduction device selection unit 302 divides the imaging area into unit areas (unit volumes when the height direction is taken into consideration), and calculates and confirms the imaging devices capturing each unit. In the example of FIG. 5, a total of four imaging devices, namely imaging devices 511-514, capture images of a unit surface 502 within the imaging area 501. Next, the reduction device selection unit 302 identifies lines connecting the coordinate points of each imaging device 511-514 and the center point of the unit surface 502, and calculates angles 521-524 between the imaging devices formed by the multiple lines (first calculation). From the calculation, it is found that the maximum angle is 143 degrees, which is the inter-imaging device angle 524. Therefore, the maximum angle (first maximum angle) between the imaging devices of unit surface 502 is 143 degrees. By having device-to-be-reduced selection unit 302 calculate the angle between the imaging devices for each unit surface, it is possible to evaluate the degree of impact on image quality at each location within the imaging area. Furthermore, in order to identify imaging devices that have a small impact on image quality, device-to-be-reduced selection unit 302 calculates the angle between the imaging devices and the maximum angle (second maximum angle) when an imaging device is removed for each unit surface (second calculation). In FIG. 5, a case is considered in which each of imaging devices 511-514 is removed one by one. That is, the second maximum angle is calculated when each of imaging devices 511-514 is removed. In this case, the only time the aforementioned maximum inter-device angle of 143 degrees is not exceeded is when imaging device 513 is removed. Therefore, device-to-be-reduced selection unit 302 selects imaging device 513 as a device that has a small impact on image quality. The reduction device selection unit 302 selects an imaging device that has the least effect on image quality for each unit surface of the imaging area, and ultimately determines the device that has the least effect on the most unit surfaces as the target for data reduction by majority vote.
[0059] In S405, the reduction device selection unit 302 outputs an instruction to the image processing device 20 connected to the image capture device 10 to be reduced to stop texture transmission. In this embodiment, the reduction device selection unit 302 stops only the texture data transmission from the image processing device 20, and transmits mask image data and the like generated as a result of foreground / background separation. Therefore, it is possible to partially reflect the shooting results of the corresponding image capture device in the three-dimensional shape information generated by the modeling device 30.
[0060] In S406, the recording setting acquisition unit 301 checks whether the user has stopped recording. If the user has issued a command to stop transmission to the database 40 via the operation unit 206, the reduction device selection unit 302 stops data transmission to each image processing device. If the recording setting acquisition unit 301 cannot confirm that the recording has stopped, the process returns to S401 and continues to monitor the free space on the recording medium.
[0061] According to this embodiment, it is possible to reduce the amount of recording data recorded on a recording medium and extend the available recording time while suppressing the effect on the image quality of the virtual viewpoint image.
[0062] <Example 2> In this embodiment, an example is described in which the image processing system 1 is used to score imaging devices that have little impact on image quality based on three-dimensional coordinate information and posture information, and data recording is stopped to extend the recording time of virtual viewpoint images.
[0063] The configuration of the image processing system 1 according to this embodiment is the same as that of the first embodiment described above.
[0064] Furthermore, the recording data reduction flow shown in FIG. 4 is also the same.
[0065] In this embodiment, in S404, an imaging device with a small impact on image quality is selected using a method for determining imaging devices with similar angles of view based on three-dimensional coordinate information and orientation information. This method determines whether the angles of view are similar by assigning a score to each imaging device relative to the three-dimensional coordinates and orientation. In this embodiment, the similarity is compared using a demerit system. The score that takes into account the three-dimensional coordinates and orientation is called the similarity score.
[0066] Figure 6 shows a top view of an arrangement of three imaging devices. While Figure 6 only shows coordinate and angle information on the XY plane, coordinate and angle information in the Z (height) direction is also included in practice. In this embodiment, if the coordinates and the pan, yaw, and tilt angles of the imaging devices are the same, each receives a perfect score (100 points), and subtractions are made from these. Therefore, the similarity score is calculated based on the total of 200 points for positions and angles. In this embodiment, 5 points are subtracted for every 1 m of distance between the imaging devices, and 5 points are subtracted for every 9 degrees of angle between the imaging devices. For example, if the distance between two imaging devices is 5 m, 75 points are awarded, and if the angle between the imaging devices is 17 degrees, 95 points are awarded. In this case, the total of the distance score and angle score, 170 points, is the similarity score for the combination of target imaging devices. In practice, the total score and the weight of the subtractions may be set according to the importance and influence of each element. Furthermore, in this embodiment, if the calculated distance or angle between the imaging devices is not divisible, the result is rounded to the nearest whole number. Under the conditions described above, the image capture devices 601-603 would receive 175 points, the image capture devices 601-602 would receive 75 points, and the image capture devices 602-603 would receive 75 points. Therefore, of the combinations of image capture devices, image capture device 603, which has the highest similarity score, is selected as the target for data reduction. Note that, because the data actually includes orientation information in the Z direction, scores are calculated for the angle on the XY plane, the angle on the XZ plane, and the angle on the XY plane. This method has the advantage of being able to change the importance of coordinates and orientation, allowing for flexible settings and changes to suit the system configuration and shooting purpose. This makes it possible to select image capture devices whose coordinates and orientation are similar to those of other image capture devices, making it possible to reduce recorded data while minimizing degradation in image quality.
[0067] Example 3 In this embodiment, an example will be described in which the image processing system 1 is used to determine which imaging device has the least impact on image quality based on overlapping imaging ranges, and data recording is stopped to extend the recording time of the virtual viewpoint image.
[0068] The configuration of the image processing system 1 according to this embodiment is the same as that of the first embodiment described above.
[0069] Furthermore, the recording data reduction flow shown in FIG. 4 is also the same.
[0070] In this embodiment, first, data for each imaging device is acquired in S403, as in the first embodiment. Next, in S404, the angle of view for each imaging device is calculated, and individuals (imaging devices) with a large overlap in angle of view with other imaging devices are selected. As an example of selection, FIG. 7 shows a top view of an arrangement of three imaging devices. Like FIG. 6, FIG. 7 actually includes information in the Z direction. The angle of view for each imaging device is calculated, and the area where the angle of view overlaps for each combination of imaging devices (actually, the volume, since information in the Z direction is included) is calculated. In other words, the size of the area where the imaging ranges overlap for each combination of imaging devices is calculated. The lower part of FIG. 7 shows the area where the angles of view overlap. The area of area where the respective angles of view overlap is largest for imaging device 702 & imaging device 703, followed by imaging device 701 & imaging device 703, and smallest for imaging device 701 & imaging device 702. In this embodiment, imaging device 703 has the largest combination of areas that overlap with other imaging devices. Therefore, imaging device 703 is the target for texture data reduction. This method can consider similarity based on the actual angle of view, which avoids the risk of eliminating image capture devices that have similar coordinates and orientations but significantly different angles of view due to different focal lengths. This allows us to select image capture devices whose actual angles of view are similar to other image capture devices, making it possible to reduce recorded data while minimizing the impact on image quality.
[0071] Example 4 In this embodiment, an example will be described in which the image processing system 1 is used to determine an image capture device with a large amount of data and stop data recording, thereby extending the recording time of virtual viewpoint images.
[0072] The configuration of the image processing system 1 according to this embodiment is the same as that of the first embodiment described above.
[0073] Furthermore, the recording data reduction flow shown in FIG. 4 is also the same.
[0074] In this embodiment, in S403, the recording setting acquisition unit 301 acquires a foreground ratio from the image processing device 20 to identify imaging devices with large texture data. The foreground ratio is the ratio of the foreground (the subject) to the entire image. A system that generates virtual viewpoint images using multiple imaging devices may include imaging devices that zoom in on specific areas to improve texture quality. Examples include an area where the subject is likely to be captured (the center of the imaging area) or a height where the subject's face is located (e.g., approximately 1.5 to 2 meters above the ground). Because the imaging device in question zooms in, it tends to have a high foreground ratio. Therefore, even with foreground separation, the foreground area is large, and the amount of texture data transmitted and recorded is inevitably relatively large. Therefore, in this embodiment, in S404, the foreground ratios of each imaging device are compared to identify imaging devices that meet the above-mentioned conditions. As a result, imaging devices with high foreground ratios (large texture data) are selected as targets for texture data reduction. This method has the advantage of reducing a large amount of data relative to the number of imaging devices to be reduced. Furthermore, as in the method of Example 1, coordinate information, attitude information, sensor size, and lens hardware information are acquired along with the foreground ratio in S403, and in S404 the foreground ratios are compared and the angle of view of each imaging device is confirmed, thereby effectively reducing the amount of recorded data while minimizing the impact on image quality.
[0075] <Example 5> In this embodiment, an example will be described in which the image processing system 1 is used to determine an imaging device that does not affect the virtual viewpoint image and stop data recording, thereby extending the recording time of the virtual viewpoint image.
[0076] The configuration of an image processing system 1 according to this embodiment is shown in Figure 8. The difference in configuration from Figure 1 described above is the information processing device 800, which additionally acquires virtual viewpoint information from the virtual viewpoint input device 50 and uses it to select targets for texture data reduction. Figure 9 shows the configuration of an information processing device 100. The difference in configuration from Figure 3 described above is the reduction device selection unit 902. The reduction device selection unit 902 additionally acquires virtual viewpoint information from the virtual viewpoint input device 50. The virtual viewpoint information shown here is at least three-dimensional coordinate information and orientation information of the image capture device in the virtual space that generates the virtual viewpoint image. Figure 10 also shows the recorded data reduction flow after the configuration change.
[0077] In this embodiment, in step S1001, which is newly added, the reduction device selection unit 902 acquires virtual viewpoint information to be used in generating a virtual viewpoint image from the virtual viewpoint input device 50. In step S404, the virtual viewpoint image capture device, which captures an area not used in generating the virtual viewpoint image, is selected as a target for texture data reduction. FIG. 11 is a diagram showing the instantaneous relationship between a virtual imaging device and a subject in generating a virtual viewpoint image. In FIG. 11, for a virtual imaging device 1101, a rendering area 1102 is the area used in rendering the generated virtual viewpoint image and contains necessary texture data. On the other hand, a non-rendering area 1103 is an area not visible from the virtual imaging device 1101, and the texture data of that area does not contribute to the image quality of the generated virtual viewpoint image. Therefore, even if the texture data of the imaging device capturing the non-rendering area 1103 is reduced, the image quality does not deteriorate.
[0078] FIG. 12 illustrates a method for selecting an imaging device according to this embodiment. FIG. 12 is a top view showing a virtual imaging device 1201 and imaging devices 1202-1204 that are actually placed. Any imaging device with an angular difference of a certain amount or more relative to the orientation of the virtual imaging device 1201 is determined not to be used for generating a virtual viewpoint image. In this embodiment, imaging devices with an angular difference of 135 degrees or more are targeted for reduction, so imaging device 1203 is the reduction target. Data reduction during recording can be achieved by stopping data recording for imaging device 1203. While FIG. 12 compares only angles on the XY plane, more detailed imaging device selection may be performed by also comparing angles on the XZ and YZ planes. Furthermore, if the three-dimensional coordinates or orientation of virtual imaging device 1201 change over time, the imaging device to be reduced may be changed as appropriate. This makes it possible to reduce the recording data of imaging devices capturing areas that are not visible in the virtual viewpoint image that is ultimately output.
[0079] The present disclosure can also be realized by providing a program that implements one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more functions. [Explanation of symbols]
[0080] 10. Imaging device 20 Image processing device 30 Modeling Equipment 40 databases 100 Information processing device
Claims
1. an acquisition means for acquiring free space of a recording medium for recording a plurality of pieces of image data output from a plurality of image capture devices used to generate a virtual viewpoint image; a transmitting means for transmitting instruction information for preventing the image data output from the specific imaging device from being recorded on the recording medium when the available capacity becomes equal to or less than a threshold value; An information processing device comprising:
2. the instruction information is information indicating an instruction not to output the image data generated from the specific imaging device, 2. The information processing apparatus according to claim 1, wherein the transmitting means transmits the instruction information to the specific image capturing device.
3. the instruction information is information indicating an instruction to stop generating the image data, 2. The information processing apparatus according to claim 1, wherein the transmitting means transmits the instruction information to the specific image capturing device.
4. the instruction information is information indicating an instruction not to record the image data generated by the specific imaging device on the recording medium, 2. The information processing apparatus according to claim 1, wherein said transmitting means transmits said instruction information to said recording medium.
5. 2. The information processing apparatus according to claim 1, wherein among the plurality of image capturing devices, an image capturing device other than the specific image capturing device outputs the image data even when the free space falls below a threshold value.
6. The information processing apparatus according to claim 1 , wherein the specific image capturing device is determined based on the positions and orientations of the plurality of image capturing devices.
7. 7. The information processing apparatus according to claim 6, wherein the specific imaging device is an imaging device among the plurality of imaging devices whose imaging range overlaps with a large area of other imaging devices.
8. a first calculation means for calculating a first maximum angle, which is the maximum angle among a plurality of angles formed by lines connecting a coordinate on an XY plane with the coordinates of the plurality of image capturing devices; a second calculation means for calculating, for each of the plurality of image capture devices, a second maximum angle that is the largest among a plurality of angles formed by a line connecting a certain coordinate and the remaining image capture devices, excluding one image capture device, thereby calculating the second maximum angle when only one of the plurality of image capture devices is excluded; 7. The information processing apparatus according to claim 6, wherein the specific imaging device is an imaging device that is excluded from the plurality of imaging devices when the second maximum angle is smaller than the first maximum angle.
9. a calculation means for calculating a similarity score in each of the plurality of image capture devices, the similarity score indicating a degree of similarity between the image capture device and the other image capture devices in terms of three-dimensional coordinates and orientation; The information processing apparatus according to claim 6 , wherein the specific image capturing device is an image capturing device having a large similarity score among the plurality of image capturing devices.
10. The specific imaging device is a plurality of imaging devices, 2. The information processing apparatus according to claim 1, further comprising a determination unit that determines the number of the specific image capturing devices based on the free space.
11. an acquisition step of acquiring free space of a recording medium for recording a plurality of pieces of image data output from a plurality of image capture devices used to generate a virtual viewpoint image; a transmitting step of transmitting an instruction not to output image data to a specific image capture device among the plurality of image capture devices when the available capacity is equal to or less than a threshold value; An information processing method comprising:
12. 11. A computer program for controlling each means of the information processing apparatus according to claim 1 by a computer.
Citation Information
Patent Citations
Drive recorder and video recording method for drive recorder
JP2012019450A