A good viewpoint image selection device, and a program to make a computer function as a good viewpoint image selection device.
The device and program analyze scenes to determine optimal viewpoints for multiple objects, improving image selection efficiency by generating high-quality images or videos from complex scenes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-08-04
- Publication Date
- 2026-04-08
AI Technical Summary
Existing technologies lack a method for automatically detecting the optimal viewpoint in complex scenes containing multiple objects, making the selection of desired images time-consuming.
A device and program that analyze a scene, extract feature quantities, calculate the centroid position of attention-attracting objects, evaluate viewpoint quality using a feature database, and generate an image from the optimal viewpoint based on camera center and sphere coordinates.
Enables efficient detection of the optimal viewpoint in complex scenes, allowing for the automatic generation of high-quality images or videos from multiple viewpoints, enhancing production efficiency in 3D image content creation.
Smart Images

Figure 0007842659000001 
Figure 0007842659000002 
Figure 0007842659000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus for selecting an image from a preferred viewpoint, and a program therefor. [Background technology]
[0002] Traditionally, volumetric capture technology has made it possible to measure real-world scenes as 3D data. By placing virtual cameras in the captured scene, it is possible to generate images from any arbitrary position afterward. However, selecting the desired image on a computer screen is time-consuming. Therefore, automatically detecting and providing the optimal viewpoint for capturing a scene with a virtual camera in 3D space is crucial for the efficient production of 3D image content.
[0003] Several techniques have been proposed to detect the optimal viewpoint for a single object (for example, Non-Patent Document 1). However, a technique for detecting the optimal viewpoint in an environment where multiple objects exist in a scene has not yet been established. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] PPVa (the preceding letter is an acute 'a')zquez, M. Feixas, M. Sbert, and W. Heidrich: “Viewpoint Selection using Viewpoint Entropy”, In Proceedings of the Vision Modeling and Visualization Conference, Stuttgart, Germany, 21-23 November 2001; pp. 273-280 [Non-Patent Document 2] Xiaodi Hou and Liqing Zhang:”Saliency Detection: A Spectral Residual Approach”, 2007 IEEE Conference on Computer Vision and Pattern Recognition, 17-22 June 2007 [Overview of the project] [Problems that the invention aims to solve]
[0005] In complex scenes containing multiple objects, the technology for detecting the optimal viewpoint has not yet been established.
[0006] The present invention aims to provide a device that detects the optimal viewpoint in a complex scene containing multiple objects and generates an image viewed from that detected optimal viewpoint. [Means for solving the problem]
[0007] The inventors of this invention have determined which object to place at the center of an image based on the feature quantities of multiple objects, evaluated the quality of the image from each viewpoint using the importance of the feature quantities, and found that the viewpoint with the best image quality is the optimal viewpoint, thus completing the present invention.
[0008] While the adjective "good" is generally used for images, in this specification, it may be used for viewpoints to mean "the image viewed from that viewpoint is a good image" or "the image viewed from that viewpoint is a desirable image."
[0009] (1) The good viewpoint image selection device according to the present invention comprises: a scene analysis unit that analyzes a scene from which an optimal viewpoint is to be selected and decomposes it into a plurality of objects; a feature extraction unit that extracts a plurality of types of feature quantities from each object; a focus object calculation unit that calculates an object that is likely to attract attention using some of the feature quantities among the plurality of types of feature quantities extracted by the feature extraction unit; a camera center coordinate generation unit that generates camera center coordinates, which are the coordinates of the centroid position of the object that is likely to attract attention calculated by the focus object calculation unit; a feature database that stores the importance of each type of feature quantity for selecting a preferred viewpoint; a good viewpoint calculation unit that calculates the quality of observation positions of a plurality of points on the surface of a sphere centered on the camera center coordinates using the feature quantities extracted by the feature extraction unit and the feature database; a camera sphere coordinate generation unit that selects a good viewpoint as an observation position from among the plurality of points on the surface of a sphere centered on the camera center coordinates and registers camera sphere coordinates, which are the coordinates of the good viewpoint as an observation position; and a good image generation unit that generates a camera image from the good viewpoint as an observation position based on the camera center coordinates and the camera sphere coordinates.
[0010] (2) The quality of the observation position is evaluated by the amount of information in the image viewed from the observation position, and the amount of information in the image viewed from the observation position may include an amount obtained by multiplying the amount of information for each object by the importance of the amount of information for each object and adding the results for multiple objects.
[0011] (3) The objects that are likely to attract attention may be calculated based on the notability of the objects.
[0012] (4) The importance of the features stored in the feature database may be determined by multiple regression analysis such that the goodness of the observation location calculated by the good viewpoint calculation unit is most highly correlated with the data on human preference for the observation location obtained by experiment.
[0013] (5) The program according to the present invention comprises a computer, a scene analysis unit that analyzes a scene from which an optimal viewpoint is to be selected and decomposes it into multiple objects, a feature extraction unit that extracts multiple types of feature quantities from each object, a focus object calculation unit that calculates an object that is likely to attract attention using some of the multiple types of feature quantities extracted by the feature extraction unit, a camera center coordinate generation unit that generates camera center coordinates which are the coordinates of the centroid position of the object that is likely to attract attention calculated by the focus object calculation unit, and feature data in which the importance of each type of feature quantity for selecting a preferred viewpoint is stored. This device is designed to function as a good viewpoint image selection device, comprising: a base; a good viewpoint calculation unit that calculates the quality of observation positions of multiple points on the surface of a sphere centered on the camera center coordinates using the feature quantities extracted by the feature quantity extraction unit and the feature quantity database; a camera sphere coordinate generation unit that selects a good viewpoint from among the multiple points on the surface of a sphere centered on the camera center coordinates and registers the camera sphere coordinates, which are the coordinates of the good viewpoint; and a good image generation unit that generates a camera image from the good viewpoint based on the camera center coordinates and the camera sphere coordinates.
[0014] (6) In the program according to the present invention, the quality of the observation position is evaluated by the amount of information in the image viewed from the observation position, and the amount of information in the image viewed from the observation position may include an amount obtained by multiplying the amount of information for each object by the importance of the amount of information for each object and adding the results for multiple objects.
[0015] (7) In the program according to the present invention, the object that is likely to attract attention may be calculated based on the notability of the object.
[0016] (8) In the program according to the present invention, the importance of the features stored in the feature database may be determined by multiple regression analysis such that the goodness of the observation location calculated by the good viewpoint calculation unit is most highly correlated with the data on human preference for the observation location obtained by experiment.
Advantages of the Invention
[0017] According to the present invention, in a situation where a plurality of objects exist, an image from an optimal viewpoint can be obtained.
Brief Description of the Drawings
[0018] [Figure 1] It is a diagram illustrating a preferred viewpoint and a non-preferred viewpoint of the present invention. [Figure 2] It is a diagram showing the configuration of a good viewpoint image selection device according to an embodiment of the present invention. [Figure 3] It is a diagram showing the processing procedure of a good viewpoint image selection device according to an embodiment of the present invention. [Figure 4] It is a diagram showing the flow of a method for selecting a good viewpoint of a good viewpoint image selection device according to an embodiment of the present invention. [Figure 5] It is a diagram for explaining the camera center coordinates of the present invention. [Figure 6] It is a diagram illustrating the feature amount database of the present invention. [Figure 7] It is a diagram for explaining the camera spherical coordinates of the present invention. [Figure 8] It is a diagram showing an example of a method for evaluating the goodness as an observation position of the present invention. [Figure 9] It is a diagram showing an example of a method for evaluating the goodness as an observation position of the present invention.
Modes for Carrying Out the Invention
[0019] Hereinafter, an example of an embodiment of the present invention will be described. FIG. 1 is a diagram showing an example of an image viewed from a preferred viewpoint and an image viewed from a non-preferred viewpoint in the present invention.
[0020] In the image on the right side of Figure 1, which is viewed from an unfavorable viewpoint, no information about the depth of the table is obtained. The pattern on the tabletop is also unclear. Furthermore, a large area is obscured by the chair in the foreground. In contrast, the image on the left side discloses more information than the image on the right side. In this invention, an image that discloses more information in this way is determined to be an image from a favorable viewpoint.
[0021] Figure 2 shows the configuration of the good viewpoint image selection device 200 according to this embodiment. The CG scene 100 is input to the good viewpoint image selection device 200, and the good viewpoint image selection device 200 outputs an image 300 from a preferred viewpoint. The good viewpoint image selection device 200 comprises a scene analysis unit 10, a feature extraction unit 20, a target of interest calculation unit 30, a camera center coordinate generation unit 40, a feature database 50, a good viewpoint calculation unit 60, a camera sphere coordinate generation unit 70, and a good image generation unit 80.
[0022] The scene analysis unit 10 analyzes the input scene 100 and breaks it down into multiple objects. The feature extraction unit 20 extracts multiple types of features for each object. Examples of multiple types of features include volume and sampling. The object of interest calculation unit 30 uses some of the features extracted by the feature extraction unit 20 to calculate objects that are likely to attract attention. For example, it uses sampling to calculate objects that are likely to attract attention.
[0023] The camera center coordinate generation unit 40 sets the position of the center of gravity of an object that is likely to attract attention, calculated by the object of interest calculation unit 30, as the center coordinate of the camera subject. Here, we will explain this "center coordinate of the camera subject" using Figure 7. In Figure 7, we assume a sphere with its center at the position of the center of gravity of an object that is likely to attract attention. Then, we place a virtual camera on the surface of this sphere and take a picture in the direction of the center of the sphere. In this case, even if we change the position of the virtual camera on the surface of the sphere in various ways, the object that is likely to attract attention will always be in the center of the captured image. Thus, the aforementioned "center coordinate of the camera subject" means the coordinate of the position of the center of gravity of the object that is in the center of the image captured by the virtual camera. Hereafter, this "center coordinate of the camera subject" will be called the "camera center coordinate". The feature database 50 stores the importance of each feature for selecting a preferred viewpoint across multiple types of features. These feature importance values can sometimes be negative.
[0024] The good viewpoint calculation unit 60 assumes multiple viewpoints on a sphere centered on the camera's central coordinates, and calculates the desirability of the image viewed from each viewpoint using multiple types of features extracted by the feature extraction unit 20 and the feature database 50. This calculation calculates the amount of information in the image viewed from each viewpoint. The camera spherical coordinate generation unit 70 sets the viewpoint that has the most information calculated by the good viewpoint calculation unit 60 as the spherical coordinate representing the camera's position. This "spherical coordinate representing the camera's position" will be referred to as the "camera spherical coordinate" from now on.
[0025] The good image generation unit 80 generates a camera image based on the camera center coordinates and camera sphere coordinates, assuming the camera is positioned at the camera sphere coordinates and the camera is photographed in the direction of the camera center coordinates, and outputs it as a good viewpoint image.
[0026] Figure 3 shows the processing procedure of the good viewpoint image selection device according to this embodiment. Figure 4 shows the flow of the method for selecting a good viewpoint using the good viewpoint image selection device according to this embodiment. In step S101, the CG scene 100 to be displayed is input to the good viewpoint image selection device 200. Figure 4(a) shows an example of the input CG scene 100. In step S102, the scene analysis unit 10 analyzes the CG scene 100 and breaks it down into multiple objects.
[0027] In step S103, the feature extraction unit 20 extracts multiple types of features, such as volume and sampling, for each object. In step S104, the feature extraction unit 20 records the features for each object. Furthermore, in step S104, the object of interest calculation unit 30 uses some of the extracted features to calculate objects that are likely to attract attention. For example, it uses spleness to calculate objects that are likely to attract attention. The creation of a spleness map is described in Non-Patent Literature 2. Figures 4(b) and 4(c) illustrate the creation of the spleness map described in Non-Patent Literature 2. In Figure 4(b), areas with high spleness are displayed with high brightness. By transferring that brightness to the objects, Figure 4(c) is obtained. In Figure 4(c), objects with high spleness are displayed in white, making it possible to identify objects with high spleness. In this example, the tower in the upper right of Figure 4(c) was determined to be an object with high spleness.
[0028] In step S105, the camera center coordinate generation unit 40 sets the position of the center of gravity of an object that is likely to attract attention as the camera center coordinate. Figure 5 is a diagram illustrating the camera center coordinate. In the center of Figure 5 are an object that is likely to attract attention: a sphere, a cone, and a rectangular prism. The position of the center of gravity of these objects that are likely to attract attention is indicated by an arrow, and the coordinates of the position of the arrow are the "camera center coordinate". The virtual camera is positioned on a sphere centered on the camera center coordinate, as shown in Figure 5.
[0029] In step S106, the good viewpoint calculation unit 60 assumes multiple viewpoints on a sphere centered on the camera's center coordinates, and calculates the desirability of the image viewed from each viewpoint using multiple types of features extracted by the feature extraction unit 20 and the feature database 50. This calculation calculates the amount of information in the image viewed from each viewpoint. Details of how to calculate the amount of information in the image viewed from each viewpoint will be described later. Here, we will explain qualitatively using Figure 4(d). Figure A in Figure 4(d) has a lot of blank space and therefore little information. Figure B in Figure 4(d) has more information because information about various buildings can be obtained. Therefore, B is an image from a better viewpoint.
[0030] Figure 6 shows an example of a feature database. The importance of each feature is stored within it. Volume represents the volume of the object. Rotation represents the camera's elevation angle relative to the camera's center coordinates. Distance from camera represents the distance from the camera to the object. Distance from camera center coordinates represents the distance from the camera's center coordinates to the object. These are examples, and other features may be used.
[0031] Furthermore, in step S106, the camera sphere coordinate generation unit 70 sets the viewpoint with the most information calculated by the good viewpoint calculation unit 60 as the camera sphere coordinate. Figure 7 shows the relationship between the camera center coordinate and the camera sphere coordinate. The camera center coordinate is the coordinate of the center of gravity of an object that is likely to attract attention. The camera sphere coordinate is located on the surface of a sphere centered at the camera center coordinate. In step S107, the good image generation unit 80 generates and outputs an image from the camera sphere coordinate position. Figure 4(e) is an example of an image output from a good viewpoint when multiple objects are present.
[0032] Next, using Figures 8 and 9, we will explain how to calculate the information content of the image in this embodiment. Here, s' all,i is the amount of information in the image as seen from viewpoint i by the method of this embodiment, and s all,i This represents the amount of information in the image as seen from viewpoint i using the conventional method (Non-Patent Document 1).
[0033] In the conventional method, all objects in a scene are regarded as one object and calculated. Therefore, the interaction (such as overlapping) between each object is not reflected. Therefore, in the method of this embodiment, the information amount s of object j as seen from viewpoint i i,j is weighted by the importance w i,j and an equation obtained by adding the weighted values is used. Here, the importance w i,j can be decomposed into the product of the k-th feature amount ν of object j as seen from viewpoint i i,j,k and the importance α of the k-th feature amount k (Equation (2) in FIG. 8). In this embodiment, an equation obtained by substituting Equation (2) in FIG. 8 into Equation (1) is used to calculate the information amount of an image. Then, the coordinates of the viewpoint at which the obtained information amount of the image becomes high are determined as spherical coordinates of the camera.
[0034] The importance α of the feature amount stored in the feature amount database k is calculated using an equation obtained by substituting Equation (2) into Equation (1), where s′ all,i is the data s of the preference of a person for the viewpoint obtained by experiment exp (s becomes larger for better viewpoints and smaller for worse viewpoints. exp s becomes smaller for worse viewpoints. exp ) and is determined by multiple regression analysis so as to have the highest correlation.
[0035] Incidentally, the information amount s of object j as seen from viewpoint i i,j can be calculated, for example, as follows. If n faces of object j are shown in the image and the number of pixels in which the m-th face is shown is p m , and the total number of pixels in the image is p t , then [[ID=4 .]] s i,j =-Σ m=1 n {(p m / p t )log(p m / p t )} This can be calculated as follows. The above formula is based on the definition of entropy in information theory.
[0036] The optimal viewpoint image selection device according to the present invention can detect the optimal viewpoint in a virtual scene, making it possible to select and provide the best image from multiple cameras within the virtual scene. Furthermore, with the growing interest in technologies that measure the entire space in three dimensions, such as volumetric photography studios, it becomes possible to automatically extract superior expressions from the measured scene, thereby improving the efficiency of image production.
[0037] Although this embodiment describes images, it goes without saying that it is also applicable to video footage.
[0038] In this embodiment, the amount of information contained in the image viewed from the viewpoint was used as the measure of goodness of the observation position. However, other quantities may also be used as measures of goodness of the observation position, such as the number of objects included in the image viewed from the viewpoint, or the area occupied by the object image in the image viewed from the viewpoint. [Explanation of Symbols]
[0039] 10 Scene Analysis Department 20 Feature Extraction Unit 30. Calculation Unit of the Target of Attention 40 Camera center coordinate generation unit 50 Feature Databases 60. Good Viewpoint Calculation Unit 70 Camera Sphere Coordinate Generation Unit 80 Image Quality Generation Unit 200 Good viewpoint image selection device
Claims
1. The scene analysis unit analyzes the scene from which the optimal viewpoint is to be selected and breaks it down into multiple objects, A feature extraction unit that extracts multiple types of features from each object, A focus target calculation unit calculates an object that is likely to attract attention using some of the types of features extracted by the feature extraction unit, A camera center coordinate generation unit generates camera center coordinates, which are the coordinates of the centroid position of an object that is likely to attract attention, calculated by the aforementioned object of attention calculation unit. A feature database that stores the importance of each feature in selecting a preferred viewpoint, A good viewpoint calculation unit calculates the quality of observation positions for multiple points on the surface of a sphere centered on the camera's central coordinates using the features extracted by the feature extraction unit and the feature database. A camera sphere coordinate generation unit selects a good viewpoint from among the multiple points on the surface of a sphere centered on the camera's center coordinates, and registers the camera sphere coordinates, which are the coordinates of the good viewpoint. A good image generation unit generates a camera image from a good viewpoint as the observation position based on the camera center coordinates and the camera sphere coordinates, A good viewpoint image selection device equipped with the following features.
2. The goodness of the observation position is evaluated by the amount of information in the image viewed from the observation position, and the amount of information in the image viewed from the observation position includes an amount obtained by multiplying the amount of information for each object by the importance of the amount of information for each object and adding this amount for multiple objects, as described in claim 1.
3. The aforementioned object that is likely to attract attention is calculated based on the object's prominence, according to claim 1, a good viewpoint image selection device.
4. The importance of the features stored in the feature database is determined by multiple regression analysis such that the goodness of the observation location calculated by the good viewpoint calculation unit is most highly correlated with the data on human preference for the observation location obtained experimentally, as described in claim 1, for the good viewpoint image selection device.
5. Computers, The scene analysis unit analyzes the scene from which the optimal viewpoint is to be selected and breaks it down into multiple objects, A feature extraction unit that extracts multiple types of features from each object, A focus target calculation unit calculates an object that is likely to attract attention using some of the types of features extracted by the feature extraction unit, A camera center coordinate generation unit generates camera center coordinates, which are the coordinates of the centroid position of an object that is likely to attract attention, calculated by the aforementioned object of attention calculation unit. A feature database that stores the importance of each feature in selecting a preferred viewpoint, A good viewpoint calculation unit calculates the quality of observation positions for multiple points on the surface of a sphere centered on the camera's central coordinates using the features extracted by the feature extraction unit and the feature database. A camera sphere coordinate generation unit selects a good viewpoint from among the multiple points on the surface of a sphere centered on the camera's center coordinates, and registers the camera sphere coordinates, which are the coordinates of the good viewpoint. A good image generation unit generates a camera image from a good viewpoint as the observation position based on the camera center coordinates and the camera sphere coordinates, A program to function as a good viewpoint image selection device equipped with these features.
6. The program according to claim 5, wherein the quality of the observation position is evaluated by the amount of information in the image viewed from the observation position, and the amount of information in the image viewed from the observation position includes an amount obtained by multiplying the amount of information for each object by the importance of the amount of information for each object and adding the results for multiple objects.
7. The program according to claim 5, wherein the object that is likely to attract attention is calculated based on the notability of the object.
8. The program according to claim 5, wherein the importance of the features stored in the feature database is determined by multiple regression analysis such that the goodness of the observation location calculated by the good viewpoint calculation unit is most highly correlated with the data on human preference for the observation location obtained by experiment.
Citation Information
Patent Citations
Three-dimensional shape display system, three-dimensional shape display method and three-dimensional shape display program
JP2006171889A
Scene optimization system and method
JP2012504828A
Information processing apparatus, method for controlling information processing apparatus, and program
JP2015170116A
Viewpoint position calculation device, image generation device, viewpoint position calculation method, image generation method, viewpoint position calculation program, and image generation program
JP2016099665A
Program, information processing device, and method
JP2022112197A