Good imaging position estimation device
The system addresses the inefficiency of manual camera position scoring by modeling goodness differences between adjacent positions, enabling automatic detection of optimal camera movements for improved image quality in 3D spaces.
Patent Information
- Application Number
- JP2023204588
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-06-16
AI Technical Summary
Existing methods for finding optimal camera positions in 3D spaces for capturing good images are time-consuming and require manual calculation of scores for all viewpoint positions.
A system that models the difference in goodness between adjacent camera positions, allowing for the calculation of the direction to move the camera for improved image quality and efficiently searching for a preferable viewpoint.
Automatically detects good viewpoint positions by determining the direction of camera movement for improved image quality, significantly reducing the time and effort required in video production.
Smart Images

Figure 2025089757000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an apparatus for estimating the position of a camera capable of capturing a preferable image when capturing an image of an object in a three-dimensional space.
Background Art
[0002] In video production, the use of 3DCG (3D computer graphics) is increasing. When trying to manually search for the position of a virtual camera that can capture a good image in a three-dimensional space, it takes a lot of time and effort. Therefore, automatically detecting the optimal camera position (in other words, "viewpoint") for shooting a scene in a three-dimensional space is important for efficiently producing three-dimensional content.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
[0004] Non-Patent Document 1 discloses a method of controlling a robot camera by learning the movement of a cameraman. Non-Patent Document 2 discloses extracting features that affect the quality of video by hand.
Summary of the Invention
Problems to be Solved by the Invention
[0005] Conventionally, it was possible to score the goodness as a viewpoint, but it was difficult to predict in which moving direction a good viewpoint was from the current camera position. For this reason, it was necessary to calculate scores for all viewpoint positions to estimate a good viewpoint.
[0006] An object of the present invention is to provide an apparatus that can automatically search for a good viewpoint position.
Means for Solving the Problems
[0007] The inventors of the present invention have found that by modeling the difference in goodness between adjacent shooting positions, it is possible to calculate in which direction the camera should be moved to capture a better video, and efficiently search or detect a preferable viewpoint in a scene, and thus have completed the present invention.
[0008] (1) The first good shooting position estimation apparatus according to the present invention is a good shooting position estimation apparatus that searches for a camera position suitable for shooting. When an image from a certain camera position is input, it outputs a score indicating the goodness of the image and a parameter indicating the increment of the score when the camera position is moved by a predetermined amount in each coordinate axis direction of the coordinate system. It includes a camera spherical coordinate generation unit that determines the coordinate axis direction for moving the camera position based on the parameter, and each coordinate axis direction includes both the direction in which the value of each coordinate increases and the direction in which it decreases.
[0009] (2) In the first good shooting position estimation apparatus, the coordinate system may be a spherical coordinate system.
[0010] (3) In the first good shooting position estimation device, the learning model may be a learning model that is trained using, as learning data, a set of an image from a certain camera position, a score for the image, and an increment of the score when the camera position is moved by a predetermined amount in each coordinate axis direction.
[0011] (4) The first good shooting position estimation device may repeat the process of moving the camera position according to the parameter and obtaining the parameter for the moved camera position from the learning model until the parameter indicates that a better camera position cannot be obtained regardless of the direction of movement along any coordinate axis.
[0012] (5) When the position obtained by moving the camera position according to the parameter is a position that has been moved before, the first good shooting position estimation device may determine, as the good camera position, the camera position with the highest score among the camera positions that have been moved before.
Advantages of the Invention
[0013] According to the present invention, when an image from a certain camera position is input, a learning model that outputs a score indicating the quality of the image and a parameter indicating an increment of the score when the camera position is moved by a predetermined amount in each coordinate axis direction of the coordinate system can be used to automatically search for a good viewpoint position.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Embodiments for Carrying Out the Invention
[0015] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In this specification, the expressions "good viewpoint" and "preferred viewpoint" may be used. The "good viewpoint" and "preferred viewpoint" refer to the viewpoint when the image taken from that viewpoint is a preferred image. Also, when the image taken from a certain viewpoint is an unfavorable image, that viewpoint is referred to as a "bad viewpoint" and an "unfavorable viewpoint". Also, in this specification, the terms "camera position" and "viewpoint position" may be used, but both mean the same thing.
[0016] FIG. 1 is a diagram for explaining whether a viewpoint is good or bad. FIG. 1(A) is an image taken from a preferred viewpoint because the shape when looking at the table from above, the number and arrangement of chairs can be recognized. On the other hand, in FIG. 1(B), the shape when looking at the table from above is unclear, there are many overlaps between objects, and it is unclear whether there are chairs on the other side of the table. Therefore, it can be said that FIG. 1(B) is an image taken from an unfavorable viewpoint.
[0017] FIG. 2 is a diagram showing the coordinate system used in the embodiment of the present invention. In the embodiment of the present invention, a spherical coordinate system is used. An object to be photographed is placed at the origin of the spherical coordinate system. The position coordinates of the camera are represented by (r, θ, φ) in the spherical coordinate system. The camera is assumed to photograph the object to be photographed located at the origin of the spherical coordinate system.
[0018] FIG. 3 is a diagram showing the hardware configuration of the good photographing position estimation device 1000 according to the embodiment of the present invention. In FIG. 3, the CPU 301 uses the RAM 302 as a work memory, executes programs stored in the ROM 303 and the hard disk drive 308, and controls the operations of the respective blocks described later via the system bus 312.
[0019] The HDD interface (hereinafter, the interface is abbreviated as "I / F") 304 connects secondary storage devices such as the HDD 308 and an optical disk drive. The HDD I / F 304 is, for example, an I / F such as Serial ATA (SATA). The CPU 301 can read data from the HDD 308 and write data to the HDD 308 via the HDD I / F 304. Further, the CPU 301 can expand the data stored in the HDD 308 to the RAM 302, and conversely, can also save the data expanded in the RAM 302 to the HDD 308. Then, the CPU 301 can execute the data expanded in the RAM 302 as a program.
[0020] The input I / F 305 connects input devices 309 such as a keyboard, a mouse, a digital camera, and a scanner. The input I / F 305 is, for example, a serial bus such as USB or IEEE1394. The CPU 301 can read various data such as a photographed image from the input device 309 via the input I / F 305.
[0021] The output I / F 306 connects the good shooting position estimation device 1000 and the display which is the output device 310. The output I / F 306 is a video output I / F such as DVI or HDMI (registered trademark), for example. The CPU 301 can send data to the display via the output I / F 306 and display a desired video on the display. The network I / F connects the good shooting position estimation device 1000 and the external server 311.
[0022] FIG. 4 is a diagram showing the software configuration of the good shooting position estimation device 1000 according to an embodiment of the present invention. A learned model is stored in the feature amount database 410. This learned model outputs seven types of parameters as output data when an image taken from a certain viewpoint is input as input data. Hereinafter, the seven types of parameters will be described as parameters a to g.
[0023] (1) Parameter a Parameter a is a quantity representing the quality of an image taken from a certain viewpoint. Parameter a may also be referred to as "score" or "Score".
[0024] (2) Parameter b Parameter b is the difference between the score of the image taken from the current viewpoint and the score of the image taken from the viewpoint that has moved a predetermined amount in the direction in which the coordinate value r increases. When the score of the image taken from the viewpoint that has moved a predetermined amount in the direction in which r increases is greater than the score of the image taken from the current viewpoint, parameter b takes a positive value. Parameter b may also be represented by "r+".
[0025] (3) Parameter c Parameter c is the difference between the score of the image taken from the current viewpoint and the score of the image taken from the viewpoint that has moved a predetermined amount in the direction in which the coordinate value r decreases. When the score of the image taken from the viewpoint that has moved a predetermined amount in the direction in which r decreases is greater than the score of the image taken from the current viewpoint, parameter c takes a positive value. Parameter c may also be represented by "r-".
[0026] (4) Parameter d Parameter d is the difference between the score of the image taken from the current viewpoint and the score of the image taken from the viewpoint that has moved a predetermined amount in the direction in which the coordinate value θ increases. When the score of the image taken from the viewpoint that has moved a predetermined amount in the direction in which θ increases is greater than the score of the image taken from the current viewpoint, parameter d takes a positive value. Parameter d may also be represented as "θ+".
[0027] (5) Parameter e Parameter e is the difference between the score of the image taken from the current viewpoint and the score of the image taken from the viewpoint that has moved a predetermined amount in the direction in which the coordinate value θ decreases. When the score of the image taken from the viewpoint that has moved a predetermined amount in the direction in which θ decreases is greater than the score of the image taken from the current viewpoint, parameter e takes a positive value. Parameter e may also be represented as "θ-".
[0028] (6) Parameter f Parameter f is the difference between the score of the image taken from the current viewpoint and the score of the image taken from the viewpoint that has moved a predetermined amount in the direction in which the coordinate value φ increases. When the score of the image taken from the viewpoint that has moved a predetermined amount in the direction in which φ increases is greater than the score of the image taken from the current viewpoint, parameter f takes a positive value. Parameter f may also be represented as "φ+".
[0029] (7) Parameter g Parameter g is the difference between the score of the image taken from the current viewpoint and the score of the image taken from the viewpoint that has moved a predetermined amount in the direction in which the coordinate value φ decreases. When the score of the image taken from the viewpoint that has moved a predetermined amount in the direction in which φ decreases is greater than the score of the image taken from the current viewpoint, parameter g takes a positive value. Parameter g may also be represented as "φ-".
[0030] When inputting an image into the learned model, the image adjustment unit 420 adjusts the image to an appropriate size. The good viewing point calculation unit 430 inputs the image whose size has been adjusted by the image adjustment unit 420 into the learned model stored on the feature amount database 410, and acquires parameters a to g.
[0031] The viewing point score output unit 440 records the parameters a to g acquired by the good viewing point calculation unit 430 in the viewing point score recording unit 450. The viewing point score recording unit 450 records the parameters a to g together with the coordinate information of the viewing point.
[0032] The camera spherical coordinate generation unit 460 generates a coordinate position at which a better image can be captured according to the parameters a to g acquired by the good viewing point calculation unit 430. When two or more of the parameters b to g have positive values, one of them is selected. The image generation unit 470 generates an image captured from the coordinate position generated by the camera spherical coordinate generation unit 460.
[0033] The reprocessing calculation unit 480 determines whether the coordinate position of the next camera created by the camera spherical coordinate generation unit 460 is a position where it has moved before. When the coordinate position of the next camera created by the camera spherical coordinate generation unit 460 is not a position where it has moved before, the image generated by the image generation unit 470 is resized by the image adjustment unit 420 and input into the learned model on the feature amount database 410. When the coordinate position of the next camera created by the camera spherical coordinate generation unit 460 is a position where it has moved before, a process of determining the viewing point position with the highest score is performed based on the record of the viewing point score recording unit 450.
[0034] FIG. 5 is a diagram showing the creation of learning data for the learning model of the embodiment of the present invention. For images taken on the spherical coordinate system (r, θ, φ) centered on the shooting target in various scenes, the quality of the images is evaluated and scored. For example, divide all the images into groups each composed of six images, and in each group, let a human select the best image. According to the selection results, quantify the quality of each image.
[0035] FIG. 6 is a diagram showing learning data in an embodiment of the present invention. By the process of FIG. 5, parameter a is determined for each image. In addition, based on the position coordinates of the viewpoint at which the image was taken and parameter a, the difference in parameter a for adjacent coordinate positions can be obtained, so that parameters b to g can also be determined for each image. Use the set of the thus obtained images and parameters a to g as learning data, and generate a learning model using a machine learning algorithm such as a neural network.
[0036] FIG. 7 is a diagram showing an example of data recorded in the viewpoint score recording unit 450. In the viewpoint score recording unit 450, together with the spherical coordinates of the camera, the values of seven parameters for the image taken from that position are recorded. If the next camera coordinate position created by the camera spherical coordinate generation unit 460 is not a position where it has moved before, the reprocessing calculation unit 480 adjusts the size of the image generated by the image generation unit 470 with the image adjustment unit 420 and inputs it to the learned model on the feature amount database 410. Therefore, the generation of a better coordinate position in the camera spherical coordinate generation unit 460 may be performed multiple times. Also, it is not essential to record the image data in the viewpoint score recording unit 450, but it may be recorded.
[0037] FIG. 8 is a flowchart showing an example of the processing of the good shooting position estimation device 1000 according to an embodiment of the present invention. In step S801, an image from the current viewpoint position is input to the good shooting position estimation device 1000. The image adjustment unit 420 adjusts the size of the image.
[0038] In step S802, the image whose size has been adjusted by the image adjustment unit 420 is input into the learning model on the feature amount database 410. As the output from the learning model, seven parameters a to g are obtained. In step S803, it is determined whether any of the parameters b to g are positive. If all of the parameters b to g are zero or negative, since the current viewpoint position is optimal, the process moves to step S808. In step S808, the current viewpoint position is determined as the optimal viewpoint position. If there is a positive value among the parameters b to g in step S803, the process proceeds to step S804.
[0039] In step S804, the viewpoint position is moved by a predetermined amount in any direction among one or more positive values. In step S805, an image from the moved viewpoint position is generated, for example, by computer graphics. In step S806, the reprocessing calculation unit 480 determines whether the current camera position is a position that has been moved to before. If the current camera position is not a position that has been moved to before, the process proceeds to step S801 and the image generated in step S805 is input. If the current camera position is a position that has been moved to before, the process proceeds to step S807.
[0040] In step S807, the camera position corresponding to the parameter a with the largest value among one or more parameters a recorded in the viewpoint score recording unit 450 is determined. In step S809, an image from the determined camera position is generated, for example, by computer graphics.
[0041] FIG. 9 is an example of the video displayed on the display of the output device 310. The numerical value displayed at the position of "Score" in Fig. 9 is the score value. The numerical value displayed at the position of "Front" in Fig. 9 is the value of r+. The numerical value displayed at the position of "Back" in Fig. 9 is the value of r-. The numerical value displayed at the position of "Right" in Fig. 9 is the value of θ+. The numerical value displayed at the position of "Left" in Fig. 9 is the value of θ-. The numerical value displayed at the position of "Up" in Fig. 9 is the value of φ+. The numerical value displayed at the position of "Down" in Fig. 9 is the value of φ-.
[0042] Input the image of the virtual camera in the computer graphics space into the learning model to obtain seven parameters a~g. When the values of parameters b~g contain positive values, move in the positive direction by a predetermined amount. After the movement, input the image of the camera again to obtain the parameters. Here, when no positive number is included in parameters b~g, set the current position as the viewpoint suitable for shooting. Also, when the destination of the movement is a viewpoint position where there has been a movement in the past, set the position with the highest score among the past movement viewpoint positions as the viewpoint position suitable for shooting.
[0043] In Fig. 9(B), since the value of θ+ is positive, the viewpoint position will move by a predetermined amount in the direction in which θ increases (a). And when it becomes impossible to expect an improvement in the score (c), it will be located at a good viewpoint.
[0044] Since the present invention can detect the viewpoint position suitable for shooting based on the image information, it is possible to detect and provide the viewpoint position suitable for shooting an object based on the video of the virtual camera in the 3DCG space or the aerial video in the actual shooting scene. Since there is no need to manually search for the viewpoint position suitable for shooting in a vast scene, the efficiency of video production can be improved.
Explanation of Signs
[0045] 301 CPU 302 RAM 303 ROM 304 HDD Interface 305 Input Interface 306 Output Interface 307 Network Interface 308 HDD 309 Input Device 310 Output Device 311 External Server 312 System Bus 410 Feature Database 420 Image Adjustment Unit 430 Good Viewpoint Calculation Unit 440 Viewpoint Score Output Unit 450 Viewpoint Score Recording Unit 460 Camera Spherical Coordinate Generation Unit 470 Image Generation Unit 480 Reprocessing Calculation Unit 1000 Good Shooting Position Estimation Device
Claims
1. A good shooting position estimation device for searching for a camera position suitable for shooting, When an image from a certain camera position is input, a learning model that outputs a score indicating the quality of the image and parameters indicating the increment of the score when the camera position is moved by a predetermined amount in each coordinate axis direction of the coordinate system, A camera spherical coordinate generation unit that determines the coordinate axis direction for moving the camera position based on the parameters, The good shooting position estimation device, wherein each of the coordinate axis directions includes both the direction in which the value of each coordinate increases and the direction in which it decreases.
2. The good shooting position estimation device according to claim 1, wherein the coordinate system is a spherical coordinate system.
3. The learning model of the good shooting position estimation device according to claim 1 is a learning model that is trained using, as learning data, a set of an image from a certain camera position, a score for the image, and an increment of the score when the camera position is moved by a predetermined amount in each coordinate axis direction.
4. The process of moving the camera position according to the parameters and obtaining the parameters for the moved camera position from the learning model is repeated until the parameters indicate that moving in any coordinate axis direction will not result in a better camera position, for the good shooting position estimation device according to claim 1.
5. When the position where the camera position is moved according to the parameters is a position that has been moved before, the good shooting position is determined as the camera position with the highest score among the camera positions that have been moved before, for the good shooting position estimation device according to claim 4.
Citation Information
Patent Citations
JP1145355008A