Robot camera control device, program therefor, and imaging system

The robot camera control device addresses the impracticality of real-space viewpoint estimation by using three-dimensional scene recognition and position adjustment to maintain high-quality imaging of moving objects.

JP2025113527APending Publication Date: 2025-08-04NIPPON HOSO KYOKAI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024007727
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-08-04

AI Technical Summary

Technical Problem

Conventional methods require 3D information measurement and conversion for estimating high-quality viewpoints, which is impractical when objects move, and lack a method for real-space viewpoint estimation without human learning.

Method used

A robot camera control device that includes three-dimensional scene recognition, subject designation, subject position estimation, virtual space arrangement correction, and camera position control to estimate and maintain high-quality viewpoints for moving objects in real space.

Benefits of technology

Enables high-quality viewpoint estimation and imaging of moving objects in real space without prior learning, predicting future positions and adjusting camera positions accordingly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025113527000001_ABST
    Figure 2025113527000001_ABST
Patent Text Reader

Abstract

To provide a robot camera control apparatus capable of performing imaging at a viewpoint perceived by a human as having high quality using a robot camera.SOLUTION: A robot camera control apparatus 2 comprises: three-dimensional scene recognition means 20 that acquires point group data that is distance information in real space and associates the point group data with a virtual space; subject designation means 21 that displays the point group data as a plan view on a display device and receives designation of a subject; subject position estimation means 220 that tracks the subject and estimates a future position from a past position of the subject; virtual space arrangement correction means 221 that corrects a position of the point group data of the subject so as to arrange the subject at the future position; viewpoint estimation means 23 that estimates the viewpoint position and orientation of the virtual camera that optimizes a viewpoint entropy; and camera position control means 24 that controls the viewpoint position and the orientation of a robot camera in accordance with the estimated viewpoint position and the orientation of the virtual camera.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a robot camera control device, its program, and a photographing system.

Background Art

[0002] When a person observes an object (subject), it is known that there are common points in the viewpoints and compositions that many people feel are of high quality (see Non-Patent Document 1). Also, algorithms for estimating viewpoints that people feel are of high quality using a computer have been proposed (see Non-Patent Documents 2 and 3). When observing an object in this way, if the three-dimensional information of the object can be handled in a computer, it is possible to estimate viewpoints and compositions that are felt to be of high quality and are close to human perception with a certain degree of accuracy. Conventionally, methods for photographing an object from a viewpoint that people feel is of high quality in a complex scene composed of a large number of objects have been proposed (see Patent Document 1 and Non-Patent Document 4). Also, a method has been proposed in which the relationship between a cameraman's camera operation and two-dimensional video information is pre-learned by machine learning, and camera work is estimated from the captured video (see Patent Document 2).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Non-Patent Documents

[0004]

Non-Patent Document 1

[0005] Conventionally, when taking a photo from a viewpoint that people feel is of high quality, it is necessary to know the 3D information of the object to be photographed. That is, in the conventional method, it is necessary to measure the 3D information, convert it into data, and then estimate the viewpoint. However, when the object moves, it takes time to measure and estimate the viewpoint. Therefore, when the estimation is completed, the shooting situation, such as the position of the object, may change, and there is a problem of lack of practicality.

[0006] In addition, since the object in the invention described in Patent Document 1 is 3DCG (3D computer graphics), a shooting method for performing viewpoint estimation in the real space has been desired. In addition, since the invention described in Patent Document 2 requires learning of camera work by a skilled cameraman, a shooting method for performing viewpoint estimation without performing learning through human hands has been desired.

[0007] The present invention has been made in view of such problems and demands, and an object of the present invention is to provide a robot camera control device, a program thereof, and a shooting system capable of performing shooting from a viewpoint that a person feels to be of high quality in the real space using a robot camera.

Means for Solving the Problems

[0008] In order to solve the above problems, a robot camera control device according to the present invention is a robot camera control device for controlling a robot camera, and includes a three-dimensional scene recognition unit, a subject designation unit, a subject position estimation unit, a virtual space arrangement correction unit, a viewpoint estimation unit, and a camera position control unit.

[0009] In such a configuration, the robot camera control device acquires point cloud data, which is distance information of the real space, at a predetermined time interval by the three-dimensional scene recognition unit and associates it with a three-dimensional virtual space. Then, the robot camera control device displays the point cloud data of the virtual space on a display device by the subject designation unit, and accepts the designation of the subject when the operator externally instructs the subject to be photographed.

[0010] Then, the robot camera control device tracks the subject designated by the subject designation unit in the virtual space by the subject position estimation unit, and estimates the future position after a predetermined time has elapsed from the past positions of the subject in time series. Then, the robot camera control device corrects the position of the point cloud data that constitutes the subject so that the subject is arranged at a future position in the virtual space by the virtual space arrangement correction means. As a result, only the subject with movement will move and be arranged at a future position after a predetermined time has elapsed in the virtual space.

[0011] Then, the robot camera control device estimates the viewpoint position and orientation of the virtual camera that optimizes the viewpoint entropy in the virtual space by the viewpoint estimation means. As a result, the viewpoint estimation means can estimate a high-quality viewpoint on the virtual space. Then, the robot camera control device controls the viewpoint position and orientation of the robot camera in the real space corresponding to the viewpoint position and orientation of the virtual camera estimated by the viewpoint estimation means by the camera position control means. As a result, the camera position control means can reflect the viewpoint position and orientation of the virtual camera that becomes a high-quality viewpoint on the virtual space onto the real space and control the viewpoint position and orientation of the robot camera.

[0012] This robot camera control device may control a plurality of robot cameras. In this case, the robot camera control device estimates the viewpoint position and orientation of the virtual camera that is the global optimal solution of the viewpoint entropy as the viewpoint position and orientation of the first robot camera by the viewpoint estimation means, and sequentially estimates the viewpoint position and orientation of the second and subsequent robot cameras in order from the one with the higher importance among the plurality of local optimal solutions. Then, the robot camera control device controls the viewpoint position and orientation of the plurality of robot cameras in the real space corresponding to the viewpoint position and orientation of the plurality of virtual cameras estimated by the viewpoint estimation means by the camera position control means. Note that the robot camera control device can operate a computer with a program for causing the computer to function as the robot camera control device.

[0013] Also, in order to solve the above problems, the imaging system according to the present invention is an imaging system using a robotic camera, and is configured to include a robotic camera and a robotic camera control device.

[0014] In such a configuration, the imaging system is instructed by the robotic camera to obtain control information for specifying the viewpoint position and orientation of the camera in the real space, and drives the camera to be in that viewpoint position and orientation, and performs imaging with that camera. Also, the imaging system instructs the robotic camera control device to specify control information for specifying the viewpoint position and orientation of the camera that provides a high-quality viewpoint for the robotic camera.

[0015] Also, the imaging system may be configured to include a plurality of robotic cameras. In this case, the imaging system instructs the robotic camera control device to specify control information for specifying the viewpoint position and orientation of the camera for each of the plurality of robotic cameras in order of high quality for each of the plurality of robotic cameras.

Advantages of the Invention

[0016] According to the present invention, it is possible to estimate a viewpoint that a person feels has high quality for an object with movement in the real space and image the object.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Embodiments for Carrying Out the Invention

[0018] Hereinafter, embodiments of the present invention will be described with reference to the drawings. ≪Imaging System≫ With reference to FIG. 1, the configuration of an imaging system 100 according to an embodiment of the present invention will be described.

[0019] The imaging system 100 uses a robot camera 1 to image a subject to be imaged within an object existing in the real space. Here, it is assumed that a plurality of objects SO1, SO2, and MO exist in the real space. The imaging system 100 moves the robot camera 1 to a viewpoint where a person feels the quality is high and performs imaging, taking as the imaging target a subject that is a moving object specified among a plurality of objects SO1, SO2, and MO. Here, the objects SO1 and SO2 are stationary objects (structures), and the object MO is an object that moves (for example, a person). The imaging system 100 includes a robot camera 1 and a robot camera control device 2.

[0020] The robot camera 1 moves the position of the camera C in the real space according to an instruction from the robot camera control device 2 to image the subject. Although a general robot camera 1 can be used, here, it is configured to include a distance measuring device 11.

[0021] <Schematic Configuration of Robot Camera> For reference, the schematic configuration of the robot camera 1 will be described with reference to FIGS. 2 and 3. As shown in FIG. 2, the robot camera 1 includes a camera C, a carriage 10, a distance measuring device 11, a robot arm 12, and a control device 13.

[0022] The camera C is a general imaging device that captures a subject. The camera C is provided at the tip of the robot arm 12 and captures the subject in a posture corresponding to the movement of the robot arm 12.

[0023] The carriage 10 mounts the distance measuring device 11, the robot arm 12, and the control device 13, and moves by driving the wheels according to a drive signal output from the control device 13. The distance measuring device 11 measures the distance from the distance measuring device 11 to an object in the real space. The distance measuring device 11 can measure the distance by parallax using a stereo camera. Also, for example, the distance measuring device 11 can use a sensor such as LiDAR (Light Detection And Ranging) that measures the time from when light is reflected from an object by an optical sensor until it is received. For example, when a 3D (three-dimensional) LiDAR is used for the distance measuring device 11, the laser light irradiation range is set to approximately -105 degrees (left end) to +105 degrees (right end) in the horizontal direction (azimuth angle) and approximately -5 degrees (lower end) to +35 degrees (upper end) in the vertical direction (elevation angle), and the distance is measured in 0.25-degree units. The distance measuring device 11 performs measurements at a predetermined time interval (for example, 25 ms) and outputs the measured distance information to the control device 13.

[0024] The robot arm 12 drives the joints of each arm according to a drive signal output from the control device 13. For example, a 6-axis vertical articulated robot can be used as the robot arm 12. The robot arm 12 drives and controls each arm so that the camera C operates in the indicated position and posture.

[0025] The control device 13 transmits and receives information to and from the robot camera control device 2, and controls the robot camera 1 based on the information. As shown in FIG. 3, the control device 13 includes a distance information transmitting means 130, a carriage driving means 131, and a robot arm driving means 132.

[0026] The distance information transmitting means 130 transmits the distance information input from the distance measuring device 11 to the robot camera control device 2. The distance information transmitting means 130 transmits the distance information to the robot camera control device 2 via a communication control unit (not shown), for example, by Wi-Fi.

[0027] The carriage driving means 131 receives the carriage control information transmitted from the robot camera control device 2, and operates the carriage 10 to a specified position. The carriage driving means 131 receives the carriage control information via a communication control unit (not shown), generates a drive signal according to the carriage control information, and outputs it to the carriage 10.

[0028] The robot arm driving means 132 receives the robot arm control information transmitted from the robot camera control device 2, and operates the robot arm 12 so that the camera C assumes a specified viewpoint position and posture. The robot arm driving means 132 receives the robot arm control information via a communication control unit (not shown), generates a drive signal according to the robot arm control information, and outputs it to the robot arm 12. With the configuration described above, the robot camera 1 transmits the distance information to the robot camera control device 2, and photographs the subject at a specified position based on the carriage control information and the robot arm control information received from the robot camera control device 2. Returning to FIG. 1, the description of the configuration of the imaging system 100 will be continued.

[0029] The robot camera control device 2 controls the robot camera 1 to photograph a specified object MO as the subject. For example, when the object MO, which is the subject, moves from the position of A MO to the position of B MO the robot camera control device 2 predicts the movement of the object MO and moves the robot camera 1 (from the position A2 to the position B2) to perform shooting. Here, the robot camera control device 2 receives distance information from the robot camera 1 and transmits trolley control information and robot arm control information to the robot camera 1 so as to move the robot camera 1 to a viewpoint position where people feel that the quality is high.

[0030] <Configuration of Robot Camera Control Device> Here, with reference to FIG. 4 (and FIG. 1 as appropriate), the configuration of the robot camera control device 2 according to the embodiment of the present invention will be described. The robot camera control device 2 includes a three-dimensional scene recognition means 20, a subject designation means 21, a motion prediction means 22, a viewpoint estimation means 23, and a camera position control means 24.

[0031] The three-dimensional scene recognition means 20 acquires point cloud data, which is distance information in the real space, at a predetermined time interval and associates it with a three-dimensional virtual space. Here, the three-dimensional scene recognition means 20 includes a distance information acquisition means 200 and a virtual space arrangement means 201.

[0032] The distance information acquisition means 200 acquires distance information. The distance information is information measured by the distance measurement device 11 mounted on the robot camera 1. The distance information acquisition means 200 receives the distance information transmitted from the robot camera 1 via a communication control unit (not shown).

[0033] As shown in FIG. 5, the distance information is two-dimensional point cloud data G1. The point cloud data G1 corresponds to two-dimensional coordinates and is information associating the two-dimensional coordinates with the distance (depth). Values corresponding to the distances from the distance measurement device 11 to the respective objects SO1, SO2, and MO are set in the point cloud data G1. Note that the point cloud data G1 shown in FIG. 5 shows an example of two-dimensional coordinates with the horizontal axis as the horizontal coordinate and the vertical axis as the vertical coordinate as distance information obtained by the parallax of the stereo camera in the distance measurement device 11. Of course, when LiDAR is used in the distance measurement device 11, it becomes two-dimensional coordinates with the horizontal axis as the azimuth coordinate and the vertical axis as the elevation coordinate. The distance information acquisition means 200 outputs the acquired distance information to the virtual space arrangement means 201.

[0034] The virtual space arrangement means 201 arranges an object in a three-dimensional virtual space from the distance information acquired by the distance information acquisition means 200. The virtual space arrangement means 201 converts and associates distance information (point cloud data) in which a distance (depth) is associated with two-dimensional coordinates into a coordinate space of three-dimensional coordinates (XYZ). At this time, the virtual space arrangement means 201 samples the point cloud data at a predetermined sampling interval and associates it with the coordinate space of the three-dimensional coordinates. Also, the virtual space arrangement means 201 may not arrange points farther than a predetermined depth in the virtual space.

[0035] For example, the virtual space arrangement means 201 samples the point cloud data G1 shown in FIG. 5 and associates it with the three-dimensional coordinates of the XYZ axes as shown in FIGS. 6(a) and 6(b). Note that FIG. 6(a) shows the XY plane and FIG. 6(b) shows the XZ plane. As shown in FIGS. 6(a) and 6(b), the objects SO1, SO2, and MO are arranged on the three-dimensional virtual space at the sampled points. Note that, as shown in FIG. 6(b), in the virtual space, there is no information in the depth direction (Z direction), and there is only information on the front surface of the object.

[0036] The virtual space arrangement means 201 sequentially stores the sampled point cloud data arranged in the three-dimensional virtual space in a storage medium (not shown). The point cloud data stored in this storage medium is referred to by the subject designation means 21, the motion prediction means 22, and the camera position control means 24.

[0037] The subject specifying means 21 displays the point cloud data arranged in the three-dimensional virtual space by the virtual space arranging means 201 on the display device M, and accepts the specification of the subject to be photographed by the operator according to the operator's instruction. Here, the subject specifying means 21 includes an object display means 210 and an object selection means 211.

[0038] The object display means 210 displays the object arranged in the virtual space on the display device M. The object display means 210 schematizes each point of the point cloud data arranged by the virtual space arranging means 201 with a predetermined three-dimensional figure (such as a sphere, a cube, a cylinder, etc.) and displays it on the display device M. For example, the object display means 210 schematizes each point corresponding to the XZ plane shown in Fig. 6(b) with a sphere, and displays the objects SO1, SO2, and MO schematized with the three-dimensional figures shown in Fig. 7 on the display device M. Fig. 7 shows a state in which the object display means 210 also displays the position of the camera C when the point cloud data is acquired on the display device M.

[0039] Note that the display of the schematized object does not have to be a top view arranged in the XZ plane. For example, the object display means 210 may rotate the virtual space in an arbitrary direction, project it onto a two-dimensional plane, and display it. In this way, even for the sampled point cloud data, by schematizing and displaying each point with a three-dimensional figure, the operator can grasp the approximate arrangement of each object.

[0040] The object selection means 211 selects, according to the operator's instruction, an object with movement that becomes the subject to be photographed from the objects displayed on the display device M. The object selection means 211 recognizes the point cloud of the object selected from the objects displayed on the display device M by the operator via an input device such as a mouse or a touch panel (not shown) as the object of the subject. For example, as shown in FIG. 7, assume that a plurality of objects SO1, SO2, MO are displayed on the display device M. At this time, when the operator selects the object MO with an input device such as a mouse, the object selection means 211 specifies the point cloud data constituting the object MO as the subject.

[0041] Note that the object selection method is not particularly limited. For example, the object selection means 211 draws a rectangle, an ellipse, a polygon, etc. according to the mouse operation, and specifies the point cloud data inside thereof as the point cloud data constituting the object MO. Alternatively, when the object selection means 211 selects a figure (a sphere (circle) in the case of FIG. 7) indicating one point of the object by clicking the mouse, it searches for figures that overlap or are in contact with the figure, and specifies the point cloud data corresponding to all the searched figures as the point cloud data constituting the object MO. The object selection means 211 outputs the point cloud data of the selected subject to the motion prediction means 22.

[0042] The motion prediction means 22 predicts the future position of the subject specified by the subject specifying means 21 after a predetermined time has elapsed at a predetermined time interval from the point cloud data arranged in the three-dimensional virtual space by the three-dimensional scene recognition means 20. Here, the motion prediction means 22 includes a subject position estimation means 220 and a virtual space arrangement correction means 221.

[0043] The subject position estimation means 220 tracks the subject specified by the subject specifying means 21 in the virtual space, and estimates the future position of the subject after a predetermined time has elapsed from the past positions of the subject in time series. Here, first, the subject position estimation means 220 detects each object from the point cloud data arranged in the virtual space by the three-dimensional scene recognition means 20. For example, the subject position estimation means 220 detects an object by clustering according to the distance between points. Then, the subject position estimation means 220 tracks the object that is the subject designated by the subject designation means 21, and estimates the future position after a predetermined period of time. Note that since the tracking (tracking algorithm) of the object in the point cloud data is a general method (for example, the ICP [Iterative closest point] algorithm, etc.), the description thereof is omitted here. The subject position estimation means 220 holds the positions (centroid positions) of the subject for a predetermined past period in time series at predetermined time intervals, and estimates the future position after a predetermined time by linear interpolation.

[0044] Specifically, the subject position estimation means 220 estimates the future position of the subject according to the following procedure. Here, let the position of the subject at time t = 0 (here, the centroid position) be the three-dimensional vector p(0). In this case, the positions of the subject sampled at the original time point (t = 0) and the past T samples are p(0), p(-1), p(-2),..., p(-T). Here, assume that the movement of the position (centroid position) of the subject is a linear movement expressed as p(t)=wt + p(0). w is the slope (three-dimensional vector).

[0045] Also, let the position vectors for the past T samples be the matrix P V =[p(-1)-p(0) … p(-2)-p(0) … p(-T)-p(0)]′ (′ means transpose; the same applies hereinafter). Also, let the time for the past T samples be the matrix t V =[-1 -2 … -T]′. In this case, since P V =t V w′, the subject position estimation means 220 calculates the three-dimensional slope w by calculating w = P V ′t V / t V ′t V by the least squares method. The subject position estimation means 220 calculates the position (centroid position) p(τ) of the subject after the sample at the future time t = τ using the calculated slope w by p(τ)=wτ + p(0). The subject position estimation means 220 outputs the estimated position (center of gravity position) p(τ) of the subject at the future time t = τ to the virtual space arrangement correction means 221.

[0046] The virtual space arrangement correction means 221 corrects the positions of the point cloud data constituting the subject so as to arrange the subject at the future position estimated by the subject position estimation means 220 in the virtual space. Specifically, the virtual space arrangement correction means 221 calculates the center of gravity position of the subject at the current time from the positions of the point cloud constituting the subject. Then, the virtual space arrangement correction means 221 translates the point cloud constituting the subject among the point clouds arranged in the virtual space by the three-dimensional scene recognition means 20 so as to move the calculated center of gravity position to the position (center of gravity position) estimated by the subject position estimation means 220. Thereby, the virtual space arrangement correction means 221 can predict the position of the subject at the future time t = τ and correct the arrangement position of the subject. The virtual space arrangement correction means 221 outputs the corrected point cloud data of the virtual space to the viewpoint estimation means 23.

[0047] The viewpoint estimation means 23 estimates the viewpoint position and orientation of a virtual camera that optimizes the viewpoint entropy in the virtual space. That is, the viewpoint estimation means 23 estimates the viewpoint position and orientation of a virtual camera that a person feels has high quality from the virtual space in which the future position of the subject after a predetermined time has elapsed is corrected by the motion prediction means 22. Here, the viewpoint estimation means 23 estimates the viewpoint position and orientation of the virtual camera from a virtual space composed of objects modeled by simple solid figures such as spheres for the point cloud.

[0048] In addition, a method for estimating a viewpoint that a person feels has high quality using viewpoint entropy as an index from a three-dimensional virtual space can use known methods. For example, the viewpoint estimation means 23 estimates the position and orientation of the optimal viewpoint from the arrangement of a plurality of objects including the subject using the method described in Patent Document 1 and Non-Patent Document 4 with viewpoint entropy as an index. The viewpoint estimation means 23 outputs the estimated viewpoint position and orientation of the virtual camera to the camera position control means 24.

[0049] The camera position control means 24 controls the viewpoint position and orientation of the robot camera in the real space corresponding to the viewpoint position and orientation of the virtual camera estimated by the viewpoint estimation means 23. Here, the camera position control means 24 includes a cart position control means 240 and a robot arm control means 241.

[0050] The cart position control means 240 controls the moving direction and distance of the cart 10 so as to move the robot camera 1 to the viewpoint position estimated by the viewpoint estimation means 23. The cart position control means 240 transmits cart control information for moving a predetermined reference position of the cart 10 (for example, the center position of the cart 10) to a position on the plane corresponding to the viewpoint position to the robot camera 1 via a communication control unit (not shown).

[0051] In addition, the cart position control means 240 moves the cart 10 to the viewpoint position (a position where the cart 10 close to the viewpoint position can move) while avoiding the objects arranged in the virtual space by the three-dimensional scene recognition means 20. For example, the cart position control means 240 uses the XZ plane of the virtual space shown in FIG. 7 as an object map and searches for a path to avoid each object by a known method (for example, the * 〔A-star〕 path search algorithm), and moves the cart 10.

[0052] Also, the cart position control means 240 outputs the difference between the position of the cart 10 and the viewpoint position estimated by the viewpoint estimation means 23 to the robot arm control means 241. The robot arm control means 241 controls the robot arm 12 of the robot camera 1 so as to photograph the camera C at the viewpoint position and posture estimated by the viewpoint estimation means 23.

[0053] The robot arm control means 241 moves the camera C to the viewpoint position from the difference between the position of the carriage 10 input from the carriage position control means 240 and the viewpoint position, and transmits robot arm control information for operating the robot arm 12 so as to have a specified posture to the robot camera 1 via a communication control unit (not shown).

[0054] With the configuration described above, even when the shooting situation such as the position of the object changes, the robot camera control device 2 can predict the position of the future subject and control the viewpoint of the robot camera 1 in the real space. Note that the robot camera control device 2 can be operated by a program (robot camera control program) for causing a computer (not shown) to function as each of the above-described means.

[0055] <Operation of the robot camera control device> Next, with reference to FIG. 8 (appropriately refer to FIGS. 1 and 4), the operation of the robot camera control device 2 according to the embodiment of the present invention will be described.

[0056] In step S1, the three-dimensional scene recognition means 20 arranges an object in a three-dimensional virtual space from the distance information transmitted by the robot camera 1. Here, the three-dimensional scene recognition means 20 acquires the distance information measured by the distance measuring device 11 mounted on the robot camera 1 by the distance information acquisition means 20'. Then, the three-dimensional scene recognition means 20 converts the distance information (point cloud data) into a coordinate space of three-dimensional coordinates (XYZ) and associates them by the virtual space arrangement means 201.

[0057] In step S2, the subject designating means 21 displays the object arranged in the virtual space on the display device M by the object display means 210. At this time, the object display means 210 schematizes each point of the point cloud data arranged in the virtual space in step S1 with a predetermined solid figure, and displays it on the display device M. In step S3, the subject designation means 21 selects, by the object selection means 211, an object to be a subject to be photographed from among the objects displayed on the display device M according to the instruction of the operator.

[0058] In step S4, the motion prediction means 22, by the subject position estimation means 220, tracks the current position (center of gravity position) of the subject selected in step S3, and estimates the position after a predetermined time has elapsed (next time point) from the past position. In step S5, the motion prediction means 22, by the virtual space arrangement correction means 221, corrects the arrangement position of the subject at the current time in the three-dimensional virtual space to the future position estimated in step S4. In step S6, the viewpoint estimation means 23 estimates a viewpoint position and posture that a person feels to be of high quality from the virtual space (a space composed of objects schematized by a group of solid figures) in which the future position of the subject after a predetermined time has elapsed in step S5 is corrected.

[0059] In step S7, the camera position control means 24 controls the movement of the robot camera 1 so as to photograph the subject at the viewpoint position and posture estimated in step S6. Here, the camera position control means 24 controls the moving direction and distance of the carriage 10 so that the carriage 10 of the robot camera 1 is moved to the viewpoint position (a position where the carriage 10 close to the viewpoint position can move) estimated in step S6 by the carriage position control means 240. Further, the camera position control means 24 moves the camera C to the viewpoint position from the difference between the position of the carriage 10 and the viewpoint position by the robot arm control means 241, and operates the robot arm 12 so as to achieve the specified posture.

[0060] In step S8, the robot camera control device 2 determines whether an end instruction from the outside is instructed by an operator via a control screen (not shown), a switch, or the like. Here, when there is no end instruction (No in step S8), in step S9, the three-dimensional scene recognition means 20 arranges an object in a three-dimensional virtual space from the distance information transmitted by the robot camera 1. The operation of this step S9 is the same as the operation of step S1. Then, the robot camera control device 2 returns to step S4 and repeats the operation. On the other hand, when there is an end instruction (Yes in step S8), the robot camera control device 2 ends the operation.

[0061] Through the above operations, the robot camera control device 2 can track a subject in a virtual space based on the distance information sequentially transmitted from the robot camera 1, predict the future position, and thus estimate and capture a high-quality viewpoint in the real space even for a moving subject.

[0062] The embodiments of the present invention have been described above, but the present invention is not limited to these embodiments. For example, here, the distance measuring device 11 is configured to be provided in the robot camera 1. However, the distance measuring device 11 may be fixedly placed at a position where it is possible to perform measurement by overlooking the entire space where shooting is performed. In that case, the distance measuring device 11 may be configured to further include the distance information transmitting means 130 of the control device 13.

[0063] Also, here, the imaging system 100 is configured to include one robot camera 1. However, the imaging system 100 may include a plurality of robot cameras 1 and perform shooting at a plurality of viewpoint positions. In that case, the distance measuring device 11 is mounted only on any one of the robot cameras 1. Of course, when the distance measuring device 11 is provided at a fixed position, it is not necessary to mount it on the robot camera 1. Then, the viewpoint estimation means 23 estimates different viewpoints for each individual robot camera 1 based on the degree of importance using the viewpoint entropy as an index by a known method. For example, the viewpoint whose viewpoint entropy is the importance of the overall optimal solution is assigned to the first robot camera 1, and in order from the one with the highest importance of the local optimal solution, that viewpoint is assigned to the second and subsequent robot cameras 1. Also, a plurality of camera position control means 24 may be provided according to the number of robot cameras 1. As a result, the imaging system 100 can perform imaging from a plurality of viewpoints that people feel have high quality, and can assist the editor in performing video editing.

Explanation of Signs

[0064] 100 Imaging system 1 Robot camera 10 Cart 11 Distance measuring device 12 Robot arm 13 Control device 2 Robot camera control device 20 3D scene recognition means 200 Distance information acquisition means 201 Virtual space arrangement means 21 Subject designation means 210 Object display means 211 Object selection means 22 Motion prediction means 23 Viewpoint estimation means 24 Camera position control means 240 Cart position control means 250 Robot arm control means C Camera M Display device

Claims

1. A robot camera control device for controlling a robot camera, comprising: a three-dimensional scene recognition means for acquiring point cloud data, which is distance information in the real space, at a predetermined time interval and associating it with a three-dimensional virtual space; a subject designation means for displaying the point cloud data of the virtual space on a display device and accepting a designation of a subject according to an operator's instruction; a subject position estimation means for tracking the subject designated by the subject designation means in the virtual space and estimating a future position of the subject after a predetermined time has elapsed from the past positions of the subject in time series; a virtual space arrangement correction means for correcting the positions of the point cloud data constituting the subject so that the subject is arranged at the future position in the virtual space; a viewpoint estimation means for estimating the viewpoint position and orientation of a virtual camera that optimizes the viewpoint entropy in the virtual space; a camera position control means for controlling the viewpoint position and orientation of the robot camera in the real space corresponding to the viewpoint position and orientation of the virtual camera estimated by the viewpoint estimation means; A robot camera control device characterized by comprising the above.

2. The robot camera control device according to claim 1, wherein the three-dimensional scene recognition means samples the point cloud data and associates it with the virtual space.

3. The robot camera control device according to claim 2, wherein the subject designation means arranges a predetermined three-dimensional shape at each point of the point cloud data after sampling and displays it on a display device.

4. The robot camera control device according to claim 1, wherein the subject position estimation means holds the centroid positions of the point cloud data constituting the subject in time series, and estimates the future centroid position by linear interpolation from the past centroid positions in time series, and sets the estimated future centroid position as the future position of the subject.

5. A robot camera control device for controlling a plurality of robot cameras, comprising: a three-dimensional scene recognition means for acquiring point cloud data, which is distance information in the real space, at a predetermined time interval and associating it with a three-dimensional virtual space; a subject designation means for displaying the point cloud data of the virtual space on a display device and accepting a designation of a subject; a subject position estimation means for tracking the subject designated by the subject designation means in the virtual space and estimating a future position of the subject after a predetermined time has elapsed from the past positions of the subject in time series; In the virtual space, virtual space arrangement correction means for correcting the positions of the point group data constituting the subject so as to arrange the subject at the future position; Viewpoint estimation means for estimating the viewpoint position and orientation of a virtual camera that is the global optimal solution of the viewpoint entropy as the viewpoint position and orientation of the first robot camera, and sequentially estimating the viewpoint positions and orientations of the second and subsequent robot cameras in order from the ones with higher importance among the plurality of local optimal solutions; Camera position control means for controlling the viewpoint positions and orientations of the plurality of robot cameras in the real space corresponding to the viewpoint positions and orientations of the plurality of virtual cameras estimated by the viewpoint estimation means; A robot camera control device characterized by comprising the above.

6. A program for causing a computer to function as the robot camera control device according to any one of Claims 1 to 5.

7. A photographing system using a robot camera, The robot camera that is driven to be in the viewpoint position and the orientation by being instructed control information for specifying the viewpoint position and the orientation of the camera in the real space, and performs photographing with the camera; The robot camera control device according to Claim 1, which instructs the robot camera to control information for specifying the viewpoint position and the orientation; A photographing system comprising the above.

8. A photographing system using a plurality of robot cameras, The plurality of robot cameras that are driven to be in the viewpoint position and the orientation by being instructed control information for specifying the viewpoint position and the orientation of the camera in the real space, and perform photographing with the camera; The robot camera control device according to Claim 5, which instructs each of the plurality of robot cameras to control information for specifying the viewpoint position and the orientation for each of the plurality of robot cameras; A photographing system comprising the above.

Citation Information

Patent Citations

  • Optimum viewpoint detecting device and program thereof

    JP2022173917A

  • Automatic camerawork generation device, automatic camerawork generation program, camerawork learning device, and camerawork learning program

    JP2023023510A