Information processing apparatus, information processing method, and program

The information processing device enhances the control of virtual cameras by integrating automatic and manual operations, addressing the operational challenges in generating virtual viewpoint images and improving user convenience.

JP2026010132APending Publication Date: 2026-01-21CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025175750
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-01-21

AI Technical Summary

Technical Problem

Existing methods for controlling virtual cameras in generating virtual viewpoint images can be difficult to operate, especially when significant or complex movements are required, leading to a lack of convenience in viewing these images.

Method used

An information processing device that combines automatic and manual operations for virtual camera control, using point of interest and attitude information to generate virtual viewpoint images, allowing for seamless adjustments without changing the camera's position.

Benefits of technology

Improves the convenience of viewing virtual viewpoint images by enabling smooth and intuitive control of virtual camera movements through a combination of automatic and manual operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026010132000001_ABST
    Figure 2026010132000001_ABST
Patent Text Reader

Abstract

To improve convenience when browsing a virtual viewpoint image.SOLUTION: The information processing apparatus acquires target point information indicating a position of a target point of the virtual camera corresponding to each time in a predetermined time and orientation information indicating an orientation of the virtual camera, acquires an input corresponding to a user operation for changing the orientation of the virtual camera in the predetermined time, and changes the orientation of the virtual camera indicated by the orientation information by the input without changing the position of the target point of the virtual camera indicated by the target point information based on the plurality of captured images, the target point information, the orientation information, and the input, thereby generating a first virtual viewpoint image.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and a program related to a virtual viewpoint image. [Background technology]

[0002] In recent years, attention has been focused on technology for generating virtual viewpoint images, which generate images from any viewpoint from multiple images taken with multiple cameras. The concept of a virtual camera is used to conveniently explain the virtual viewpoint specified to generate virtual viewpoint images. Unlike physical cameras, virtual cameras are capable of various movements in three-dimensional space, such as translation, rotation, and scaling, without being subject to physical constraints. In order to properly control a virtual camera, multiple operation methods corresponding to each movement have been devised.

[0003] Patent Document 1 discloses a method for controlling the operation of a virtual camera in a configuration in which a virtual viewpoint image is displayed on an image display device equipped with a touch panel, depending on the number of fingers used to perform a touch operation on the touch panel. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent No. 6419278 Summary of the Invention [Problem to be solved by the invention]

[0005] The method described in Patent Document 1 can be difficult to operate depending on the intended movement of the virtual viewpoint, such as when the user wants to move the virtual viewpoint significantly or when the user wants to make complex changes. In view of the above-mentioned problems, an object of the present invention is to improve the convenience of viewing virtual viewpoint images. [Means for solving the problem]

[0006] In order to solve the above-mentioned problems, one aspect of the information processing device according to the present invention is a first acquisition means for acquiring point of interest information indicating the position of a point of interest of a virtual camera corresponding to each time instant within a predetermined period of time and attitude information indicating the attitude of the virtual camera; a second acquisition means for acquiring an input corresponding to a user operation for changing the attitude of the virtual camera at the predetermined time; a generating means for generating a first virtual viewpoint image based on a plurality of captured images, the attention point information, the attitude information, and the input, by changing the attitude of the virtual camera indicated by the attitude information in accordance with the input, without changing the position of the attention point of the virtual camera indicated by the attention point information; The present invention is characterized by having the following. [Effects of the Invention]

[0007] According to the present invention, it is possible to improve convenience when viewing a virtual viewpoint image. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating the configuration of a virtual viewpoint image generation system. [Figure 2] FIG. 1 is a diagram illustrating the configuration of an image display device. [Figure 3] 1 is a diagram showing the position, orientation, and focus point of a virtual camera. [Figure 4] 10A and 10B are diagrams showing an example of automatic operation of a virtual camera. [Figure 5] 4A and 4B are diagrams showing a manual operation screen and manual operation information; [Figure 6] 10 is a flowchart showing an example of a process for operating a virtual camera. [Figure 7] 10A to 10C are diagrams showing operation examples and display examples. [Figure 8] 10A and 10B are diagrams showing a tag generation and tag reproduction method according to a second embodiment. [Figure 9] 10 is a flowchart showing an example of a process for operating a virtual camera according to the second embodiment. [Figure 10] FIG. 11 is a diagram showing an example of manual operation according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0010] First Embodiment In this embodiment, a system is described that generates a virtual viewpoint image representing a view from a specified virtual viewpoint based on a plurality of images captured by a plurality of imaging devices and a specified virtual viewpoint. The virtual viewpoint image in this embodiment is not limited to an image corresponding to a viewpoint freely (arbitrarily) specified by a user, and also includes, for example, an image corresponding to a viewpoint selected by a user from a plurality of candidates. Furthermore, in this embodiment, the case where the virtual viewpoint is specified by a user operation is mainly described, but the virtual viewpoint may also be specified automatically based on the results of image analysis, etc.

[0011] In this embodiment, the term "virtual camera" will be used for explanation. The virtual camera is a virtual camera that is different from the multiple imaging devices actually installed around the imaging area, and is a concept for conveniently explaining the virtual viewpoint related to the generation of a virtual viewpoint image. In other words, the virtual viewpoint image can be considered to be an image captured from a virtual viewpoint set in a virtual space associated with the imaging area. The position and orientation of the viewpoint in this virtual imaging can be expressed as the position and orientation of the virtual camera. In other words, the virtual viewpoint image can be said to be an image that simulates an image captured by a camera when it is assumed that the camera exists at the position of the virtual viewpoint set in space.

[0012] The explanation will be given in the following order: Fig. 1 explains the entire virtual viewpoint image generation system, and Fig. 2 explains the image display device 104 therein. This image display device 104 performs virtual camera control processing that effectively combines automatic operation and manual operation. Fig. 6 explains the virtual camera control processing, and Fig. 7 explains an example of the control and an example of displaying a virtual viewpoint image using the virtual camera.

[0013] 3 to 5, the configuration of the virtual camera required for the explanation, as well as automatic processing and manual processing, will be explained.

[0014] (Configuration of virtual viewpoint image generation system) First, the configuration of a virtual viewpoint image generation system 100 according to this embodiment will be described with reference to FIG.

[0015] The virtual viewpoint image generation system 100 has n sensor systems, from sensor system 101a to sensor system 101n, and each sensor system has at least one camera, which is an imaging device. Hereinafter, unless otherwise specified, the n sensor systems will not be distinguished from one another and will be referred to as multiple sensor systems 101.

[0016] FIG. 1(B) is a diagram showing an example of installation of multiple sensor systems 101. The multiple sensor systems 101 are installed to surround an area 120, which is the area to be photographed, and each sensor system photographs the area 120 from a different direction. In the example of this embodiment, the area to be photographed 120 is assumed to be the field of a stadium where a soccer match is played, and n (e.g., 100) sensor systems 101 are installed to surround the field. However, the number of sensor systems 101 to be installed is not limited, and the area to be photographed 120 is not limited to the stadium field. For example, the area 120 may include stadium seating, or the area 120 may be an indoor studio, stage, etc.

[0017] Furthermore, the multiple sensor systems 101 do not have to be installed around the entire periphery of the area 120, and may be installed only in a part of the periphery of the area 120 depending on restrictions on installation location, etc. Furthermore, the multiple cameras of the multiple sensor system 101 may include imaging devices with different functions, such as a telephoto camera and a wide-angle camera.

[0018] The multiple cameras included in the multiple sensor systems 101 capture images synchronously. The multiple images obtained by capturing images from these cameras are called multi-view images. Note that each of the multi-view images in this embodiment may be a captured image, or may be an image obtained by performing image processing on the captured image, such as a process of extracting a predetermined area.

[0019] The multiple sensor systems 101 may have a microphone (not shown) in addition to a camera. The microphones of the multiple sensor systems 101 synchronously collect audio. Based on this collected audio, an acoustic signal can be generated that is played back together with the display of an image on the image display device 104. For the sake of simplicity, a description of audio will be omitted below, but it is assumed that images and audio are basically processed together.

[0020] The image recording device 102 acquires multi-viewpoint images from multiple sensor systems 101 and stores them together with the time code used for shooting in a database 103. The time code is information that uniquely identifies the time when the image was captured by the imaging device, and can be specified in a format such as "day:hour:minute:second.frame number."

[0021] The image display device 104 provides images based on the multi-viewpoint images corresponding to the time code from the database 103 and the user's manual operation of the virtual camera.

[0022] The virtual camera 110 is set in a virtual space associated with the area 120, and can view the area 120 from a viewpoint different from that of any of the cameras of the multiple sensor systems 101. Details of the virtual camera 110 and its operation will be described later with reference to FIG.

[0023] In this embodiment, both automatic and manual operation of virtual camera 110 is used, and automatic operation information for automatically operating the virtual camera will be described later with reference to Fig. 4. Also, a manual operation method for manually operating virtual camera 110 will be described later with reference to Fig. 5.

[0024] The image display device 104 executes a virtual camera control process that effectively combines automatic operation and manual operation, and generates a virtual viewpoint image from a multi-viewpoint image based on the controlled virtual camera and time code. This virtual camera control process is described in detail in Fig. 6, and an example of this control and an example of a virtual viewpoint image generated using the virtual camera are described in detail in Fig. 7.

[0025] The virtual viewpoint image generated by the image display device 104 is an image that represents the view from the virtual camera 110. The virtual viewpoint image in this embodiment is also called a free viewpoint image, and is displayed on the touch panel, liquid crystal display, or the like of the image display device 104.

[0026] The configuration of the virtual viewpoint image generation system 100 is not limited to that shown in Fig. 1(A). The image display device 104 may be configured to be separate from the operation device and display device. Alternatively, a plurality of display devices may be connected to the image display device 104, and a virtual viewpoint image may be output to each of them.

[0027] In the example of FIG. 1(A), the database 103 and the image display device 104 are described as separate devices, but the database 103 and the image display device 104 may be configured as an integrated unit. Also, a configuration may be used in which only important scenes are copied from the database 103 to the image display device 104 in advance. Furthermore, depending on the settings of the time codes accessible by the database 103 and the image display device 104, it may be possible to switch between allowing access to all time codes of the game or only to some of them. In a configuration in which some data is copied, it may be possible to allow access to only the time codes of the copied section.

[0028] In this embodiment, the virtual viewpoint image is mainly a moving image, but the virtual viewpoint image may be a still image.

[0029] (Functional configuration of image display device) Next, the configuration of the image display device 104 will be described with reference to FIG.

[0030] 2(A) is a diagram showing an example of the functional configuration of the image display device 104. The image display device 104 uses the functions shown in the figure to execute virtual camera control processing that effectively combines automatic operation and manual operation. The image display device 104 includes a manual operation unit 201, an automatic operation unit 202, an operation control unit 203, a tag management unit 204, a model generation unit 205, and an image generation unit 206.

[0031] The manual operation unit 201 is a receiving unit for receiving input information manually operated by the user with respect to the virtual camera 110 or the time code. Manual operation of the virtual camera includes operation of at least one of a touch panel, a joystick, and a keyboard, but the manual operation unit 201 may also obtain input information via operation of another input device. Details of the manual operation screen will be described later with reference to FIG. 5.

[0032] The automatic operation unit 202 automatically operates the virtual camera 110 and the time code. An example of automatic operation in this embodiment will be described in detail with reference to FIG. 4, in which the image display device 104 operates or sets the virtual camera or time code separately from manual operation by the user. The user experiences the virtual viewpoint image as if it were being automatically operated, separate from operations performed by the user. Information used by the automatic operation unit 202 for automatic operation is called automatic operation information, and uses, for example, coordinates of three-dimensional models of players and the ball, which will be described later, and position information measured by GPS or the like. Note that the automatic operation information is not limited to these, and may be any information that can specify information related to the virtual camera and time code without relying on user operation.

[0033] The operation control unit 203 controls the virtual camera or time code by meaningfully combining the user input information acquired from the manual operation screen by the manual operation unit 201 and the automatic operation information used by the automatic operation unit 202.

[0034] Details of the processing by the operation control unit 203 will be described later with reference to FIG. 6, and an example of a virtual viewpoint image generated by the controlled virtual camera will be described later with reference to FIG.

[0035] Furthermore, the operation control unit 203 sets the virtual camera operation to be operated for each of the automatic operation and manual operation, which is called operation setting processing, and will be described in detail later with reference to FIGS. 4(A) and 4(B).

[0036] Model generation unit 205 generates a three-dimensional model representing the three-dimensional shape of an object in area 120 based on multi-viewpoint images acquired by specifying a time code from database 103. Specifically, a foreground image in which a foreground area corresponding to an object such as a ball or a person is extracted, and a background image in which a background area other than the foreground area is extracted, are acquired from the multi-viewpoint images. Then, model generation unit 205 generates a three-dimensional model of the foreground based on the multiple foreground images.

[0037] The three-dimensional model is generated by a shape estimation method such as visual hull analysis and is composed of a point cloud. However, the format of the three-dimensional shape data representing the shape of the object is not limited to this. The three-dimensional model of the background may be acquired in advance by an external device. The model generation unit 205 may be configured to be included in the image recording device 102 rather than the image display device 104. In this case, the three-dimensional model is recorded in the database 103, and the image display device 104 reads the three-dimensional model from the database 103. For the three-dimensional model of the foreground, the coordinates of each ball, person, and object may be calculated and stored in the database 103. The coordinates of each object may be specified as automatic operation information, which will be described later with reference to FIG. 4(C), and used for automatic operation.

[0038] The image generation unit 206 generates a virtual viewpoint image from the three-dimensional model based on the controlled virtual camera. Specifically, appropriate pixel values ​​are acquired from the multi-viewpoint image for each point constituting the three-dimensional model, and coloring processing is performed. The colored three-dimensional model is then placed in a three-dimensional virtual space, projected onto the virtual viewpoint, and rendered, thereby generating a virtual viewpoint image. However, the method for generating the virtual viewpoint image is not limited to this, and various methods can be used, such as a method for generating a virtual viewpoint image by projective transformation of a captured image without using a three-dimensional model.

[0039] The tag management unit 204 manages, as tags, the automatic operation information used by the automatic operation unit 202. Tags will be described later in a second embodiment, and in this embodiment, a configuration that does not use tags will be described.

[0040] (Hardware configuration of image display device) 2(B), the hardware configuration of the image display device 104 will be described. The image display device 104 includes a CPU (Central Processing Unit) 211, a RAM (Random Access Memory) 212, and a ROM (Read Only Memory) 213. The image display device 104 also includes an operation input unit 214, a display unit 215, and an external interface 216.

[0041] The CPU 211 performs processing using programs and data stored in the RAM 212 and the ROM 213. The CPU 211 controls the overall operation of the image display device 104 and executes processing for realizing each function shown in Fig. 2(A). Note that the image display device 104 may have one or more pieces of dedicated hardware different from the CPU 211, and at least a part of the processing by the CPU 211 may be executed by the dedicated hardware. Examples of the dedicated hardware include an ASIC (application specific integrated circuit), an FPGA (field programmable gate array), and a DSP (digital signal processor).

[0042] The ROM 213 holds programs and data. The RAM 212 has a work area for temporarily storing programs and data read from the ROM 213. The RAM 212 also provides a work area used by the CPU 211 when it executes various processes.

[0043] The operation input unit 214 is, for example, a touch panel, and acquires information operated by the user. For example, it accepts operations on a virtual camera or a time code. The operation input unit 214 may be connected to an external controller and accept input information from the user regarding operations. The external controller may be, for example, a three-axis controller such as a joystick, a mouse, or the like. The external controller is not limited to these.

[0044] The display unit 215 is a touch panel or a screen that displays the generated virtual viewpoint image. In the case of a touch panel, the operation input unit 214 and the display unit 215 are integrated into one unit.

[0045] The external interface 216 transmits and receives information to and from the database 103 via, for example, a LAN or the like. Information may also be transmitted to an external screen via an image output port such as HDMI (registered trademark) or SDI. Image data may also be transmitted via Ethernet or the like.

[0046] (Virtual camera movement) Next, the operation of virtual camera 110 (or virtual viewpoint) will be described with reference to Fig. 3. To explain this operation, the position, orientation, view frustum, and point of interest of the virtual camera will first be described.

[0047] The virtual camera 110 and its movement are specified using a coordinate system, which is a typical Cartesian coordinate system in three-dimensional space consisting of X, Y, and Z axes, as shown in Fig. 3(A).

[0048] This coordinate system is set for the subject and used. The subject is a stadium field, a studio, etc. As shown in FIG. 3(B), the subject includes the entire stadium field 391, as well as a ball 392 and players 393 on the field. The subject may also include spectator seats around the field.

[0049] The coordinate system for the subject is set with the center of field 391 as the origin (0,0,0). The X axis is the long side direction of field 391, the Y axis is the short side direction of field 391, and the Z axis is the vertical direction to the field. Note that the method for setting the coordinate system is not limited to these.

[0050] Next, the virtual camera will be explained using Figures 3(C) and 3(D). The virtual camera serves as a viewpoint for drawing a virtual viewpoint image. In the quadrangular pyramid shown in Figure 3(C), the vertices represent the position 301 of the virtual camera, and the vectors extending from the vertices represent the orientation 302 of the virtual camera. The position of the virtual camera is expressed by coordinates (x, y, z) in three-dimensional space, and the orientation is expressed by a unit vector with the components of each axis as scalars.

[0051] The virtual camera pose 302 passes through the center points of the front clip plane 303 and the rear clip plane 304. The space 305 sandwiched between the front clip plane 303 and the rear clip plane is called the virtual camera's view frustum, and is the range in which the image generator 203 generates a virtual viewpoint image (or the range in which the virtual viewpoint image is projected and displayed; hereinafter, this range is referred to as the display area of ​​the virtual viewpoint image). The virtual camera pose 302 is expressed by a vector, and is also called the optical axis vector of the virtual camera.

[0052] The movement and rotation of the virtual camera will be explained using Figure 3(D). The virtual camera moves and rotates within a space expressed in three-dimensional coordinates. The movement 306 of the virtual camera is the movement of the virtual camera position 301, and is expressed by the components of each axis (x, y, z). The rotation 307 of the virtual camera is expressed by Yaw, which is a rotation around the Z axis, Pitch, which is a rotation around the X axis, and Roll, which is a rotation around the Y axis, as shown in Figure 3(A).

[0053] These allow the virtual camera to move and rotate freely in the three-dimensional space of the subject (field), and any area of ​​the subject can be generated as a virtual viewpoint image.In other words, by specifying the virtual camera's X, Y, and Z coordinates (x, y, z) and the rotation angles of the X, Y, and Z axes (pitch, roll, yaw), the shooting position and shooting direction of the virtual camera can be controlled.

[0054] Next, the position of the target point of the virtual camera will be described with reference to Fig. 3(E), which shows a state in which the orientation 302 of the virtual camera is directed toward the target plane 308.

[0055] The plane of interest 308 is, for example, the XY plane (field plane, Z=0) described in Figure 3(B). The plane of interest 308 is not limited to this, and may be a plane parallel to the XY plane and at the player's height (for example, Z=1.7m) where people naturally turn their gaze. Note that the plane of interest 308 does not have to be parallel to the XY plane.

[0056] The position (coordinates) of the virtual camera's point of interest indicates the intersection 309 of the virtual camera's orientation 302 (vector) and the plane of interest 308. The coordinates 309 of the virtual camera's point of interest can be calculated as unique coordinates once the virtual camera's position 301, orientation 302, and plane of interest 308 are determined. Note that calculation of the intersection of the vector and plane used in this calculation (virtual camera's point of interest 309) can be achieved by a known method, so this is omitted here. Note that a virtual viewpoint image may be generated so that the coordinates 309 of the virtual camera's point of interest are at the center of the drawn virtual viewpoint image.

[0057] In this embodiment, the virtual camera movements that can be realized by combining the movement and rotation of the virtual camera are called virtual camera movements. There are an infinite number of virtual camera movements that can be realized by combining the above-mentioned movements and rotations, and among them, combinations that are easy for the user to see or operate may be defined in advance as virtual camera movements. There are several representative virtual camera movements according to this embodiment, such as enlargement (zoom in), reduction (zoom out), translation of the point of interest, and horizontal or vertical rotation around the point of interest, which are different from the above rotations. These movements will be described next.

[0058] Scaling up and down is the operation of enlarging or reducing the display of an object in the display area of ​​the virtual camera. Scaling up and down of an object is achieved by moving the virtual camera position 301 back and forth along the virtual camera's attitude direction 302 (optical axis direction). When the virtual camera moves forward along the attitude direction (optical axis direction) 302, it approaches the object in the display area without changing its direction, resulting in the object being displayed as enlarged. Conversely, when the virtual camera moves backward along the attitude direction (optical axis direction), the object is displayed as reduced. Note that scaling up and down is not limited to these methods, and it is also possible to use changes in the focal length of the virtual camera, etc.

[0059] Translation of the point of interest is an operation of moving the point of interest 309 of the virtual camera without changing the attitude 302 of the virtual camera. This operation will be explained with reference to FIG. 3(F). The translation of the point of interest 309 of the virtual camera is a trajectory 320, and refers to the movement of the virtual camera from point of interest 309 to point of interest 319. With translation of the point of interest 320, the position of the virtual camera changes but the attitude vector 302 does not change, so this is a virtual camera operation in which only the point of interest moves and the subject continues to be viewed from the same direction. Note that since the attitude vector 302 does not change, translation 320 of the point of interest 309 of the virtual camera and translation 321 of the position 301 of the virtual camera follow the same movement trajectory (movement distance).

[0060] Next, horizontal rotation of the virtual camera around the point of interest will be described with reference to FIG. 3(G). Horizontal rotation around the point of interest is an operation in which the virtual camera rotates 322 along a plane parallel to the plane of interest 308, with point of interest 309 as the center, as shown in FIG. 3(G). During horizontal rotation around point of interest 309, the orientation vector 302 changes while continuing to face point of interest 309, and the coordinates of point of interest 309 do not change. Because point of interest 309 does not move when the user rotates around the point of interest, the user can view surrounding objects from various angles without changing the height from the plane of interest. This operation does not shift the point of interest and does not change the height distance from the object (plane of interest), allowing the user to easily view the object of interest from various angles.

[0061] Note that the trajectory 322 of the virtual camera position 301 when it rotates is on a plane parallel to the plane of interest 308, and the center of rotation is 323. That is, assuming that the X and Y coordinates of the point of interest are (X1, Y1) and the X and Y coordinates of the virtual camera are (X2, Y2), the position of the virtual camera is determined so that it moves on a circle of radius R, where R^2 = (X1 - X2)^2 + (Y1 - Y2)^2. Also, by determining the attitude so that the attitude vector is in the X and Y coordinate directions of the point of interest from the virtual camera position, it is possible to view the subject from various angles without shifting the point of interest.

[0062] Next, vertical rotation of the virtual camera around the point of interest will be described with reference to FIG. 3(H). Vertical rotation around the point of interest is an operation in which the virtual camera rotates 324 along a plane perpendicular to the plane of interest 308, with point of interest 309 as the center, as shown in FIG. 3(H). In vertical rotation around the point of interest, the orientation vector 302 changes while continuing to point toward point of interest 309, and the coordinates of point of interest 309 do not change. Unlike the horizontal rotation 322 described above, this operation changes the height of the virtual camera from the plane of interest. In this case, the position of the virtual camera can be determined so that the distance between the virtual camera position and the point of interest does not change on a plane that includes the virtual camera position and the point of interest and is perpendicular to the XY plane. Furthermore, by determining the orientation so that the orientation vector points from the virtual camera position to the point of interest, the subject can be viewed from various angles without shifting the point of interest.

[0063] By combining the horizontal rotation 322 and vertical rotation 324 about the point of interest described above, the virtual camera can realize an operation that allows the subject to be viewed from various angles of 360 degrees without changing the point of interest 309. An example of virtual camera operation control using this operation will be described later with reference to FIG.

[0064] The virtual camera operations are not limited to these, and any operations can be realized by combining the movement and rotation of the virtual camera. The virtual camera operations described above may be performed automatically or manually.

[0065] (Operation setting process) In this embodiment, the virtual camera operations described in Fig. 3 can be freely assigned to automatic operations and manual operations. This assignment is called an operation setting process.

[0066] Next, the operation setting process will be described with reference to Figures 4(A) and 4(B). The operation setting process is executed by the operation control unit 203 of the image display device 104.

[0067] FIG. 4(A) is a list in which identifiers are assigned to virtual camera operations, and is used in the operation setting process. In FIG. 4(A), the virtual camera operations described in FIG. 3 are set for each operation identifier (ID). For example, identifier=2 is set to the translation of the point of interest described with reference to FIG. 3(F), identifier=3 is set to the horizontal rotation about the point of interest described with reference to FIG. 3(G), and identifier=4 is set to the vertical rotation about the point of interest described with reference to FIG. 3(H). Similar settings can be made for other identifiers. The operation control unit 203 uses this identifier to specify the virtual camera operations to be assigned to automatic operation or manual operation. Note that, in one example, multiple operation identifiers may be assigned to the same operation. For example, identifier=5 may be set to the translation of the point of interest described with reference to FIG. 3(F), with a movement distance of 10 m, and identifier=6 may be set to the translation of the point of interest described with reference to FIG. 3(F), with a movement distance of 20 m.

[0068] Figure 4(B) is an example of operation settings. In Figure 4(B), the translation of the virtual camera's focus point (identifier = 2) is set as the target of automatic operation. In addition, the following are set as targets of manual operation: zoom in / out (identifier = 1), translation of the focus point (identifier = 2), and horizontal rotation (identifier = 3) and vertical rotation (identifier = 4) around the focus point.

[0069] The operation control unit 203 of the image display device 104 controls the virtual camera operation by automatic operation and manual operation according to this operation setting. In the example of FIG. 4(B), the focus point of the virtual camera is translated according to the time code by automatic operation, and the subject can be viewed at various angles and magnifications by manual operation by the user without changing the position of the focus point specified by automatic operation. Detailed explanations of the virtual camera operation according to this operation setting example and display examples of virtual viewpoint images will be given later with reference to FIGS. 4(C), 4(D), and 7. Note that the combinations that can be set for operation are not limited to the example of FIG. 4(B), and the contents of the operation setting can also be changed midway. Furthermore, more than one operation can be set for automatic operation or manual operation, and multiple virtual camera operations can be set for each.

[0070] Furthermore, the combination of set operations may basically be exclusive control between automatic operation and manual operation. For example, in the example of Fig. 4(B), the target of automatic operation is the translation of the focus point of the virtual camera, and during automatic operation, even if a translation of the focus point is operated by manual operation, processing may be performed to cancel (ignore) it. Furthermore, processing may be performed to temporarily release the exclusive control.

[0071] The combination of operations to be set may be preset by the image display device 104 as setting candidates, or may be set by accepting a selection operation from the user to assign virtual camera operations to automatic operation and manual operation.

[0072] (Automated Operation Information) Regarding the virtual camera operation specified for automatic operation in the operation setting process described with reference to Figures 4(A) and 4(B), information indicating specific details for each time code is called automatic operation information. Here, automatic operation information will be described with reference to Figures 4(C) and 4(D). The identifier (type) of the virtual camera operation specified in the automatic operation information is determined by the operation setting process as described with reference to Figures 4(A) and 4(B). Figure 4(C) is an example of automatic operation information for the operation setting example of Figure 4(B).

[0073] The automatic operation information 401 in Fig. 4(C) specifies the position and orientation of the virtual camera's focus point for each time code as virtual camera operation. A series of time codes (2020-02-02 13:51:11.020 to 2020-02-02 13:51:51.059), such as the time codes from the first line to the last line in Fig. 4(C), is called a time code section. The start time (13.51.11.020) is called the start time code, and the end time (13:51:51.059) is called the end time code.

[0074] The coordinates of the point of interest are specified as (x, y, z), with z being a constant value (z=z01) and only (x, y) changing. When these continuous coordinate values ​​are plotted in the three-dimensional space 411 shown in FIG. 4(D), they become a trajectory 412 that moves parallel to the field surface at a constant height (identifier = 2 in FIG. 4(B)). Meanwhile, only the initial value is specified for the posture of the virtual camera in the corresponding time code section. Therefore, automatic operation using the automatic operation information 401 operates the virtual camera so that the posture does not change from the initial value (-Y direction) in the specified time code section, and the point of interest follows the trajectory 412 (421-423).

[0075] For example, if the trajectory 412 is set to follow the position of the ball in a field sport such as rugby (position of the point of interest = position of the ball), this automatic operation will cause the virtual camera to always follow the ball and surrounding subjects without changing its posture.

[0076] 4(E) to 4(G) show examples of virtual viewpoint images when the virtual camera is positioned at 421 to 423 when the automatic operation information shown in FIG. 4(D) is used. As shown in these figures, although the subject that appears in the angle of view of the virtual camera changes as the time code and position change, all virtual viewpoint images have the same orientation vector, i.e., the -Y direction. In this embodiment, automatic operation using such automatic operation information 401 is also called automatic playback.

[0077] The coordinates of the ball and players, which are foreground objects, calculated when the three-dimensional model is generated as explained in FIG. 2 can be used as the coordinates of continuous values ​​such as the trajectory 412.

[0078] Data acquired from a device other than the image display device 104 may also be used for the automatic operation information. For example, a positioning tag may be attached to a ball or a player's clothing included in the subject, and the position information may be used as a point of interest. Furthermore, the coordinates of a foreground object calculated when generating a three-dimensional model or the position information acquired by the positioning tag may be used directly for the automatic operation information, or information calculated separately from such information may also be used. For example, with regard to the coordinate values ​​and positioning values ​​of a foreground object in a three-dimensional model, the z value varies depending on the player's posture and the position of the ball. However, z may be a fixed value, and only the x and y values ​​may be used. The z value may be set to the height of the field (z = 0) or a height to which viewers naturally turn (e.g., z = 1.7 m). The device that acquires the automatic operation information is not limited to the positioning tag. For example, information received from an image display device other than the image display device 104 currently displaying the information may also be used. For example, an operator operating the automatic operation information may input the automatic operation information by operating an image display device other than the image display device 104.

[0079] The time code section that can be specified in the automatic operation information may be the entire game of a field sport, or a part of it.

[0080] Note that the examples of the automatic operation information are not limited to the examples described above. The items of the automatic operation information are not limited as long as they are information related to the operation of the virtual camera. Other examples of the automatic operation information will be described later.

[0081] (Manual operation screen) Manual operation of the virtual camera according to this embodiment will be described with reference to Fig. 5. The virtual camera operation specified by manual operation is determined by the operation setting process shown in Fig. 4(A) and Fig. 4(B). Here, a manual operation method for the virtual camera operation assigned by the operation setting process and a manual operation method for the time code will be described.

[0082] FIG. 5A is a diagram for explaining the configuration of a manual operation screen 501 of the operation input unit 214 of the image display device 104. As shown in FIG.

[0083] 5(A), the manual operation screen 501 is broadly composed of two areas: a virtual camera operation area 502 and a time code operation area 503. The virtual camera operation area 502 accepts the user's manual operation of the virtual camera, and the time code operation area 503 is a graphical user interface (GUI) for accepting the user's manual operation of the time code.

[0084] First, a description will be given of the virtual camera operation area 502. The virtual camera operation area 502 executes the virtual camera operation set in the operation setting process of Fig. 4(A) and Fig. 4(B) according to the type of operation received.

[0085] In the example of Figure 5, a touch panel is used, so the types of operations are touch operations such as tapping and swiping. On the other hand, if a mouse is used to input user operations, a click operation may be performed instead of a tap operation, and a drag operation may be performed instead of a swipe operation. To each of these touch operations, operations such as translation of the point of interest and horizontal / vertical rotation around the point of interest, which are the virtual camera operations shown in Figure 4, can be assigned. The table that determines such assignments is called the manual operation setting and is shown in Figure 5(B).

[0086] The manual operation settings will be explained using Fig. 5(B). In Fig. 5(B), the items include touch operation, number of touches, touch area, and virtual camera operation identifier (ID).

[0087] The touch operation item lists touch operations possible in the virtual camera operation area 502. Examples include tapping and swiping. The touch count item defines the number of fingers required for a touch operation. The touch area item specifies the area to be processed by the touch operation. For example, the virtual camera operation area (entire) or the virtual camera operation area (right edge) may be specified. The content of the manual operation accepted from the user is distinguished according to the contents of the three items, the touch operation item, the touch count item, and the touch area item. Then, by assigning a virtual camera operation identifier (ID) to each manual operation, one of the virtual camera operations is executed when each manual operation is accepted. The virtual camera operation identifier (ID), as described with reference to FIG. 4(A), identifies the type of virtual camera operation described in FIG. 3.

[0088] In the allocation example of Fig. 5(B), for example, by specifying the second line (No. 2), horizontal rotation (action identifier = 3) is executed around the point of interest when a horizontal swipe operation with one finger is accepted on the virtual camera operation area 502. Also, by specifying the third line (No. 3), vertical rotation (action identifier = 4) is executed around the point of interest when a vertical swipe operation with one finger is accepted on the virtual camera operation area 502.

[0089] As for the type of touch operation used for manual operation of the virtual camera, gesture operations according to the number of fingers may be used, or relatively simple gesture operations shown in a third embodiment described later may be used. Either of these can be specified in the manual operation settings of Fig. 5(B). Note that the operations that can be specified in the manual operation settings are not limited to touch operations, and operations using devices other than a touch panel can also be specified.

[0090] Next, the time code operation area 503 will be described. The time code operation area 503 is made up of elements 512 to 515 for operating the time code. The main slider 512 can be operated for all time codes of the shooting data. By selecting the position of the knob 522 by dragging or the like, it is possible to specify any time code within the main slider 512.

[0091] The sub-slider 513 displays an enlarged portion of the total time code, allowing for finer time code manipulation than the main slider 512. The position of the knob 523 can be selected by dragging or other operations to specify the time code. The main slider 512 and the sub-slider 513 appear to be the same length on the screen, but the selectable time code width differs. For example, the main slider 512 allows for selection from three hours, which is the length of a single game, while the sub-slider 513 allows for selection from 30 seconds, which is a portion of that time. In other words, each slider has a different scale, and the sub-slider 513 allows for finer time code specification, such as in frame units. In one example, the sub-slider 513 ranges from 15 seconds before to 15 seconds after the time indicated by the knob 522.

[0092] The time code specified using the knob 522 of the main slider 521 or the time code specified using the knob 523 of the sub-slider 513 may be displayed as a number in the format of "day:hour:minute:second.frame number." The sub-slider 513 does not have to be displayed all the time. For example, it may be displayed after a display instruction is received, or when a specific operation such as pausing is instructed. The time code section that can be selected with the sub-slider 513 may be variable. When a specific operation such as pausing is received, a section of about 15 seconds before and after the time of the knob 523 at the time the pause instruction was received may be displayed.

[0093] The slider 514 for specifying the playback speed can specify playback speeds such as normal playback or slow playback. The count-up interval of the time code is controlled according to the playback speed selected with the knob 524. An example of controlling the count-up interval of the time code will be described with reference to the flowchart in FIG.

[0094] The cancel button 515 may be used to cancel each operation related to the time code, or to clear the pause and return to normal playback. Note that the button is not limited to the cancel button as long as it is a button for performing manual operations related to the time code.

[0095] The manual operation screen may have areas other than the virtual camera operation area 502 and the time code operation area 503. For example, match information may be displayed in area 511 at the top of the screen. Match information may include the venue, date and time, match cards, and score status. However, the match information is not limited to these.

[0096] Furthermore, an exceptional operation may be assigned to the area 511 at the top of the screen. For example, when the area 511 at the top of the screen receives a double tap, the operation may be to move the position and orientation of the virtual camera to a position where the entire subject can be viewed from above. Manual operation of the virtual camera can be difficult for an inexperienced user, and the user may lose track of where they are. In such a case, an operation may be assigned to return the camera to a bird's-eye view point where the position and orientation are easy for the user to understand. A bird's-eye view image is a virtual viewpoint image looking down on the subject from the Z axis, as shown in FIG. 3(B).

[0097] Note that the configuration is not limited to these as long as it is possible to operate the virtual camera or the time code, and the virtual camera operation area 502 and the time code operation area 503 do not have to be separated. For example, a double tap on the virtual camera operation area 502 may be processed as an operation on the time code, such as pausing. Note that although the case where the operation unit input unit 214 is a tablet 500 has been described, the operation and display device are not limited to this. For example, a double tap on the right half of the display area 502 may be set to fast forward 10 seconds, and a double tap on the left half may be set to rewind 10 seconds.

[0098] (Operation control processing) Next, a control process for a virtual camera according to this embodiment that effectively combines automatic operation and manual operation will be described with reference to the flowchart of FIG.

[0099] The virtual camera operations set for automatic operation and manual operation are determined by the operation setting process described above, and automatic operation is performed according to the automatic operation information described in Fig. 4, while manual operation is performed according to the manual setting details described in Fig. 5. This flowchart describes virtual camera control processing that meaningfully combines both types of operation. The flowchart shown in Fig. 6 is implemented by the CPU 211 using the RAM 212 as a workspace and executing a program stored in the ROM 213.

[0100] The operation control unit 203 executes processing related to the time code and playback speed in S602, and then executes processing for meaningfully combining automatic operation and manual operation in S604 to S608. The operation control unit 203 also executes the loop processing of S602 to S611 on a frame-by-frame basis. For example, if the frame rate of the output virtual viewpoint image is 60 FPS, one loop (one frame) of S602 to S611 is processed at intervals of approximately 16.6 ms. The interval of one loop may be achieved by setting the update rate (refresh rate) of the image display on a touch panel or the like in the image display device to 60 FPS and synchronizing the processing with that.

[0101] In S602, the operation control unit 203 counts up the time code. The time code according to this embodiment can be specified by "day:hour:minute:second:frame number" as explained in Fig. 1, and the count up is performed in units of frames. In other words, since the loop processing from S602 to S610 is performed in units of frames as explained above, the number of frames in the time code is counted up by one for each loop processing.

[0102] The count-up interval of the time code may be changed according to the selected value 514 of the playback speed slider described in Fig. 5. For example, if a playback speed of 1 / 2 is specified, the count-up may be performed by counting up one frame for every two loop processes of steps S601 to S610.

[0103] Next, the operation control unit 203 advances the process to S603, and acquires a three-dimensional model of the counted-up or specified time code via the model generation unit 205. The model generation unit 205 generates a three-dimensional model of the subject from the multi-viewpoint images, as described in FIG.

[0104] Next, in S604, the operation control unit 203 determines whether or not there is automatic operation information of the automatic operation unit 202 at the counted-up or specified time code. If there is automatic operation information at the time code (Yes in S604), the operation control unit 203 proceeds to S605, and if there is no automatic operation information (No in S604), the operation control unit 203 proceeds to S606.

[0105] In S605, the operation control unit 203 automatically operates the virtual camera using the automatic operation information at the corresponding time code. The automatic operation of the virtual camera has been described with reference to FIG. 4, so a description thereof will be omitted.

[0106] In S606, the operation control unit 203 accepts a manual operation from the user and switches processing depending on the accepted operation. If a manual operation regarding the time code has been accepted ("time code operation" in S606), the operation control unit 203 proceeds to S609. If a manual operation regarding the virtual camera has been accepted ("virtual camera operation" in S606), the operation control unit 203 proceeds to S607. If a manual operation has not been accepted ("none" in S606), the operation control unit 203 proceeds to S609.

[0107] In S607, the operation control unit 203 determines whether the virtual camera operation specified by the manual operation (S606) interferes with the virtual camera operation specified by the automatic operation (S605). The comparison of the operations may use the identifier of the virtual camera operation shown in Fig. 4(A), and may determine that the automatic operation and the manual operation interfere with each other if the identifiers are the same. Note that the comparison is not limited to the identifier comparison, and it may also be possible to determine whether or not there will be interference by comparing the amount of change in the position or attitude of the virtual camera.

[0108] If it is determined that the virtual camera operation specified by manual operation will interfere with the operation specified by automatic operation (Yes in S607), the operation control unit 203 proceeds to S609, where the accepted manual operation is canceled. In other words, if an automatic operation interferes with a manual operation, the automatic operation takes precedence. Note that parameters that do not change due to automatic operation are operated based on the manual operation. For example, if a translation of the focus point is performed by automatic operation, the position of the focus point changes but the attitude of the virtual camera does not. Therefore, horizontal or vertical rotation of the virtual camera, which does not change the position of the focus point but changes the attitude of the virtual camera, can be performed by manual operation because it does not interfere with the translation of the focus point. If it is determined that the specified virtual camera operation will not interfere with the operation specified by automatic operation (No in S607), the operation control unit 203 proceeds to S608, where it executes virtual camera control processing that combines automatic and manual operations.

[0109] For example, if the automatic operation specified in S605 is a translation of the point of interest (identifier=2), and the manual operation specified in S606 is a translation of the point of interest (identifier=2), and it is determined that viewing is to be performed, the manual translation of the point of interest is canceled. On the other hand, for example, if the automatic operation specified in S605 is a translation of the point of interest (identifier=2), and the manual operation specified in S606 is a vertical / horizontal rotation around the point of interest (identifier=3, 4), and it is determined that these are different operations. In this case, the combined operations are executed in S608, as will be described later with reference to FIG. 7.

[0110] Next, in S608, the operation control unit 203 executes manual operation of the virtual camera. The manual operation of the virtual camera has been described with reference to FIG. 5, and therefore its description will be omitted. Next, the operation control unit 203 advances the process to S609, where it generates and draws a virtual viewpoint image when capturing an image at a position and orientation of the virtual camera operated by at least one of automatic operation and manual operation. The drawing of the virtual viewpoint image has been described with reference to FIG. 2, and therefore its description will be omitted.

[0111] In S610, the operation control unit 203 updates the time code to be displayed to the time code manually specified by the user. The manual operation for specifying the time code has been described with reference to FIG. 5, and therefore will not be described here.

[0112] In S611, the operation control unit 203 determines whether or not display of all frames has been completed, i.e., whether or not the end time code has been reached. If the end time code has been reached (Yes in S611), the operation control unit 203 ends the processing shown in Fig. 6, and if the end time code has not been reached (No in S611), the processing returns to S602. An example of virtual camera control using the above flowchart and an example of generating a virtual viewpoint image will be described with reference to the following Fig. 7.

[0113] (Example of virtual camera control and virtual viewpoint image display) An example of controlling the virtual camera by combining automatic operation and manual operation according to this embodiment, and an example of displaying a virtual viewpoint image will be described with reference to FIG.

[0114] Here, as an example of controlling a virtual camera, we will explain a case where the automatic operation information described in Figures 4(C) and 4(D) is used for automatic operation, and the manual setting contents described in Figure 5(B) are used for manual operation.

[0115] In the automatic operation of Figures 4(C) and 4(D), the virtual camera is operated to translate (identifier = 2) so that the coordinates of the virtual camera's point of interest follow a trajectory 412 as the time code elapses without changing its posture. In the manual operation of Figure 5(B), the user operates the virtual camera to perform horizontal and vertical rotation around the point of interest (Figures 3(G) and 3(H), identifiers = 3, 4) by at least one of swiping and dragging.

[0116] For example, if the trajectory 412 in Fig. 4(D) is for tracking a ball in a field sport, the automatic operation of S605 will translate the virtual camera so that the ball is always captured at the point of interest. In addition to this operation, by manually controlling horizontal and vertical rotation around the point of interest, the user can unconsciously track the ball at all times and view the surrounding play from various angles.

[0117] This series of actions will be described with reference to Figures 7(A) to 7(D). In the example of Figure 7, the sport is rugby, and a scene is being captured of a player making an offload pass (a pass made before falling down). In this case, the subject is at least one of the rugby ball and a player on the field.

[0118] FIG. 7(A) shows a bird's-eye view of the three-dimensional space of the field, in which the focus point of the virtual camera is translated (identifier=2) to coordinates 701 to 703 by automatic operation (trajectory 412).

[0119] First, we will explain the case where the virtual camera's point of interest is at coordinate 701. In this case, the virtual camera is located at coordinate 711, and its orientation is approximately in the -X direction. The virtual viewpoint image in this case is one in which the play is viewed in the -X direction along the long side of the field, as shown in Figure 7(B).

[0120] Next, a case will be described in which the virtual camera's focus point is translated to coordinate 702 by automatic operation. Here, it is assumed that while the focus point is translated from coordinate 701 to 702 by automatic operation (identifier = 2), the user manually rotates the focus point vertically (identifier = 4) around the focus point. In this case, the position of the virtual camera is translated to coordinate 712, and its posture is rotated vertically in approximately the -Z direction. The virtual viewpoint image at this time is one looking down on the subject from above the field, as shown in FIG. 7(C).

[0121] Next, let us consider a case where the focus point of the virtual camera moves in parallel to coordinate 703. Here, it is assumed that while the focus point moves in parallel from coordinate 702 to 703 (identifier = 2), the user manually rotates the virtual camera horizontally (identifier = 3) around the focus point. In this case, the position of the virtual camera moves in parallel to coordinate 713, and its posture rotates approximately in the -Y direction. The virtual viewpoint image at this time is one in which the subject is viewed in the -Y direction from outside the field, as shown in FIG. 7(D).

[0122] As described above, by effectively combining automatic and manual operation, it is possible to automatically keep the area around the ball within the virtual camera's field of view during fast-paced scenes such as an offload pass in rugby, and to view it from various angles in response to manual operation. Specifically, automatic operation uses the ball's coordinates obtained from a 3D model and positioning tags to keep the area around the ball within the virtual camera's field of view at all times, without placing any operational burden on the user. Furthermore, as a manual operation, the user can simply perform a simple drag operation (rotating horizontally or vertically around a point of interest by dragging with one finger) to view noteworthy plays from various angles.

[0123] If no manual operation is performed during the section (trajectory 412) where automatic operation is performed, the virtual camera attitude is maintained at the attitude at which it was last manually operated. For example, if no manual operation is performed while the position of the point of interest is translated from coordinate 703 to 704 (identifier = 2) due to automatic operation, the virtual camera attitude remains the same (same attitude vector) at coordinates 713 and 714. Therefore, as shown in Figure 7(E), the virtual camera attitude is in the -Y direction, and from Figure 7(D) onwards, as long as time passes without manual operation, the image display device 104 displays the scene in which the player makes a successful off-load pass and sprints towards a try, in the same attitude.

[0124] 6 and 7, a process may be added to temporarily disable or enable the settings made by at least one of manual and automatic operations. For example, when an instruction to pause the playback of the virtual branch office image is received, the operation setting process made by manual and automatic operations may be reset and the position and coordinates of the default virtual camera may be set. The default virtual camera position may be changed, for example, by changing the position and coordinates of the virtual camera from above the field in the -Z direction so that the entire field fits within the field of view.

[0125] In another example, when an instruction to pause playback of a virtual viewpoint image is received, the content set by manual operation may be invalidated, and when the pause is released and playback is resumed, the content set by manual operation may be reflected again.

[0126] Furthermore, during automatic operation, parallel movement of the focus point (identifier = 2) is set for automatic operation, and if the same operation (identifier = 2) is manually operated and the manual operation is canceled in S607, manual operation of parallel movement of the focus point may be possible during pause regardless of the automatic operation information.

[0127] For example, consider the case where a pause is received during automatic playback when the virtual camera's focus point is at coordinate 701. The virtual viewpoint image at this time is as shown in Figure 7(B), but the user wants to check the state of the opposing team's player who made the tackle that triggered the offload pass, but the outside player is outside the field of view. In such a case, if the user pauses the playback and manually operates translational movement downward on the screen, the focus point can be translated to coordinate 705, etc.

[0128] Then, when the pause is released and normal playback resumes, the operation settings are enabled, the automatic parallel movement of the point of interest (identifier=2) is executed, and the manual parallel movement of the point of interest (identifier=2) may be canceled. At this time, the amount of parallel movement by the manual operation (the distance moved from coordinate 701 to coordinate 705) may also be canceled, and processing may be performed to return the point of interest to coordinate 701 specified by the automatic operation.

[0129] As described above, by controlling the virtual camera through a meaningful combination of automatic and manual operations, it is possible to keep the desired location of the subject within the field of view of the virtual camera, even in scenes where the subject moves quickly and in a complex manner, and to view the subject from various angles.

[0130] For example, in field sports like rugby, even in scenes where the play unfolds rapidly across a wide area of ​​the field and leads to a score, automatic operation can be used to keep the area around the ball within the virtual camera's field of view, and manual operation can be added to view it from various angles.

[0131] Second Embodiment In this embodiment, a description will be given of tag generation processing and tag reproduction processing in the image display device 104. The same reference numerals will be used for the configurations, processing, and functions as those in the first embodiment, including the virtual viewpoint image generation system in Fig. 1 and the image display device 104 in Fig. 2, and descriptions thereof will be omitted.

[0132] The tag in this embodiment allows the user to select a scene of interest, such as a sports score, and set automatic operation for the selected scene. The user can also perform additional manual operation during automatic playback using the tag, and can use the tag to execute a virtual camera control process that combines automatic and manual operation.

[0133] Such tags are used in sports program commentary, which is a typical use case of virtual viewpoint images. In such use cases, it is required to be able to view and comment on multiple important scenes, such as scores, from various positions and angles. In addition to the automatic operation of setting the positions of the ball and players as the focus point of the virtual camera as shown in the example of embodiment 1, this embodiment can also generate tags that set any position other than the ball, players, or other objects as the focus point, depending on the intention of the commentator, etc.

[0134] In this embodiment, the tag will be described as a tag in which an automatic operation for setting the position and orientation of a virtual camera is associated with a time code. However, in one example, a tag may be a tag in which a time code section (period) having a start and end time code is associated with a plurality of automatic operations within the time code section.

[0135] 1, in a configuration in which data of a notable period, such as scores, is copied from the database 103 to the image display device 104 for use, tags may be added to the copied time code section. This makes it easy to access only the copied data, and in sports programs and the like, the image display device 104 can be used for commentary using virtual viewpoint images of important scenes.

[0136] Tags in this embodiment are managed by a tag management unit 204 in the image display device 104 (FIG. 2), and tag generation and tag playback processes are executed by an operation control unit 203. Each process is realized based on the automatic operation in FIG. 4 and the manual operation in FIG. 5 described in the first embodiment, and the virtual camera control process that combines the automatic and manual operations in FIG. 6, etc.

[0137] (Tag generation process) Next, a method for generating a tag in the image display device 104 will be described with reference to FIG. 8(A).

[0138] 8A has the same configuration as the manual operation screen of the image display device 104 described in FIG. 5A, and the touch panel 501 is roughly divided into a virtual camera operation area 502 and a time code operation area 503. In FIG.

[0139] In user operations for tag generation, the user first selects a time code in the time code operation area 503, and then specifies the virtual camera operation for the time code in the virtual camera operation area 502. By repeating these two steps, it is possible to create automatic operations for any time code section.

[0140] When generating a tag, the operation for specifying the time code in the time code operation area 503 is the same as that explained in Fig. 5(A). That is, the time code for which you want to generate a tag can be specified by dragging the knob on the main slider or by dragging the knob on the sub-slider 811.

[0141] It should be noted that when specifying a time code, a rough time code can be specified on the main slider 512, followed by a pause instruction, and then the time code can be specified in detail using the sub-slider 513, which displays a range of several tens of seconds before and after the specified time code. It should be noted that the sub-slider 513 may be displayed when a tag generation instruction is given.

[0142] When the user selects a time code, a virtual viewpoint image (bird's-eye view image) for that time code is displayed. In virtual camera operation area 502 in Fig. 8(A), a virtual viewpoint image that overlooks the subject is displayed. This is because, while viewing the bird's-eye view image that displays a wide range of the subject, it is easy to grasp the playing situation and to specify a point of interest.

[0143] Next, while visually observing the overhead image (virtual viewpoint image) at that time code being displayed, the user specifies the point of interest of the virtual camera at that time code by tapping 812. The image display device 104 converts the tapped position into coordinates on the plane of interest of the virtual camera in three-dimensional space, and records those coordinates as the coordinates of the point of interest at that time code. Note that in this embodiment, the XY coordinates are the tapped position, and the Z coordinate is assumed to be Z=1.7 m. However, the Z coordinate may also be obtained based on the tapped position in the same way as the XY coordinates. This method of coordinate conversion from the tapped position is well known, and a description thereof will be omitted.

[0144] By repeating the above two steps (821 and 822 in FIG. 8), it is possible to specify the focus point of the virtual camera in a given time code section in chronological order.

[0145] Note that the time codes specified by the above operations may not be continuous and may have gaps, but the coordinates between them may be interpolated using manually specified coordinates. Interpolation between multiple points is performed using various curve functions (spline curves, etc.) and will not be described here.

[0146] The tag generated by the above procedure has the same format as the automatic operation information explained in Fig. 4(A). Therefore, the tag generated by this embodiment can be used in the virtual camera control process that meaningfully combines automatic operation and manual operation explained in the first embodiment.

[0147] The virtual camera operation that can be specified by the tag is not limited to the coordinates of the point of interest, and may be any of the virtual camera operations shown in the first embodiment.

[0148] Using the above method, it is possible to generate multiple tags for multiple scoring scenes in a single game, each with its own unique tag. It is also possible to generate multiple tags with different automatic operation information for the same time code.

[0149] It is also possible to provide a button for switching between the normal virtual camera operation screen described in embodiment 1 and the tag generation operation screen of this embodiment. Although not shown, when the operation screen is switched to the tag generation operation screen, a process may be performed to move the virtual camera to a position and posture that provides a bird's-eye view of the subject.

[0150] Using the tag generation process described above, it is possible to easily generate tags for automatic operations in any time code section where a notable scene, such as a score, occurred in a sports game. Next, we will explain the tag playback process, which allows the user to select a tag generated in this way and easily play it automatically.

[0151] (Tag regeneration processing) Next, tag reproduction in the image display device 104 will be described with reference to the flowchart in FIG. 9 and the operation screens in FIGS. 8(B) and 8(C).

[0152] Fig. 9 is a flowchart showing an example of virtual camera control processing executed by the image display device 104 during tag playback. In Fig. 9, similar to Fig. 6, the operation control unit 203 repeatedly executes the processing of S902 to S930 on a frame-by-frame basis. Each time processing is performed, the number of frames in the time code is counted up by one, thereby processing consecutive frames. Note that a description of steps that are the same as those in the flowchart of the virtual camera control processing in Fig. 6 will be omitted here.

[0153] In S902, the operation control unit 203 performs the same process as in S603 to obtain a three-dimensional model of the current time code.

[0154] In step S903, the operation control unit 203 determines the display mode. The display mode is the display mode on the operation screen of Fig. 5(A), and in this case, it determines whether it is normal playback mode or tag playback mode. If the display mode determination result is normal playback, the process proceeds to S906, and if it is tag playback, the process proceeds to S904.

[0155] The display modes to be determined are not limited to these. For example, the tag generation described in Fig. 8(A) may also be included in the mode determination.

[0156] Next, the image display device 104 advances the process to S904, and the operation control unit 203 reads the automatic operation information specified by the time code in the selected tag because the display mode is tag playback. As explained above, the automatic operation information is that shown in Figure 4(C) etc. Next, the image display device 104 advances the process to S905, and the operation control unit 203 performs the same process as S605 to automatically operate the virtual camera.

[0157] Next, the image display device 104 proceeds to S906, where the operation control unit 203 accepts a manual operation from the user and switches processing depending on the accepted operation. If a manual operation regarding the time code is accepted from the user ("Time Code Operation" in S906), the image display device 104 proceeds to S912. If a manual operation regarding the virtual camera is accepted from the user ("Virtual Camera Operation" in S906), the image display device 104 proceeds to S922. If a manual operation is not accepted from the user ("None" in S906), the image display device 104 proceeds to S907. If a tag selection operation is accepted from the user ("Tag Selection Operation" in S906), the image display device 104 proceeds to S922. The manual operation screen at this time will be described with reference to FIGS. 8(B) and 8(C).

[0158] Note that the contents of the manual operation to be determined in S906 are not limited to these, and the tag generation operation described in FIG. 8(A) may also be included.

[0159] Fig. 8(B) is an extracted view of the time code operation area 503 on the manual operation screen 801. Multiple tags (831 to 833) are displayed in a distinguishable manner on the main slider 512 for operating the time code in Fig. 8(B). The tags (831 to 833) are displayed at positions on the main slider corresponding to the start time codes of the tags. The displayed tags may be tags generated by the image generation device 104 shown in Fig. 8(A), or tags generated by another device may be read.

[0160] Here, the user selects one of the multiple tags (831 to 833) displayed on the main slider 512 by tapping 834. When a tag is selected, automatic operation information set for the tag can be listed as shown in balloon 835. In balloon 835, two automatic operations are set for the time code, and two corresponding tag names (Auto1, Auto2) are displayed. Note that the number of tags that can be set for the same time code is not limited to two. When the user selects one of tags 831 to 833 by tapping 834, the image display device 104 advances the process to S922.

[0161] In S922, the operation control unit 203 sets the display mode to tag playback, and the process proceeds to S923.

[0162] In S923, the operation control unit 203 reads the start time code specified by the tag selected in S906, and proceeds to S912. Then, in S912, the operation control unit 203 updates the time code to the specified time code, and returns the process to S902. Then, in subsequent S904 and S905, the image display device 104 automatically operates the virtual camera using the automatic operation information specified by the selected tag.

[0163] Here, the screen during tag playback (automatic playback after tag selection) is shown in Fig. 8(C). In Fig. 8(C), a sub-slider 843 is displayed in the time code operation area 503. The scale of the sub-slider 843 may be the length of the time code section specified in the tag. Also, the sub-slider 843 may be displayed dynamically, for example, when tag playback is instructed. Also, during tag playback, the tag being played may be displayed as a display 832 that distinguishes it from other tags.

[0164] 8(C) displays a virtual viewpoint image generated using a virtual camera specified by the automatic operation information included in the selected tag. During tag playback, manual operation can also be accepted, and in S907 and S908, the image display device 104 executes virtual camera control processing that combines automatic operation and manual operation, similar to that described in the first embodiment.

[0165] In S907, when the operation control unit 203 receives a manual operation related to the virtual camera, it checks the current display mode. If the display mode is the normal playback mode ("normal playback" in S907), the operation control unit 203 proceeds to S909, and if the display mode is the tag playback mode ("tag playback" in S907), it proceeds to S908.

[0166] In S908, the operation control unit 203 compares the automatic operation of the selected tag with the manual operation of the virtual camera received in S906, and determines whether or not there is interference. The method of comparing the automatic operation with the manual operation is the same as in S607, so a description thereof will be omitted.

[0167] In S909, the operation control unit 203 manually operates the virtual camera in the same manner as in S608. In S910, the operation control unit 203 draws a virtual viewpoint image in the same manner as in S609. In S911, the operation control unit 203 counts up the time code. The counting up of the time code is the same as in S602, so a description thereof will be omitted.

[0168] In S921, the operation control unit 203 determines whether the time code counted up in S911 during tag playback is the end time code specified in the tag. If the counted up time code is not the end time code specified in the tag (No in S921), the operation control unit 203 proceeds to S930 and makes the same determination as in S611. If the counted up time code is the end time code specified in the tag (Yes in S921), the operation control unit returns to S923, reads the start time code of the selected tag, and performs repeat playback to repeat tag playback. In this case, the manual operation set in the previous playback may be reflected as an automatic operation.

[0169] Note that the user may press the cancel button 844 shown in FIG. 8(C) to cancel repeat playback during tag playback.

[0170] When repeating tag playback, the posture last manually operated may be maintained during the repeat. When performing operations using the automatic operation information of Figures 4(C) and 4(D) as tags, if repeat playback is performed in the time code section (trajectory 412, start point 700, end point 710), the posture at end point 710 may be maintained even when returning to start point 700. Alternatively, when returning to start point 700, the posture may be returned to the initial value specified in the automatic operation information.

[0171] As described above, the tag generation method according to this embodiment allows the image display device 104 to easily generate automatic operation information for a predetermined range of time codes in a short time.

[0172] Furthermore, the tag playback method allows the user to perform automatic playback for each scene with a simple operation of simply selecting a desired tag from multiple tags. Even if the subject's movements vary and are complex for each scene, automatic operation allows the virtual camera to keep capturing the position of interest for each scene within its viewing angle. Furthermore, by adding manual operation to the automatic operation, the user can display the subject from the user's desired angle. Furthermore, since the user can add automatic operations as tags, the user can combine operations of multiple virtual cameras to suit their preferences.

[0173] For example, for multiple scoring scenes in a sports broadcast, tags can be easily generated for multiple scenes by specifying time codes using time code operations and adding tags by specifying the coordinates of the virtual camera's focal point. In subsequent sports commentary programs, commentators can use these tags to easily display multiple scoring scenes from various angles and use them in their commentary. For example, as shown in Figures 7(B) to 7(D) of the first embodiment, by displaying a scene of interest from various angles, it is possible to provide detailed commentary on notable points.

[0174] Furthermore, tags are not limited to sports commentary. For example, a sports expert or famous athlete who is knowledgeable about the subject of a video can create tags for scenes that they find noteworthy, and other users can use those tags for tag playback. With just a few simple operations, users can enjoy notable scenes from various angles using tags set by highly specialized people.

[0175] Furthermore, the generated tag can be transmitted to another image display device, and the tag can be used for automatic playback on the other image display device.

[0176] <Third embodiment> In this embodiment, an example of manual operation of the virtual camera in the image display device 104 will be described. Compared to a method of switching the virtual camera operation in response to switching touch operations based on the number of fingers, the virtual camera operation can be switched with one finger, making it easier for the user to perform manual operation.

[0177] Descriptions of the same configurations, processes, and functions as those in the first and second embodiments will be omitted. The virtual viewpoint image generation system in Fig. 1 and the functional blocks of the image display device 104 in Fig. 2 have been explained in the first embodiment, so explanations will be omitted and only manual operations will be explained.

[0178] Manual operation of the virtual camera of this embodiment will be described with reference to Fig. 10. The screen configuration of Fig. 10, similar to Fig. 5 described in the first embodiment, has a virtual camera operation area 502 and a time code operation area 503 displayed on the touch panel 1001. The components of the time code operation area 503 are the same as those described with reference to Fig. 5, so description thereof will be omitted.

[0179] In this embodiment, touch operations in the virtual camera operation area 1002 for two virtual camera operations will be described when a touch panel is used for the operation input unit 214 of the image display device 104. One operation is translation of the focus point of the virtual camera shown in FIG. 3(F), and the other operation is zooming in and out (not shown) of the virtual camera.

[0180] In this embodiment, the touch operation for translating the focus point of the virtual camera is a tap operation 1021 (or a double tap operation) on an arbitrary position in the operation area. The image display device 104 converts the tapped position into coordinates on the focus plane of the virtual camera in three-dimensional space and translates the virtual camera so that the coordinates become the focus point of the virtual camera. For the user, this operation is equivalent to moving the virtual camera so that the tapped position becomes the center of the screen. By simply repeating the tap operation, the user can continue to display the desired focus point on the subject. Note that, in one example, the virtual camera may be translated at a predetermined speed from the coordinates 1020 of the current focus point of the virtual camera to a point specified by the tap operation 1021. Note that the speed of movement from the coordinates 1020 to the point specified by the tap operation 1021 may be constant. In this case, the speed of movement may be determined so that the movement is completed in a predetermined time, such as one second. Alternatively, the speed of movement may be proportional to the distance to the point specified by the tap operation 1021, and may become slower as the user approaches the point specified by the tap operation 1021.

[0181] However, if parallel movement is performed by a drag operation (swipe operation) using two fingers, it is difficult to adjust the amount of movement of the drag operation when the subject moves quickly.In contrast, in this embodiment, you can simply intuitively tap the destination you want to move to.

[0182] The touch operation for enlarging or reducing the virtual camera is a drag operation 1022 on a predetermined area on the right edge of the virtual camera operation area 1002. For example, a pinch-out is performed by dragging (swiping) the predetermined area on the edge (edge) of the virtual camera operation area 1002 upward, and a pinch-in is performed by dragging (swiping) it downward. Note that in the example of FIG. 10 , the pinch-in is performed on the right edge of the screen of the virtual camera operation area 1002, but it may be on the edge of the screen in another direction, and the drag direction is not limited to up and down.

[0183] Instead of pinching in and out with two fingers, zooming in and out can be done simply by dragging one finger along the edge of the screen. For example, dragging upward on the edge of the screen zooms in, and dragging downward on the edge of the screen zooms out.

[0184] The operations described above can be set to the manual settings described in FIG. 5, and can be used for the virtual camera control process that combines the automatic operation and manual operation described in the first embodiment.

[0185] As described above, by using the manual operation according to this embodiment, the user does not need to frequently switch the number of fingers used to operate the virtual camera, and can intuitively operate the image display device 104 with one finger. Furthermore, in a translation operation, translation to a predetermined position can be performed by tapping to specify a destination point, so translation can be instructed more intuitively than with operations such as dragging or swiping that require adjustment of the amount of operation.

[0186] <Fourth embodiment> In this embodiment, an example of automatic operation of a virtual camera in the image display device 104 will be described.

[0187] An example of automatic operation information is one that uses the same items as in Figure 4(C) and uses constant values ​​for both the coordinates and orientation of the virtual camera's focus point during a certain section of the time code. For example, in sports, a game may be interrupted due to a player's injury or foul, and such periods can be boring for viewers. When such an interruption is detected, the virtual camera may be automatically positioned and oriented to view the entire field and the time code may be advanced at double speed. Alternatively, it is possible to specify that the entire interrupted period be skipped. In this case, the interruption in the game may be input by an operator or specified by a user using a tag. By performing such automatic operations, the user can fast-forward or skip over boring periods during which play is interrupted.

[0188] As another example, in a scene where the action unfolds very quickly, the playback speed for a given section of the time code may be changed. For example, by setting the playback speed to a slow value such as half speed, the user can easily manually operate the virtual camera, which would be difficult at normal speed.

[0189] In the case of sports, a pause in play may be detected by image analysis based on the referee's gestures or the halt of the ball or player's movement. On the other hand, image analysis may be used to detect that play is progressing quickly based on the movement speed of the players or ball. Furthermore, audio analysis may be used to detect pauses or rapid progression of play.

[0190] <Other embodiments> The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0191] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0192] 104: Image generating device, 110: Virtual camera, 201: Manual operation unit, 202: Automatic operation unit, 203: Operation control unit, 204: Tag management unit, 206: Image generating unit

Claims

1. a first acquisition means for acquiring point of interest information indicating the position of a point of interest of a virtual camera corresponding to each time within a predetermined period of time and attitude information indicating the attitude of the virtual camera; a second acquisition means for acquiring an input corresponding to a user operation for changing the attitude of the virtual camera at the predetermined time; a generating means for generating a first virtual viewpoint image based on a plurality of captured images, the attention point information, the attitude information, and the input, by changing the attitude of the virtual camera indicated by the attitude information in accordance with the input, without changing the position of the attention point of the virtual camera indicated by the attention point information; An information processing device comprising:

2. The information processing device described in claim 1, characterized in that before accepting the input, the generation means generates a second virtual viewpoint image based on the position of the focus point of the virtual camera indicated by the focus point information and the attitude of the virtual camera indicated by the attitude information, and after accepting the input, generates the first virtual viewpoint image.

3. The point of interest of the virtual camera is a point on the optical axis of the virtual camera, 2. The information processing device according to claim 1, further comprising a viewpoint control means for controlling the change of the position of the virtual camera so that, in the attitude of the virtual camera changed by the input, the position of the focus point of the virtual camera indicated by the focus point information is positioned on the optical axis of the virtual camera.

4. 4. The information processing apparatus according to claim 3, wherein said viewpoint control means controls to change the position of said virtual camera to a position on a circle having a point of interest of said virtual camera as a center.

5. The information processing device according to claim 1, characterized in that the generating means generates the first virtual viewpoint image based on the position of the focus point of the virtual camera indicated by the focus point information and the attitude of the virtual camera changed by the input.

6. a first acquisition means for acquiring point of interest information indicating the position of a point of interest of a virtual camera corresponding to each time within a predetermined period of time and attitude information indicating the attitude of the virtual camera; a second acquisition means for acquiring an input corresponding to a user operation for changing the attitude of the virtual camera at the predetermined time; an output means for outputting information indicating the position of the point of interest of the virtual camera and the attitude of the virtual camera after the change, by changing the attitude of the virtual camera indicated by the attitude information based on the point of interest information, the attitude information, and the input, without changing the position of the point of interest of the virtual camera indicated by the point of interest information; An information processing device comprising:

7. a receiving means for receiving an input corresponding to a user operation for changing the attitude of the virtual camera; a display means for displaying a first virtual viewpoint image, the first virtual viewpoint image being generated by changing the attitude of the virtual camera based on the input without changing the position of the focus point of the virtual camera corresponding to each time point during a predetermined time period that has been acquired in advance, based on a plurality of captured images and the input; An information processing device comprising:

8. a first acquisition step of acquiring point of interest information indicating the position of a point of interest of a virtual camera corresponding to each time point within a predetermined period of time and attitude information indicating the attitude of the virtual camera; a second acquisition step of acquiring an input corresponding to a user operation for changing the attitude of the virtual camera at the predetermined time; a generating step of generating a first virtual viewpoint image based on a plurality of captured images, the attention point information, the attitude information, and the input by changing the attitude of the virtual camera indicated by the attitude information without changing the position of the attention point of the virtual camera indicated by the attention point information; An information processing method comprising:

9. a first acquisition step of acquiring point of interest information indicating the position of a point of interest of a virtual camera corresponding to each time point within a predetermined period of time and attitude information indicating the attitude of the virtual camera; a second acquisition step of acquiring an input corresponding to a user operation for changing the attitude of the virtual camera at the predetermined time; an output step of outputting information indicating the position of the point of interest of the virtual camera and the attitude of the virtual camera after the change, by changing the attitude of the virtual camera indicated by the attitude information based on the point of interest information, the attitude information, and the input, without changing the position of the point of interest of the virtual camera indicated by the point of interest information; An information processing method comprising:

10. a receiving step of receiving an input corresponding to a user operation for changing the attitude of the virtual camera; a display step of displaying a first virtual viewpoint image generated by changing the attitude of the virtual camera based on the input without changing the position of the focus point of the virtual camera corresponding to each time point at a predetermined time acquired in advance, based on the plurality of captured images and the input; An information processing method comprising:

11. A program for causing a computer to execute the information processing method according to any one of claims 8 to 10.

Citation Information

Patent Citations

  • Operation control for refrigerator

    JP1989019278A