Image display device, control method and non-transitory computer-readable storage medium
By combining user input and pre-set parameters, and utilizing image display devices and control methods, the problem of virtual camera control in virtual viewpoint image generation was solved, achieving flexible virtual viewpoint image generation and high-quality display.
Patent Information
- Application Number
- CN202110472533.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-08
- Filing Date
- 2021-04-29
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2041-04-29
AI Technical Summary
Existing technologies make it difficult for users to effectively control the virtual camera when generating virtual viewpoint images, especially in cases of complex changes or significant movement.
An image display device and control method are provided. By accepting user-input parameters and pre-set parameters, and combining automatic and manual operations, a virtual viewpoint image is generated. The device includes a manual operation unit, an automatic operation unit, an operation control unit, a label management unit, and a model generation unit. The generation of the virtual viewpoint image is controlled by the position and posture of a 3D model and a virtual camera.
It improves the convenience of browsing virtual viewpoint images and enables flexible control of the virtual camera. Users can easily specify the pose and target point of the virtual viewpoint through a combination of manual and automatic operation to generate high-quality virtual viewpoint images.
Smart Images

Figure CN113630547B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an image display device, control method, and non-transitory computer-readable storage medium based on a virtual viewpoint image. Background Technology
[0002] In recent years, a technique for generating virtual viewpoint images has attracted attention, in which arbitrary viewpoint images are generated from multiple images captured by multiple cameras. A virtual camera is used as a convenient way to describe the concept of a virtual viewpoint specified for generating virtual viewpoint images. Unlike physical cameras, virtual cameras can exhibit various behaviors in three-dimensional space without physical limitations, such as movement / rotation / zooming in / out. To properly control the virtual camera, multiple operational methods corresponding to each behavior can be considered.
[0003] Japanese Patent Application Publication No. 6419278 discloses a method for displaying a virtual viewpoint image on an image display device including a touch panel, wherein the method is used to control the movement of a virtual camera based on the number of fingers performing a touch operation on the touch panel.
[0004] The method disclosed in Japanese Patent Application Publication No. 6419278 is difficult to operate when the virtual viewpoint is expected to be moved significantly or when complex changes are expected, depending on the anticipated behavior of the virtual viewpoint. Summary of the Invention
[0005] One aspect of this disclosure is to eliminate the problems that arise from using conventional technologies.
[0006] The present disclosure is characterized by providing a technique for improving convenience when browsing virtual viewpoint images.
[0007] According to a first aspect of this disclosure, an image display apparatus is provided, the image display apparatus comprising: a receiving unit configured to receive a user operation for determining a first parameter, the first parameter specifying at least one of a pose viewed from a virtual viewpoint and a position of a target point in a virtual viewpoint image; an acquiring unit configured to acquire a second parameter, the second parameter being preset and specifying at least one of a pose viewed from a virtual viewpoint and a position of a target point in the virtual viewpoint image; and a display unit configured to display a virtual viewpoint image generated based on at least one of the first parameter and the second parameter on the display unit.
[0008] According to a second aspect of this disclosure, a control method for an image display device is provided, the control method comprising: receiving a user operation for determining a first parameter, the first parameter specifying at least one of a pose viewed from a virtual viewpoint and the position of a target point in a virtual viewpoint image; obtaining a second parameter, the second parameter being preset and specifying at least one of a pose viewed from a virtual viewpoint and the position of a target point in the virtual viewpoint image; and displaying a virtual viewpoint image generated based on at least one of the first parameter and the second parameter to a display unit.
[0009] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided for storing a program that causes a computer having a display unit to perform a control method for an image display device, the control method comprising: receiving a user operation for determining a first parameter, the first parameter specifying at least one of a pose viewed from a virtual viewpoint and a position of a target point in a virtual viewpoint image; obtaining a second parameter, the second parameter being preset and specifying at least one of a pose viewed from a virtual viewpoint and a position of a target point in the virtual viewpoint image; and displaying a virtual viewpoint image generated based on at least one of the first parameter and the second parameter to the display unit.
[0010] Other features of the invention will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0011] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments of the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.
[0012] Figure 1A This is a configuration diagram of the virtual viewpoint image generation system.
[0013] Figure 1B This is a diagram showing an example of the installation of multiple sensor systems 101.
[0014] Figure 2A This is a diagram illustrating the function of an image display device.
[0015] Figure 2B This is a diagram illustrating the structure of an image display device.
[0016] Figure 3A This is a diagram showing the coordinate system of the virtual camera.
[0017] Figure 3B This is a diagram showing the location of the virtual camera.
[0018] Figure 3C This is a diagram showing the pose of the virtual camera.
[0019] Figure 3D This is a diagram showing the movement of a virtual camera.
[0020] Figure 3E This is a diagram showing the movement of a virtual camera.
[0021] Figure 3F This is a diagram showing the movement of a virtual camera.
[0022] Figure 3G This is a diagram showing the movement of a virtual camera.
[0023] Figure 3H This is a diagram showing the movement of a virtual camera.
[0024] Figure 4A This is a diagram illustrating an example of operation of a virtual camera that operates automatically.
[0025] Figure 4B This is a diagram showing an example of a virtual camera's settings.
[0026] Figure 4C This is a diagram illustrating an example of automated operation information.
[0027] Figure 4D This is a diagram illustrating an example of automated operation information.
[0028] Figure 4E This is a diagram showing an example of a virtual viewpoint image.
[0029] Figure 4F This is a diagram showing an example of a virtual viewpoint image.
[0030] Figure 4G This is a diagram showing an example of a virtual viewpoint image.
[0031] Figure 5A This is a diagram showing the manual operation screen.
[0032] Figure 5B This is a diagram showing manual operation information.
[0033] Figure 6 This is a flowchart illustrating a processing example for operations used with a virtual camera.
[0034] Figure 7A This is a diagram illustrating an example of the operation.
[0035] Figure 7B This is a diagram showing an example.
[0036] Figure 7C This is a diagram showing an example.
[0037] Figure 7D This is a diagram showing an example.
[0038] Figure 7E This is a diagram showing an example.
[0039] Figure 8A This is a diagram showing the label generated according to the second embodiment.
[0040] Figure 8B This is a diagram illustrating a label reproduction method according to a second embodiment.
[0041] Figure 8C This is a diagram illustrating a label reproduction method according to a second embodiment.
[0042] Figure 9 This is a flowchart illustrating a processing example for the operation of a virtual camera according to a second embodiment.
[0043] Figure 10 This is a diagram illustrating a manual operation example according to the third embodiment. Detailed Implementation
[0044] In the following, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments are not intended to limit the scope of the claimed disclosure. Several features are described in the embodiments, but it is not required that this disclosure satisfy all features, and multiple such features may be appropriately combined. Furthermore, in the drawings, the same reference numerals are given the same or similar configuration, and redundant descriptions thereof are omitted.
[0045] <First Embodiment>
[0046] In this embodiment, a system is described that generates a virtual viewpoint image representing a view seen from a specified virtual viewpoint, based on multiple images captured by multiple camera devices and a specified virtual viewpoint. The virtual viewpoint image in this embodiment is not limited to an image corresponding to a viewpoint freely (arbitrarily) specified by the user; it can be, for example, an image corresponding to a viewpoint selected by the user from multiple candidates. This embodiment primarily describes the case where the virtual viewpoint is specified via user operation, but the virtual viewpoint can be automatically specified based on image analysis results, etc.
[0047] In this embodiment, the term "virtual camera" will be used for description. A virtual camera is a virtual camera distinct from multiple actual camera devices installed around the image capture area, and is a concept used to conveniently explain the virtual viewpoint used in generating virtual viewpoint images. That is, a virtual viewpoint image can be considered as an image captured from a virtual viewpoint set in a virtual space associated with the capture area. The position and orientation of the viewpoint in the virtual camera can be represented as the position and orientation of the virtual camera. In other words, if we assume the camera is located at a virtual viewpoint set in space, the virtual viewpoint image can be described as an image simulating the captured image that will be obtained by that camera.
[0048] Regarding the order of explanation, it will refer to Figure 1A and Figure 1B Describe the entire virtual viewpoint image generation system, referring to Figure 2A and Figure 2B The image display device 104 is described. The image display device 104 performs virtual camera control processing that meaningfully combines automatic and manual operation. It will use... Figure 6 Describe the processing used to control the virtual camera, and will... Figures 7A to 7E The document describes examples of controlling a virtual camera, as well as examples of using a virtual camera to display virtual viewpoint images.
[0049] They will be in Figures 3A to 3H , Figures 4A to 4G , Figure 5A and Figure 5B This describes the configuration of the virtual camera used for illustration, as well as automatic and manual processing.
[0050] (Configuration of the virtual viewpoint image generation system)
[0051] First, refer to Figure 1A and Figure 1B The configuration of the virtual viewpoint image generation system 100 according to this embodiment is described.
[0052] The virtual viewpoint image generation system 100 includes n sensor systems, from sensor system 101a to sensor system 101n, and each sensor system includes at least one camera as an imaging device. In the following text, unless otherwise specified, "sensor system 101" will refer to multiple sensor systems 101 without distinction between the n sensor systems.
[0053] Figure 1BThis is a view illustrating an example of the installation of multiple sensor systems 101. Multiple sensor systems 101 are installed around an area 120, which serves as the target area for imaging, and each sensor system 101 captures images of the area 120 from a different direction. In this example embodiment, it is assumed that the target area 120 is the field of a stadium where a football match is to be played, and it is described as having n (e.g., 100) sensor systems 101 installed around the field. However, the number of sensor systems 101 to be installed is not limited, and the target area 120 is not limited to a sports field. For example, the stadium stands may be included in area 120, and area 120 may be an indoor studio or a stage, etc.
[0054] Furthermore, the multiple sensor systems 101 can be configured such that they are not installed on the entire circumference of the coverage area 120, and can be configured such that, depending on the installation location and other limitations, the multiple sensor systems 101 are installed only in a portion of the periphery of the area 120. Additionally, among the multiple cameras included in the multiple sensor systems 101, imaging devices with different functions, such as telephoto cameras and wide-angle cameras, can be included.
[0055] The multiple cameras included in the multiple sensor system 101 perform shooting synchronously with each other. The multiple images obtained by these cameras are referred to as multi-view images. Note that each multi-view image in this embodiment can be a captured image, and can be an image obtained by performing image processing on the captured image, such as processing for extracting a predetermined region.
[0056] Note that, in addition to the camera, the multiple sensor systems 101 may also include microphones (not shown). Each microphone in the multiple sensor systems 101 collects audio synchronously with each other. An audio signal that is reproduced together with the image display in the image display device 104 can be generated based on the collected audio. In the following description, the description of the audio is omitted for ease of description, but it is generally assumed that the image and audio are processed simultaneously.
[0057] The image recording device 102 acquires multi-view images from multiple sensor systems 101 and stores them in a database 103 along with a timecode used for capturing the image. This timecode is information used to uniquely identify the time when the camera device captured the image, and can be specified, for example, in the form of "day:hour:minute:second.frame number".
[0058] The image display device 104 provides images based on multi-view images with timecodes from the database 103 and the user's manual operation relative to the virtual camera.
[0059] A virtual camera 110 is positioned within a virtual space associated with region 120 and can view region 120 from a viewpoint different from any of the cameras in the plurality of sensor systems 101. (See below for reference.) Figures 3C to 3H Describe the virtual camera 110 and its actions in detail.
[0060] In this embodiment, both automatic and manual operation are used for the virtual camera 110, but the following description... Figure 4A This describes the automatic operation information for the virtual camera. Additionally, see below for reference. Figure 5A and Figure 5B This describes the manual operation method for the virtual camera 110.
[0061] The image display device 104 performs virtual camera control processing, which meaningfully combines automatic and manual operation and generates virtual viewpoint images from multi-viewpoint images based on the controlled virtual camera and timecode. (Refer to...) Figure 6 Describe the virtual camera control process in detail, and refer to Figures 7A to 7E Provide a detailed description of its control examples and examples of virtual viewpoint images generated using a virtual camera.
[0062] The virtual viewpoint image generated by the image display device 104 is an image representing the view from the virtual camera 110. In this embodiment, the virtual viewpoint image is also called a free-viewpoint video and is displayed on a touch panel such as a liquid crystal display of the image display device 104.
[0063] Note that the configuration of the virtual viewpoint image generation system 100 is not limited to... Figure 1A The configuration shown can be such that the image display device 104 is separated from the operation device or the display device. Alternatively, it can be such that multiple display devices are connected to the image display device 104 and each outputs a virtual viewpoint image.
[0064] Note that in Figure 1A The example illustrates that database 103 and image display device 104 are different devices, but they can be configured to be integrated. Alternatively, it could be configured so that important scenes from database 103 are pre-copied to image display device 104. Furthermore, it can be configured to switch between allowing access to all timecodes in the match or only certain timecodes, depending on the settings that allow database 103 and image display device 104 to access them. By configuring the copying of some data, it is possible to configure the system to allow access only to the copied portion of the timecodes.
[0065] Note that, although this embodiment is described with the example of a moving image as the virtual viewpoint image, the virtual viewpoint image can be a still image.
[0066] (Functional configuration of the image display device)
[0067] Next, we will refer to Figure 2A and Figure 2B The configuration of the image display device 104 is explained.
[0068] Figure 2A This is a view illustrating an example of the functional configuration of the image display device 104. The image display device 104 uses... Figure 2A The functions shown in the figure are implemented, and virtual camera control processing is performed that clearly combines automatic and manual operation. The image display device 104 includes a manual operation unit 201, an automatic operation unit 202, an operation control unit 203, a tag management unit 204, a model generation unit 205, and an image generation unit 206.
[0069] The manual operation unit 201 is a receiving unit for accepting input information manually input by the user for the timecode virtual camera 110. Although manual operation of the virtual camera includes operation of at least one of a touch panel, joystick, or keyboard, the manual operation unit 201 can obtain input information via operation of other input devices. (See also...) Figure 5A and Figure 5B Describe the manual operation screen in detail.
[0070] The automatic operation unit 202 automatically operates the virtual camera 110 and the timecode. (See reference...) Figure 4A The following describes in detail an example of automatic operation in this embodiment. In addition to manual user operation, the image display device 104 also operates or sets the virtual camera or timecode. For the user, the experience is as if the virtual viewpoint image is automatically operated, unlike the operation performed by the user themselves. The information used by the automatic operation unit 202 for automatic operation is called automatic operation information, and for example, uses the coordinates and position information of a three-dimensional model related to the player and ball, as measured via GPS or the like. Note that the automatic operation information is not limited to these, but only requires specifyable information related to the virtual camera and timecode that is independent of user operation.
[0071] The operation control unit 203 meaningfully combines the user input information obtained by the manual operation unit 201 from the manual operation screen with the automatic operation information used by the automatic operation unit 202 to control the virtual camera or timecode.
[0072] The following is for reference. Figure 6 The operation control unit 203 is described in detail, and references are provided. Figures 7A to 7E An example describing a virtual viewpoint image generated by a controlled virtual camera.
[0073] Additionally, in the operation control unit 203, the virtual camera motion is set as the operation target for each of automatic and manual operations. This is called motion setting processing, and will be referred to below. Figure 4A and Figure 4B Provide a detailed description.
[0074] The model generation unit 205 generates a 3D model representing the 3D topology of the subject within region 120 based on multi-view images obtained by specifying timecodes from database 103. Specifically, it obtains foreground images from the multi-view images, from which foreground regions corresponding to subjects such as spheres and people are extracted, and background images from which background regions other than the foreground regions are extracted. Then, the model generation unit 205 generates a foreground 3D model based on the multiple foreground images.
[0075] The 3D model is generated using topological estimation methods such as the Visual Hull method and consists of a group of points. However, the form of the 3D shape data representing the topology of the subject is not limited to this. Note that the background 3D model can be obtained in advance by an external device. Note that the model generation unit 205 can be configured to be included in the image recording device 102, rather than in the image display device 104. In this case, the 3D model is recorded in the database 103, and the image display device 104 reads the 3D model from the database 103. Note that the following configuration can be adopted: with respect to the foreground 3D model, the coordinates of each of the sphere, person, and subject are calculated and accumulated in the database 103. The coordinates of each subject can be specified as references to be described later. Figure 4C The description of the automatic operation information is used for automatic operation.
[0076] The image generation unit 206 generates virtual viewpoint images from a 3D model based on a controlled virtual camera. Specifically, for each point constituting the 3D model, appropriate pixel values are obtained from the multi-viewpoint image, and shading processing is performed. Additionally, a virtual viewpoint image is generated by arranging a colored 3D model in 3D virtual space, projecting it onto the virtual viewpoint, and rendering it. However, the method for generating virtual viewpoint images is not limited to this; various methods can be used, such as generating virtual viewpoint images through projection transformation of captured images without using a 3D model. Note that the model generation unit 205 and the image generation unit 206 can be configured as devices different from and connected to the image display device 104.
[0077] The tag management unit 204 manages the automatic operation information used by the automatic operation unit 202 as tags. Tags will be described later in the second embodiment, and this embodiment will be described as a configuration without tags.
[0078] (Hardware configuration of the image display device)
[0079] Next, refer to Figure 2B The hardware configuration of the image display device 104 is described below. The image display device 104 includes a CPU (Central Processing Unit) 211, RAM (Random Access Memory) 212, and ROM (Read-Only Memory) 213. In addition, the image display device 104 includes an operation input unit 214, a display unit 215, and an external interface 216.
[0080] The CPU 211 processes data using programs and data stored in RAM 212 and ROM 213. The CPU 211 performs overall motion control of the image display device 104 and executes functions for implementing... Figure 2A The processing of each function is shown. Note that the image display device 104 may include one or more dedicated hardware components different from the CPU 211, and the dedicated hardware may perform at least a portion of the processing performed by the CPU 211. Examples of dedicated hardware include application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and digital signal processors (DSPs).
[0081] ROM 213 stores programs and data. RAM 212 includes a working area for temporarily storing programs and data read from ROM 213. In addition, RAM 212 provides a working area to be used when CPU 211 performs various processes.
[0082] The operation input unit 214 is, for example, a touch panel, and receives information about user operations. For example, it accepts operations related to a virtual camera and timecode. Note that the operation input unit 214 can be connected to an external controller and receive input information from the user related to that operation. The external controller is, for example, a three-axis controller such as a joystick or mouse. Note that the external controller is not limited to these.
[0083] The display unit 215 is a touch panel or screen, and displays the generated virtual viewpoint image. In the case of a touch panel, the operation input unit 214 and the display unit 215 are integrated.
[0084] External interface 216 performs the sending and receiving of information about database 103, for example via LAN. Additionally, information can be transmitted to an external screen via an image output port such as HDMI (registered trademark) or SDI. Furthermore, image data can be transmitted via Ethernet.
[0085] (Virtual camera animation)
[0086] Next, we will refer to Figures 3A to 3HDescribe the actions of the virtual camera 110 (or virtual viewpoint). To facilitate the description of these actions, the position and pose of the virtual camera / view frustum / target point, etc., will be described first.
[0087] Use a coordinate system to specify the virtual camera 110 and its actions. For the coordinate system, you can use... Figure 3A The example shows a typical three-dimensional orthogonal coordinate system formed by the X-axis, Y-axis, and Z-axis.
[0088] A coordinate system is set up and used for the subject. The subject is a venue such as a stadium or studio. Figure 3B As shown, the subjects include the entire stadium field 391, the ball 392 on it, and the athletes 393, etc. Note that the subjects may also include the stands around the field.
[0089] When setting up the coordinate system for the subject, the center of the field 391 is taken as the origin (0, 0, 0). Additionally, the X-axis is the direction of the longer side of the field 391, the Y-axis is the direction of the shorter side of the field 391, and the Z-axis is the direction perpendicular to the field. However, the method of setting up the coordinate system is not limited to this.
[0090] Next, we will use Figure 3C and Figure 3D Describe a virtual camera. A virtual camera is a viewpoint used to render virtual viewpoint images. Figure 3C In the square pyramid shown, the vector extending from the vertex represents the virtual camera's position 301 and its orientation 302. The virtual camera's position is represented by coordinates (x, y, z) in three-dimensional space, and its pose is represented by a unit vector that is scalarized over the components of each axis.
[0091] Assume that the virtual camera pose 302 passes through the center coordinates of the front clamping surface 303 and the rear clamping surface 304. Furthermore, the space 305 clamped between the front clamping surface 303 and the rear clamping surface 304 is referred to as the frustum of the virtual camera, and is located within the area where the image generation unit 206 generates the virtual viewpoint image (or the area where the virtual viewpoint image is projected and displayed, hereinafter referred to as the display area of the virtual viewpoint image). The virtual camera pose 302 is represented as a vector and is referred to as the optical axis vector of the virtual camera.
[0092] Will use Figure 3D Describe the movement and rotation of the virtual camera. The virtual camera moves and rotates within a space represented by three-dimensional coordinates. The movement 306 of the virtual camera is the movement of the virtual camera's position 301 and is represented by the components (x, y, z) of each axis. As shown in Figure 3, the rotation 307 of the virtual camera is represented by yaw about the Z-axis, pitch about the X-axis, and roll about the Y-axis.
[0093] In this way, the virtual camera can freely move and rotate the subject (site) in three-dimensional space, and can generate any area of the subject as a virtual viewpoint image. In other words, by specifying the virtual camera's coordinates X, Y, and Z, as well as the rotation angles (pitch, roll, yaw) along the X, Y, and Z axes, the camera's position and direction can be manipulated.
[0094] Next, refer to Figure 3E Describe the location of the target point of the virtual camera. Figure 3E The virtual camera is shown in pose 302 facing the target surface 308.
[0095] Target plane 308 is, for example Figure 3B The target surface 308 is described in the XY plane (field surface, Z = 0). The target surface 308 is not limited to this and can be a plane parallel to the XY plane and at a natural athlete's height from a human perspective (e.g., Z = 1.7 m). Note that the target surface 308 does not need to be parallel to the XY plane.
[0096] The target point position (coordinates) of the virtual camera represents the virtual camera's pose 302 (vector), target plane 308, and intersection point 309. If the virtual camera's position 301, pose 302, and target plane 308 are determined, the coordinates 309 of the virtual camera's target point can be calculated as unique coordinates. Note that since the calculation of the intersection point of the vector and the plane (the virtual camera's target point 309) can be achieved using known methods, it will not be elaborated further. Note that a virtual viewpoint image can be generated such that the coordinates 309 of the virtual camera's target point are centered within the rendered virtual viewpoint image.
[0097] In this embodiment, the movement of a virtual camera, which can be achieved through a combination of movement and rotation, is referred to as a virtual camera action. As mentioned above, an infinite number of virtual camera movements can be achieved through combinations of movement and rotation, and combinations that are convenient for the user to see and operate can be predefined as virtual camera actions. According to this embodiment, there are several representative virtual camera actions, such as zoom-in, zoom-out, translation of the target point, and horizontal or vertical rotation centered on the target point, which differs from the aforementioned rotation. These actions will be described below.
[0098] Zooming in / out is the action of enlarging or reducing the display of a subject within the display area of a virtual camera. This is achieved by moving the virtual camera's position 301 back and forth along its orientation direction 302 (optical axis). When the virtual camera moves forward along the orientation direction (optical axis) 302, the subject is enlarged because it approaches the subject within the display area without changing direction. Conversely, when the virtual camera moves backward along the orientation direction (optical axis), the subject becomes smaller. Note that zooming in / out is not limited to this method; changes such as the virtual camera's focal length can also be used.
[0099] Translation of the target point is an action in which the target point 309 of the virtual camera moves without changing the pose 302 of the virtual camera. For reference... Figure 3F Describing this action, the translation of the virtual camera's target point 309 is trajectory 320, indicating the movement from the virtual camera's target point 309 to target point 319. Since the virtual camera's position changes during the translation 320 of the target point, while the pose vector 302 remains unchanged, the virtual camera's action ensures that only the target point moves and can continuously see the subject from the same direction. Note that because the pose vector 302 does not change, the translation 320 of the virtual camera's target point 309 and the translation 321 of the virtual camera's position 301 follow the same trajectory (distance).
[0100] Next, refer to Figure 3G Describe the horizontal rotation of the virtual camera centered on the target point. For example... Figure 3G As shown, the horizontal rotation centered on the target point is the motion 322 of the virtual camera rotating along a plane parallel to the target surface 308, with the target point 309 as the center. During this horizontal rotation centered on the target point 309, the pose vector 302 changes while continuously facing the target point 309, and the coordinates of the target point 309 remain unchanged. Because the target point 309 does not move, the user can view the surrounding subjects from various angles without changing the height above the target surface through this horizontal rotation centered on the target point. Since this action does not move the target point, and the distance from the subject (target surface) does not change in height, the user can easily view the target of interest from various angles.
[0101] Note that the trajectory 322 of the virtual camera during rotation is located on a plane parallel to the target surface 308, and the center of rotation is reference number 323. In other words, the virtual camera is determined to move on a circle of radius R such that R^2 = (X1-X2)^2 + (Y1-Y2)^2, where the XY coordinates of the target point are (X1, Y1) and the XY coordinates of the virtual camera are (X2, Y2). Furthermore, by determining the pose to be a pose vector in the XY coordinate direction of the target point from the position of the virtual camera, the subject can be viewed from various angles without moving the target point.
[0102] Next, refer to Figure 3H Describe the vertical rotation of the virtual camera centered on the target point. For example... Figure 3H As shown, the vertical rotation centered on the target point is a rotation 324 of the virtual camera along a vertical plane relative to the target surface 308, centered on the target point 309. During this vertical rotation, the pose vector 302 changes while continuously facing the target point 309, and the coordinates of the target point 309 remain unchanged. Unlike the horizontal rotation 322 described above, the height above the target surface of the virtual camera changes. In this case, the position of the virtual camera can be determined such that the distance between the virtual camera's position and the target point remains unchanged on a plane perpendicular to the XY plane including both the virtual camera's position and the target point. Furthermore, by determining the pose such that the pose vector is the position of the virtual camera from the target point, the subject can be viewed from various angles without moving the target point.
[0103] By combining the vertical rotation 324 and horizontal rotation 322 centered on the target point, the virtual camera can achieve the following action: viewing the subject from various angles in 360 degrees without changing the target point 309. See below. Figures 7A to 7E Describe an example of how to control the operation of a virtual camera using this action.
[0104] Note that virtual camera actions are not limited to these; they can be any actions achieved through a combination of virtual camera movement and rotation. Whether operated automatically or manually, the aforementioned virtual camera actions can be performed in the same way.
[0105] (Action settings processing)
[0106] In this embodiment, operations can be freely allocated to automatic and manual operation. Figures 3A to 3H The virtual camera motion described. This assignment is called motion setup processing.
[0107] Next, refer to Figure 4A and Figure 4BDescribe the motion setting process. Note that the motion setting process is performed in the operation control unit 203 of the image display device 104.
[0108] Figure 4A It is a list in which identifiers are assigned to virtual camera actions, and this list is used in action setup processing. Figure 4A In the process, set each action identifier (ID) Figures 3C to 3H The virtual camera actions described in [the document]. For example, setting a reference for identifier = 2. Figure 3F The translation of the target point is described, and a reference is set for identifier = 3. Figure 3G The description refers to a horizontal rotation centered on the target point. A reference is set for identifier = 4. Figure 3H The description refers to a vertical rotation centered on the target point. Other identifiers can be set in the same way. The operation control unit 203 uses these identifiers to specify the virtual camera actions assigned to automatic or manual operation. Note that in one example, multiple action identifiers can be assigned to the same action. For example, for a reference... Figure 3F The translation of the target point can be set to a translation distance of 10m for identifier = 5, and for the reference... Figure 3F The translation of the target point can be set to a translation distance of 20m for identifier = 6.
[0109] Figure 4B This is an example where actions are set. Figure 4B In this setup, the translation of the virtual camera's target point (identifier = 2) is set as the target for automatic operation. Additionally, the zoom in / out (identifier = 1), translation (identifier = 2), and horizontal rotation (identifier = 3) and vertical rotation (identifier = 4) centered on the target point are set as targets for manual operation.
[0110] The operation control unit 203 of the image display device 104 controls the virtual camera's movement through automatic and manual operations set according to the action. Figure 4B In the example, the virtual camera's target point is panned according to the timecode via automatic operation without changing the position of the target point specified by the automatic operation, and the subject can be viewed at various angles or magnifications by the user's manual operation. See later. Figure 4C , Figure 4D and Figures 7A to 7E This section provides a detailed description of the virtual camera actions and virtual viewpoint image display examples based on this action setup example. Note that the combinations of action settings that can be performed are not limited to... Figure 4BFor example, the action settings can be changed midway through the process. Furthermore, for actions that can be set in automatic or manual operation, the configuration allows for setting each of multiple virtual camera actions, rather than just one.
[0111] The combination of action settings can essentially be an exclusive control between automatic and manual operation. For example, in Figure 4B In the example, the goal of the automatic operation is the translation of the target point of the virtual camera, and during the automatic operation, even if the translation of the target point is performed manually, processing such as cancellation (ignore) can be performed. Additionally, processing such as temporarily disabling exclusive control can be performed.
[0112] Furthermore, for combinations of actions set therein, the image display device 104 can pre-store setting candidates, and the user can accept and set the virtual camera action selection operation assigned to automatic and manual operations.
[0113] (Automatic operation information)
[0114] Regarding references Figure 4A and Figure 4B In the described motion settings process, the virtual camera motion designated as automatic operation, and the information indicating specific content for each timecode, are called automatic operation information. Here, refer to... Figure 4C and Figure 4D Describe information about automated operation. (See reference...) Figure 4A and Figure 4B The process involves determining the identifier (type) of the virtual camera action specified by the automatic operation information through action setting processing. Figure 4C Is with Figure 4B Examples of action settings and corresponding automated operation information.
[0115] for Figure 4C The automatic operation information 401 in the code specifies the position and pose of the target point of the virtual camera as the virtual camera action for each timecode. For example, from... Figure 4C The timecode from the first line to the last line, a continuous timecode (from 2020-02-02 13:51:11.020 to 2020-02-02 13:51:51.059), is called the timecode period. Similarly, the start time (13.51.11.020) is called the start timecode, and the end time (13:51:51.059) is called the end timecode.
[0116] The target point's coordinates are represented by (x, y, z), where z is a constant value (z = z01) and only (x, y) varies. When Figure 4DWhen these coordinates are shown as continuous values in the three-dimensional space 411, it becomes a translation at a constant height above the field surface. Figure 4B The trajectory 412 is defined by identifier 2 in the time code. Meanwhile, regarding the pose of the virtual camera, only an initial value is specified in the corresponding time code period. Thus, for automatic operation using automatic operation information 401, the pose will not change from the initial value (-Y direction) within the specified time code period, and the virtual camera will move in a manner that causes the target point to follow trajectory 412 (421 to 423).
[0117] For example, if trajectory 412 is set to follow the position of a ball (e.g., rugby ball) in a sport (target point position = ball position), the virtual camera will always track the ball and surrounding subjects without needing to change its posture automatically.
[0118] In use such Figure 4D In the case of the automatic operation information shown, respectively in Figures 4E to 4G Examples of virtual viewpoint images are shown when the virtual camera is positioned at reference numbers 421 to 423. As illustrated in this series of figures, the subject entering the virtual camera's field of view changes by altering the timecode and position, but all virtual viewpoint images have the same pose vector, in other words, the same -Y direction. In this embodiment, the automatic operation using this automatic operation information 401 is also referred to as automatic reproduction.
[0119] Note that for coordinates with continuous values such as trajectory 412, e.g. Figure 2A As described in the model generation unit 205, the coordinates of a ball or athlete, which are calculated as foreground objects when generating a 3D model, can be used.
[0120] Additionally, data obtained from devices other than image display device 104 can be used for automated operation information. For example, a position measurement tag can be attached to the clothing of an athlete or a ball included as a subject, and the position information can be used as a target point. Furthermore, the coordinates of the foreground object calculated during the generation of the 3D model and the position information obtained through the position measurement tag can be used as is for the automated operation information, and information calculated separately from this information can also be used. For example, regarding the coordinates and measurements of the foreground object in the 3D model, the z-value changes depending on the athlete's posture and the ball's position, but z can be set to a fixed value, or only the x and y values can be used. The z-value can be set to the height of the field (z = 0) or a height that is natural and easy for the viewer's eyes (e.g., z = 1.7 m). Note that the device used to obtain the automated operation information is not limited to, for example, a position measurement tag; information received from another image display device different from the image display device 104 that is being displayed can be used. For example, a worker operating the automated operation information can input the automated operation information by operating another image display device different from the image display device 104.
[0121] Note that the timecode period that can be assigned to automatic operation information can be the entire game or a part of it in a sport.
[0122] Note that examples of automated operation information are not limited to those mentioned above. If the item in the automated operation information relates to virtual camera movements, then the information is not limited to that. Other examples of automated operation information will be described later.
[0123] (Manual operation screen)
[0124] Will use Figure 5A and Figure 5B This describes the manual operation of the virtual camera according to this embodiment. Figure 4A and Figure 4B The motion setup process shown determines the virtual camera action specified by manual operation. Here, the manual operation methods for the virtual camera actions assigned in the motion setup process and the manual operation methods for the timecode will be described.
[0125] Figure 5A This is a view describing the configuration of the manual operation screen 501 of the operation input unit 214 of the image display device 104.
[0126] exist Figure 5A The operation screen 501 is mainly configured by the virtual camera operation area 502 and the timecode operation area 503. The virtual camera operation area 502 is a graphical user interface (GUI) for accepting manual operations of the user on the virtual camera, and the timecode operation area 503 is a graphical user interface (GUI) for accepting manual operations of the user on the timecode.
[0127] First, the virtual camera operation area 502 will be described. The virtual camera operation area 502 performs operations based on the type of operation received. Figure 4A and Figure 4B The virtual camera action is set in the action settings process.
[0128] exist Figure 5A In the example, because a touch panel is used, the types of operations are touch operations such as taps and swipes. However, if a mouse is used to input user actions, a click can replace a tap, and a drag can replace a swipe. Each of these touch operations is assigned as such... Figure 4A The virtual camera actions shown include movements such as translation of the target point and horizontal / vertical rotation centered on the target point. The table used to determine the assignment is shown in Figure 5, which is referred to as the manual operation settings.
[0129] Will use Figure 5B Describe the manual operation settings. Figure 5B In the project, there are identifiers (IDs) for touch operations, touch counts, touch areas, and virtual camera actions.
[0130] The touch operation section lists possible touch operations within the virtual camera operation area 502. Examples include clicking and swiping. The touch count section defines the number of fingers required for the touch operation. The touch area section specifies the area to be processed for the touch operation. For example, it specifies the virtual camera operation area (overall) and the virtual camera operation area (right edge). Based on the content of these three sections—touch operation, touch count, and touch area—it becomes possible to distinguish the content of the manual operation received from the user. Then, when each manual operation is received, one of the virtual camera actions is executed based on the identifier (ID) of the virtual camera action assigned to each manual operation. (See reference...) Figure 4A The aforementioned identification of virtual camera actions (IDs) Figures 3C to 3H The types of virtual camera actions described in the text.
[0131] exist Figure 5B In the allocation example, for instance, by specifying the second line (No.2), when a horizontal swipe operation is accepted on the virtual camera operation area 502 using one finger, a horizontal rotation centered on the target point is performed (motion identifier = 3). Additionally, by specifying the third line (No.3), when a vertical swipe operation is accepted on the virtual camera operation area 502 using one finger, a vertical rotation centered on the target point is performed (motion identifier = 4).
[0132] Note that regarding the type of touch operation for manual operation of the virtual camera, gesture operations based on finger counting can be used, and the relatively simple gesture operations shown later in the third embodiment can also be used. Both can be... Figure 5B The manual operation settings can be used to specify the operation. Note that the operation that can be specified through the manual operation settings is not limited to touch operation, and can also be specified when using devices other than touch panels.
[0133] Next, the timecode operation area 503 will be described. The timecode operation area 503 is configured by components 512 to 515 for operating the timecode. The main slider 512 can operate on all timecodes of the captured data. Any timecode in the main slider 512 can be specified by using the position of the selection knob 522 via dragging operations, etc.
[0134] The sub-slider 513 magnifies and displays some timecodes, and can be manipulated with a finer granularity than the main slider 512. The timecode can be specified by using dragging operations or selecting the position of the knob 523. The main slider 512 and the sub-slider 513 have the same length on the screen, but the range of timecodes they can select differ. For example, while the main slider 512 can select from 3 hours as the duration of a match, the sub-slider 513 can select from 30 seconds as a part of the duration of a match. In other words, each slider has a different scale, and the sub-slider 513 can specify the timecode with a finer granularity, such as in frames. In one example, the sub-slider 513 selects from 15 seconds before to 15 seconds after the time indicated by the knob 522.
[0135] Note that the time code specified using knob 522 of main slider 512 and knob 523 of sub-slider 513 can be displayed as a numerical value in the form of "day:hour:minute:second.frame number". Note that sub-slider 513 is not always displayed. For example, it can be displayed after a display command is received, or when a specific operation such as pause is indicated. Furthermore, the time code period selectable using sub-slider 513 is variable. When a specific operation such as pause is received, the interval from 15 seconds before to 15 seconds after the time on knob 523 can be displayed when a pause command is received.
[0136] The slider 514, which specifies the playback speed, allows you to choose between normal speed playback and slow speed playback. The timecode increment interval is controlled by the playback speed selected using knob 524. (Reference) Figure 6 The flowchart describes an example of controlling the increment interval of the timecode.
[0137] The Cancel button 515 can be used to cancel any timecode-related operations, and can also be used to remove pauses and return to normal playback. Note that if the button is a manual timecode-related operation, the button is not limited to cancel.
[0138] Note that the manual operation screen can include areas other than the virtual camera operation area 502 and the timecode operation area 503. For example, match information can be displayed in area 511 at the top of the screen. This includes the host location / date and time, match cards, and score information. Note that match information is not limited to this.
[0139] Furthermore, exceptional operations can be assigned to the upper area 511 of the screen. For example, when double-clicking is performed in the upper area 511, the position and pose of the virtual camera can be manipulated to move it to a position from above, allowing a view of the entire subject. Inexperienced users may find it difficult to manually operate the virtual camera and may be unsure of its location. In such cases, an operation can be assigned to return to a viewpoint that provides a clearer understanding of the position and pose. The viewpoint is as follows: Figure 3B The image shown is a virtual viewpoint image of the subject viewed from the Z-axis.
[0140] Note that the configuration is not limited to these, as long as the virtual camera or timecode can be operated and the virtual camera operation area 502 and the timecode operation area 503 do not need to be separated. For example, a pause operation can be performed by double-clicking on the virtual camera operation area 502, which is a timecode operation. Note that although the case where the operation input unit 214 is a tablet computer 500 has been described, the operation and display devices are not limited to this. For example, the operation input unit 214 can be configured to perform a 10-second fast forward when the right half of the display area 502 is double-clicked, and to perform a 10-second rewind when the left half is double-clicked.
[0141] (Operation control processing)
[0142] Next, use Figure 6 The flowchart describes the control processing of the virtual camera, in which automatic and manual operations are meaningfully combined according to this embodiment.
[0143] The aforementioned motion setting process determines the virtual camera motion set for each of the automatic and manual operations. Furthermore, according to... Figure 4A The automatic operation information described in the document is used to execute automatic operations, based on... Figure 5BThe manual settings described herein are executed manually. Furthermore, this flowchart describes a virtual camera control process that meaningfully combines these operations. Note that the flowchart shown in Figure 3 is a flowchart. This is achieved by the CPU 211 using RAM 212 as its workspace and executing a program stored in ROM 213. Figure 6 The flowchart shown.
[0144] In step S602, the operation control unit 203 performs processing related to the timecode and playback speed. Subsequently, in steps S604 to S608, processing is performed to meaningfully combine automatic and manual operations. Furthermore, the operation control unit 203 performs the cyclic processing of steps S602 to S611 for each frame. For example, if the frame rate of the output virtual viewpoint image is 60 FPS, then the processing of one cycle (one frame) of steps S602 to S611 is performed at intervals of approximately 16.6 [ms]. Note that the interval of one cycle can be achieved by setting the update rate (refresh rate) in the image display such as a touch panel to 60 FPS and performing synchronization processing therewith in the image display device.
[0145] In step S602, the operation control unit 203 increments the time code. According to this embodiment, it can be achieved by... Figure 4C The timecode is specified using "day:hour:minute:second.frame number" and incremented in units of frames. In other words, since the loop processing of steps S602 to S611 is performed for each frame as described above, the frame number in the timecode is incremented in each loop processing.
[0146] Note that the increment of the timecode can be based on... Figure 5A The playback speed slider 514 described herein is selected to change the increment interval. For example, when a playback speed of 1 / 2 is specified, the frame increment can be one frame every two cycles of steps S602 to S611.
[0147] Next, the operation control unit 203 causes the processing to proceed to step S603, and the model generation unit 205 obtains a three-dimensional model with incremental or specified timecodes. For example... Figure 2A As shown, the model generation unit 205 generates a 3D model of the subject from multi-viewpoint images.
[0148] Next, in step S604, the operation control unit 203 determines whether there is automatic operation information for the automatic operation unit 202 in the incrementing or specified time code. If there is automatic operation information in the time code ("Yes" in step S604), the operation control unit 203 causes the process to proceed to step S605; otherwise, if there is no automatic operation information ("No" in step S604), the operation control unit 203 causes the process to proceed to step S606.
[0149] In step S605, the operation control unit 203 performs automatic operation of the virtual camera by using the automatic operation information in the corresponding timecode. (Note: The last part, "due to reference," appears to be an error and doesn't translate directly. It likely refers to a separate process.) Figures 4A to 4C The automatic operation of the virtual camera has been described, so it will not be repeated here.
[0150] In step S606, the operation control unit 203 receives a manual operation from the user and switches the processing according to the received operation. If a manual operation is received regarding timecode ("Timecode Operation" in step S606), the operation control unit 203 proceeds to step S609. If a manual operation is received regarding the virtual camera ("Virtual Camera Operation" in step S606), the operation control unit 203 proceeds to step S607. If no manual operation is received (step S606 is "None"), the operation control unit 203 proceeds to step S609.
[0151] In step S607, the operation control unit 203 determines whether the virtual camera action specified by manual operation (step S606) interferes with the virtual camera action specified by automatic operation (step S605). The comparison of actions can be performed using... Figure 4A The virtual camera motion identifiers shown determine whether automatic and manual operations interfere with each other when they have the same identifier. Note that this is not limited to comparing identifiers, and interference can also be determined by comparing variables such as the virtual camera's position or pose.
[0152] If it is determined that the action of the virtual camera specified by manual operation interferes with the action specified by automatic operation ("Yes" in step S607), the operation control unit 203 causes the process to proceed to step S609, and the accepted manual operation is cancelled. In other words, when automatic and manual operations interfere, automatic operation takes priority. Note that manual operation is used to operate parameters that will not be changed by automatic operation. For example, when the translation of the target point is performed by automatic operation, the position of the target point changes, but the pose of the virtual camera remains unchanged. Therefore, the horizontal rotation, vertical rotation, etc., of the virtual camera will not change the position of the target point, but will change the pose of the virtual camera, and will not interfere with the translation of the target point, so it can be performed by manual operation. If it is determined that the action of the specified virtual camera does not interfere with the action specified by automatic operation ("No" in step S607), the operation control unit 203 causes the process to proceed to step S608, and performs control processing of the virtual camera that combines automatic and manual operations.
[0153] For example, if the automatic operation action specified in step S605 is a translation of the target point to be viewed (identifier = 2), and the manual operation action specified in step S606 is also a translation of the target point to be viewed (identifier = 2), then the manual translation of the target point is cancelled. On the other hand, for example, if the automatic operation action specified in step S605 is a translation of the target point (identifier = 2), and the manual operation action specified in step S606 is a vertical / horizontal rotation around the center of the target point (identifier = 3, 4), then they are determined to be different actions. In this case, as will be explained later... Figures 7A to 7E In step S608, each of the actions that are combined is performed.
[0154] In step S608, the operation control unit 203 continues to perform manual operation of the virtual camera. Because in Figure 5A and Figure 5B The manual operation of the virtual camera is described in the previous section, so it will not be repeated here. Next, the operation control unit 203 causes the processing to proceed to step S609, and, in the case of performing imaging based on the position and pose of the virtual camera operated by at least one of automatic and manual operations, a virtual viewpoint image is generated and rendered. Figure 2A The rendering of the virtual viewpoint image has been described, so it will not be repeated here.
[0155] In step S610, the operation control unit 203 updates the timecode to be displayed to the timecode manually specified by the user. Because in Figure 5A and Figure 5B The manual steps for specifying timecodes are described in the document, so they will not be repeated here.
[0156] In step S611, the operation control unit 203 determines whether the display of each frame has been completed; in other words, whether the end timecode has been reached. If the end timecode has been reached (yes in step S611), the operation control unit 203 terminates. Figure 6 The process is as shown, and if the end timecode has not yet been reached (No in step S611), the operation control unit 203 returns the process to step S602. This is achieved by using the following... Figures 7A to 7E This document describes an example of controlling a virtual camera using the flowchart described above, as well as an example of generating virtual viewpoint images.
[0157] (Examples of virtual camera control and virtual viewpoint image display)
[0158] Reference Figures 7A to 7E This describes a control example of a virtual camera that combines automatic and manual operation, and a display example of a virtual viewpoint image according to this embodiment.
[0159] Here, an example of virtual camera control is described. Figure 4C and Figure 4D The automatic operation information described in the document is used in automatic operation and in... Figure 5B The manual settings described herein are used in cases where manual operation is required.
[0160] for Figure 4C and Figure 4D The automated operation manipulates the virtual camera to perform translation (identifier = 2), so that the coordinates of the virtual camera's target point follow trajectory 412 according to the passage of timecode without changing its pose. For Figure 5B Manual operation, where the user operates the virtual camera, allows for horizontal / vertical rotation centered on the target point to be performed through at least one of sliding and dragging actions. Figure 3G and Figure 3H (e.g., identifiers = 3, 4).
[0161] For example, if Figure 4D If the focus 412 can follow the ball in a sport, then the automatic operation in step S605 is a translation action, ensuring that the ball is always captured at the target point of the virtual camera. In addition to this action, by performing a manual operation of horizontal / vertical rotation centered on the target point, the user can unconsciously track the ball while viewing athletes around it from various angles.
[0162] By using Figures 7A to 7D Examples describing this series of actions. Figures 7A to 7E In the example, suppose the sport is rugby, and the scene is filmed as follows: a player is about to make an offload pass (a pass just before falling). In this case, the subjects are the rugby ball and at least one player on the field.
[0163] Figure 7A It is a three-dimensional view overlooking the sports field, and it is the case where the target point of the virtual camera is translated (identifier = 2) to coordinates 701 to 703 by automatic operation (track 412).
[0164] First, let's describe the case where the virtual camera's target point is at coordinates 701. In this case, the virtual camera is positioned at coordinates 711, and its pose is primarily in the -X direction. For example... Figure 7B As shown, the virtual viewpoint image at this time is a view of the game along the long side of the sports field in the -X direction.
[0165] Next, we describe the case where the virtual camera's target point is automatically translated to coordinate 702. Here, we assume that while the target point is automatically translated from coordinate 701 (identifier = 2) to coordinate 702, the user performs a manual operation of vertical rotation centered on the target point (identifier = 4). In this case, the virtual camera's position is translated to coordinate 712, and the virtual camera's pose is generally rotated vertically along the -Z direction. Figure 7C As shown, the virtual viewpoint image at this time is an overhead view of the subject from above the sports field.
[0166] Furthermore, the case where the virtual camera's target point is translated to coordinate 703 is described. Here, it is assumed that when the target point is translated from coordinate 702 (identifier = 2) to coordinate 703, the user performs a manual operation of horizontal rotation centered on the target point (identifier = 3). In this case, the virtual camera's position is translated to coordinate 713, and the virtual camera's pose is rotated vertically approximately in the -Y direction. Figure 7D As shown, the virtual viewpoint image at this time is a view of the subject from outside the sports field along the -Y direction.
[0167] As described above, by meaningfully combining automatic and manual operations, in rapidly evolving scenarios such as rugby passing, continuous automatic filming of the vicinity of the ball can be performed within the virtual camera's field of view, while manual operation allows viewing the scene from various angles. Specifically, by using the ball's coordinates obtained from a 3D model or position measurement tags, continuous filming of the vicinity of the ball can be performed within the virtual camera's field of view without imposing an operational burden on the user. Furthermore, as a manual operation, the user can view the game of interest from various angles by simply applying a single drag operation (horizontal / vertical rotation, where the target point is made the gaze point by dragging a finger).
[0168] Note that in the section performing automatic operations (trajectory 412), the virtual camera's pose remains the same as the pose at the last manual operation performed, unless manual operation is performed. For example, when the target point's position is translated from coordinate 703 (identifier = 2) to coordinate 704 via automatic operation without manual operation, the virtual camera's pose remains the same at coordinates 713 and 714 (the same pose vector). Accordingly, as... Figure 7E As shown, the virtual camera's pose is in the -Y direction, while the following text... Figure 7D During the passage of time without manual operation, the image display device 104 displays a scene in the same posture in which the athlete successfully performs a pass and sprints toward a try to score with the ball.
[0169] Note that you can... Figure 6 and Figures 7A to 7EThe operation control adds processing to temporarily invalidate or enable settings through at least one of manual and automatic operations. For example, upon receiving a pause command for reproducing a virtual viewpoint image, the action setting processing configured through manual and automatic operations can be reset to set the default virtual camera position and coordinates. For the default virtual camera position, for example, for a pose above the field of view that fits the entire field of view, the virtual camera position and coordinates can be changed in the -Z direction.
[0170] In another example, upon receiving an instruction to pause the playback of the virtual viewpoint image, the content set manually can be invalidated, and upon releasing the pause and resuming playback, the content set manually can be reflected again.
[0171] Furthermore, when the translation of the target point (identifier = 2) is set to automatic operation during automatic operation, the manual operation is cancelled in step S607 as a result of the manual operation for the same action (identifier = 2). During pause, the target point can be translated manually regardless of the automatic operation information.
[0172] For example, consider the following scenario: during automatic playback, the session is paused when the virtual camera's target point is at coordinates 701. The virtual viewpoint image at this time would look like... Figure 7B As shown; however, the appearance of a trackling player from the opposing team who triggered the pass needs to be confirmed, but the outside player is not in the field of view. In this situation, after performing a manual panning operation in the downward direction of the screen, the user can pause and pan the target point to coordinates 705, etc.
[0173] Additionally, when the pause is released and normal playback resumes, the action settings can be made active, automatically performing the translation of the target point (identifier = 2), and manually canceling the translation of the target point (identifier = 2). Note that at this time, the amount of movement already performed manually (the distance moved from coordinate 701 to coordinate 705) can also be canceled, and processing to return the target point to coordinate 701 specified by the automatic operation can be performed.
[0174] As described above, by meaningfully combining automatic and manual operation and controlling the virtual camera, it becomes possible to continuously capture the desired location of the subject within the virtual camera's viewpoint, even in scenes where the subject moves rapidly and complexly, and to browse from various angles.
[0175] For example, in sports such as rugby, even in scenarios where a game suddenly develops and leads to a score over a wide field, it is possible to continuously film the vicinity of the ball within the field of view of a virtual camera through automatic operation, and to view the scene from various angles when manual operation is applied.
[0176] <Second Embodiment>
[0177] This embodiment describes the label generation process and label reproduction process in the image display device 104. (Including...) Figure 1A Virtual viewpoint image generation system and Figure 2A The image display device 104, including the same configuration, processing and functions, uses the same reference numerals as in the first embodiment, and will not be described again.
[0178] The tags in this embodiment enable users to select a target scene when scoring, for example, in a game, and set automated operations within the selected scene. Users can also perform additional manual operations during automatic playback based on the tags, and can perform virtual camera control processing, in which automatic and manual operations are combined using the tags.
[0179] These tags are used in sports commentary and other similar applications, which is a typical use case for virtual viewpoint images. In this usage scenario, it is necessary to be able to browse from various positions and angles and provide commentary on multiple important scenes, such as scoring moments. In addition to automatically operating on the position of the ball or athlete as the target point for the virtual camera, as shown in the example of the first embodiment, this embodiment can also generate tags to make any position other than the subject, such as the ball or athlete, the target point, based on the intention of the commentary, etc.
[0180] In this embodiment, the label is described under the assumption that the automatic operations used to set the position and pose of the virtual camera are associated with timecode. However, in one example, a configuration may be adopted in which a timecode period (time cycle) including a start timecode and an end timecode is associated with multiple automatic operations within that timecode period.
[0181] Note that in Figure 1A In the virtual viewpoint image generation system, the following configuration can be used: data for a target time period, such as scores, is copied from the database 103 to the image display device 104 and used, with tags added to the copied timecode period. This facilitates easy access to only the copied data, and the image display device 104 can be used for narration using virtual viewpoint images of important scenes in sports programs, etc.
[0182] In this embodiment, the label is on the image display device 104 ( Figure 2AThe label management unit 204 manages the label generation and label reproduction processes, while the operation control unit 203 performs these processes. Each process is implemented based on the following: as described in the first embodiment. Figure 4A Automatic operation and Figure 5B Manual operation, and Figure 6 The system combines automatic and manual operations with virtual camera control processing.
[0183] (Label generation and processing)
[0184] Next, refer to Figure 8A A method for generating tags in an image display device 104 is described.
[0185] Figure 8A Is with Figure 5A The image display device 104 described herein has the same configuration as the manual operation screen, and the touch panel 501 is roughly divided into a virtual camera operation area 502 and a timecode operation area 503.
[0186] In the user actions used for tag generation, the user first selects a timecode in the timecode operation area 503, and then specifies a virtual camera action for that timecode in the virtual camera operation area 502. By repeating these two steps, automatic actions can be created within any timecode period.
[0187] Even during tag generation, the operation used to specify the timecode in the timecode operation area 503 is also consistent with... Figure 5A The content described herein is the same. In other words, the desired timecode for label generation can be specified by dragging the knob on the main slider or by dragging the knob on the sub-slider.
[0188] Note that regarding the specification of the time code, a coarse time code can be specified in the main slider 512 to execute the pause command, and then a detailed time code can be specified using the sub-slider 513, which displays a range of tens of seconds before and after that time. Note that the sub-slider 513 can be displayed when the indicator label is generated.
[0189] When the user selects a timecode, a virtual viewpoint image (overhead view) is displayed at that timecode. Figure 8A On the virtual camera operation area 502, a virtual viewpoint image overlooking the subject is displayed. This is because it is easier to identify the game situation when viewing an overhead image that displays a wide range of the subject, and it is easier to specify target points.
[0190] Next, while viewing the overhead image (virtual viewpoint image) of the timecode being displayed, a target point of the virtual camera for the timecode is specified via a tap operation 812. The image display device 104 converts the tap location into coordinates on the target surface of the virtual camera in three-dimensional space and records the coordinates as the coordinates of the target point at that timecode. Note that in this embodiment, the XY coordinates are assumed to be the tap location, and the Z coordinate is described as Z = 1.7m. However, the Z coordinate can be obtained based on a tap location similar to the XY coordinates. The method of converting from the tap location to the target is well known and will not be described further.
[0191] By repeating the two steps above ( Figure 8A (821 and 822 in the code) can specify the target point of the virtual camera in chronological order within any time code period.
[0192] Note that although the timecodes specified by the above operations are sometimes discontinuous and have intervals, coordinates can be generated by interpolation using manually specified coordinates during the intervals. The method of interpolating between multiple points using their respective curve functions (spline curves, etc.) is well-known and therefore omitted.
[0193] The tags generated by the above steps have the same characteristics as... Figure 4A The automatic operation information shown is in the same format. Therefore, the tags generated in this embodiment can be used in the virtual camera control process described in the first embodiment, in which automatic and manual operations are meaningfully combined.
[0194] Note that the virtual camera action to be specified by the label is not limited to the coordinates of the target point, and can be any virtual camera action described in the first embodiment.
[0195] Using the above method, multiple different tags can be generated for multiple scoring scenarios in a competition. Additionally, multiple tags with different automated operation information can be generated for timecodes within the same time interval.
[0196] Note that the following configuration can be adopted: a button can be prepared for switching between the normal virtual camera operation screen described in the first embodiment and the tag generation operation screen in this embodiment. Note that, although not shown, it can be configured such that after switching to the tag generation operation screen, processing to move the virtual camera to a position and pose overlooking the subject can be performed.
[0197] Using the tag generation process described above, automated actions can be easily generated as tags at any time code period when a target scenario, such as a score, occurs during a sports competition. Next, the tag reproduction process will be described, through which users can easily and automatically reproduce the scenario by selecting tags generated in this way.
[0198] (Label reproduction processing)
[0199] Next, we will use Figure 9 The flowchart and Figure 8B and Figure 8C The operation screen describes the display of labels in the image display device 104.
[0200] Figure 9 This is a flowchart illustrating an example of the control processing of a virtual camera performed by the image display device 104 during label reproduction. Figure 9 In, similar to Figure 6 The operation control unit 203 repeatedly executes the processing of S902 to S930 frame by frame. Each time, the frame number in the timecode is incremented by 1, and then consecutive frames are processed. Note that the description of... Figure 6 The flowchart for virtual camera control processing in the process follows the same steps.
[0201] In step S902, the operation control unit 203 performs the same process as in S603 and obtains a three-dimensional model of the current timecode.
[0202] In step S903, the operation control unit 203 determines the display mode. The display mode is for... Figure 5A The system determines the display mode of the operation screen and identifies either the normal playback mode or the label playback mode. When the display mode is determined to be normal playback, the process proceeds to step S906; otherwise, it proceeds to step S904.
[0203] Note that the display modes to be determined are not limited to these. For example, you can... Figure 8A The label generation shown is added to the target determined by the pattern.
[0204] Next, the image display device 104 initiates the process at step S904, and since the display mode is label reproduction, the operation control unit 203 reads the automatic operation information specified by the time code from the selected label. Figure 4C The automatic operation information shown is as described above. Next, the image display device 104 initiates the process to step S905, and the operation control unit 203 performs the same process as in step S605, and automatically operates the virtual camera.
[0205] Next, the image display device 104 initiates the process to step S906, and the operation control unit 203 accepts a manual operation from the user and switches the process according to the accepted operation. If a manual operation regarding timecode is accepted from the user ("Timecode Operation" in step S906), the image display device 104 initiates the process to step S912. If a manual operation related to the virtual camera is accepted from the user ("Virtual Camera Operation" in step S906), the image display device 104 initiates the process to step S922. If no manual operation is accepted from the user ("None" in step S906), the image display device 104 initiates the process to step S907. If a tag selection operation is accepted from the user ("Tag Selection Operation" in step S906), the image display device 104 initiates the process to step S922. Figure 8B and Figure 8C Describe the manual operation screen at this time.
[0206] Note that the manual operations determined in step S906 are not limited to these, and may include... Figure 8A The label generation operation shown is taken as the target.
[0207] exist Figure 8B In the process, extract the timecode operation area 503 from the manual operation screen 801. During operation... Figure 8B On the main slider 512 for the time code, multiple labels (831 to 833) are displayed in a distinguishable manner. Regarding the display position of the labels (831 to 833), the start time code of the label is displayed at the corresponding position on the main slider. Note that for the displayed labels, a time code can be used as follows: Figure 8A The label shown is generated in the image generation device 104, and can also read labels generated on other devices.
[0208] Here, the user can select any of the multiple labels (831 to 833) displayed on the main slider 512 by tapping operation 834. When a label is selected, as shown in the dialog bubble 835, the automatic operation information set in the label can be listed. In the dialog bubble 835, two automatic operations are set in timecodes, and two label names (Auto1 and Auto2) corresponding to them are displayed. Note that the number of labels that can be set to the same timecode is not limited to two. When the user selects any label 831 to 833 by tapping operation 834, the image display device 104 causes the processing to proceed to step S922.
[0209] In step S922, the operation control unit 203 sets the label reproduction to display mode and causes the process to proceed to step S923.
[0210] In step S923, the operation control unit 203 reads the start timecode specified by the tag selected in step S906 and proceeds to step S912. Then, in step S912, the operation control unit 203 updates the timecode to the specified timecode and returns the process to step S902. Furthermore, after S904 and S905, the image display device 104 uses the automatic operation information specified by the selected tag to perform automatic operation of the virtual camera.
[0211] Here, in Figure 8C The image shows the screen when the label is reproduced (automatic reproduction after label selection). Figure 8C In this context, a sub-slider 843 is displayed on the timecode operation area 503. Note that the scale of the sub-slider 843 can be the length of the timecode period specified by the label. Furthermore, the sub-slider 843 can be dynamically displayed when a label is indicated for reproduction. Moreover, during label reproduction, the label being reproduced can be replaced with a display 842 that is distinguishable from other labels.
[0212] exist Figure 8C In the virtual camera operation area 502, a virtual viewpoint image generated by the virtual camera specified by the automatic operation information included in the selected label is displayed. When the label is reproduced, manual operation can be added, and in S907 and S908, the image display device 104 performs virtual camera control processing, in which, similar to that described in the first embodiment, automatic operation and manual operation are combined.
[0213] In step S907, when the operation control unit 203 receives a manual operation related to the virtual camera, the operation control unit 203 confirms the current display mode. If the display mode is the label playback mode ("normal playback" in step S907), the operation control unit 203 causes the process to proceed to step S909, and if the display mode is the normal playback mode ("label playback" in step S907), the operation control unit 203 causes the process to proceed to step S908.
[0214] In step S908, the operation control unit 203 compares the automatic operation in the selected tag with the manual operation of the virtual camera accepted in step S906 and determines whether they interfere with each other. The method for comparing automatic and manual operations is similar to that in step S607, and therefore will not be described again.
[0215] In step S909, similar to step S608, the operation control unit 203 performs manual operation of the virtual camera. In step S910, similar to step S609, the operation control unit 203 renders the virtual viewpoint image. In step S911, the operation control unit 203 increments the time code. The incrementing of the time code is similar to that in step S602, and therefore will not be described again.
[0216] In step S921, the operation control unit 203 determines whether the time code incremented in step S911 during label reproduction is the end time code specified for the label. If the incremented time code is not the end time code specified in the label ("No" in step S921), the operation control unit 203 causes the process to proceed to step S930 and performs a determination similar to that in step S611. If the incremented time code is the end time code specified in the label ("Yes" in step S921), the operation control unit causes the process to return to step S923, reads the start time code of the selected label, and performs repeated reproduction to repeat the label reproduction. Note that in this case, the manual operation set in the previous reproduction can be reflected as an automatic operation.
[0217] Note that to disable repeated reproduction during label reproduction, the user can press [button / button]. Figure 8C The cancel button shown is 844.
[0218] Note that when the repeated label reappears, the posture of the last manual operation can be retained during the repetition. This can be configured to allow for different operation postures. Figure 4C and Figure 4D When the automated operation information is used as a tag, and repeated reproduction is performed within the timecode period (track 412, start point 700, and end point 710), the pose at end point 710 can be maintained even when returning to start point 700. Alternatively, the following configuration can be adopted: when returning to start point 700, the pose returns to the initial value specified in the automated operation information.
[0219] As described above, by means of the tag generation method according to this embodiment, automated operation information about a predetermined range of time codes can be easily generated in the image display device 104 in a short time.
[0220] Furthermore, thanks to the label-based reproduction method, users can automatically reproduce each scene with a simple action of selecting one desired label from multiple options. Additionally, even if the subject's behavior varies and is complex depending on the scene, the system can continuously capture images of the desired location within the virtual camera's viewpoint in each scene through automatic operation. Moreover, users can apply manual operations to automatic operations to display the subject at the user's desired angle. Furthermore, since users can add labels to automatic operations, multiple virtual camera operations can be combined according to user preferences.
[0221] For example, for multiple scoring scenarios in sports broadcasting, timecode operations can be used to specify timecodes, and the specified coordinates of the target point of the virtual camera can be added as tags, easily generating tags for multiple scenarios. By using these tags in post-match commentary programs, commentators can easily display multiple scoring scenarios from various angles and use them for commentary. For example, as in the first embodiment... Figures 7B to 7D As shown, by displaying the scene to be focused on from various angles, the points to be focused on can be explained in detail.
[0222] Furthermore, tags are not limited to sports commentary. For example, tags can be generated for scenes that experts or well-known athletes in a target sport should focus on, and other users can use these tags to reproduce the content. With just a few simple steps, users can also view scenes from various angles using tags set by highly specialized individuals.
[0223] Alternatively, the generated tags can be sent to other image display devices and the scene can be automatically reproduced by using tags on other image display devices.
[0224] <Third Implementation Method>
[0225] In this embodiment, an example of manual operation of the virtual camera in the image display device 104 will be described. Compared to methods that switch virtual camera actions based on changes in the number of fingers in a touch operation, since the virtual camera action can be switched with a single finger, users can more easily perform manual operations.
[0226] Configurations, processes, and functions similar to those in the first and second embodiments will not be described again. These have already been described in the first embodiment. Figure 1A Virtual viewpoint image generation system and Figure 2A The functional blocks of the image display device 104 are not described in detail here, and only the manual operation will be described.
[0227] Reference Figure 10 This section will describe the manual operations related to the virtual camera in this embodiment. This is in contrast to the operation described in the first embodiment. Figure 5A In similar cases, Figure 10The screen configuration includes a virtual camera operation area 502 and a timecode operation area 503 displayed on the touch panel 1001. The components of the timecode operation area 503 are as follows: Figure 5A The descriptions are similar, so I will not repeat them here.
[0228] In this embodiment, touch operations on the virtual camera operation area 1002 for two virtual camera actions will be described when the touch panel is used as the operation input unit 214 of the image display device 104. One action is Figure 3F The translation of the target point of the virtual camera shown is another action, which is the zooming in / out of the virtual camera (not shown).
[0229] In this embodiment, the touch operation for performing translation of the target point of the virtual camera is a tap operation 1021 (or a double tap operation) at any location within the operation area. The image display device 104 converts the tap position into coordinates on the target surface of the virtual camera in three-dimensional space and translates the virtual camera so that the coordinates become the target point of the virtual camera. For the user, this operation moves the tap position to the center of the screen, and by simply repeating the tap operation, the position on the subject that the user wants to focus on can continue to be displayed. Note that in one example, the virtual camera can be translated at a predetermined speed from the coordinates 1020 of the current virtual camera's target point to the point specified by the tap operation 1021. Note that the movement speed from coordinates 1020 to the point specified by the tap operation 1021 can be a constant speed. In this case, the movement speed can be determined to complete the movement within a predetermined time period, such as one second. Alternatively, the movement speed can be configured to increase proportionally to the distance to the point specified by the tap operation 1021 and slow down as it gets closer to the point specified by the tap operation 1021.
[0230] Note that when performing translation using a two-finger dragging operation (slide operation), it is difficult to adjust the amount of movement when the subject is moving rapidly. In contrast, in this embodiment, only a direct tap on the desired destination is required.
[0231] Touch operations for zooming in or out of the virtual camera are drag operations 1022 within a predetermined area on the right edge of the virtual camera operation area 1002. For example, a pinch-out is performed by dragging (sliding) the predetermined area at the end (edge) of the virtual camera operation area 1002 in an upward direction, and a pinch-in is performed by dragging (sliding) in a downward direction. Note that... Figure 10 The example shows the right edge of the screen of the virtual camera operation area 1002, but it can be the edge of the screen in another direction, and the dragging direction is not limited to up and down.
[0232] Compared to zooming in / out by pinching with two fingers, this only requires dragging with one finger along the edge of the screen. For example, dragging upwards at the edge of the screen zooms in, and dragging downwards at the edge of the screen zooms out.
[0233] Note that the above operation can be set to Figure 5A and Figure 5B The manual settings described herein can be used for virtual camera control processing, in which automatic and manual operations are combined, as described in the first embodiment.
[0234] As described above, through manual operation according to this embodiment, the user does not need to switch fingers to operate the virtual camera in small increments, and can intuitively operate the image display device 104 with one finger. Furthermore, in panning operations, since panning to a predetermined position can be achieved by tapping the target point of the specified moving destination, panning can be intuitively indicated compared to operations that require adjusting the amount of operation such as dragging and sliding.
[0235] <Fourth Implementation Method>
[0236] In this embodiment, an example of the automatic operation of the virtual camera in the image display device 104 will be described.
[0237] As an example of automated operation information, using... Figure 4C For the same event, for a specific timecode interval, the coordinates and pose of the virtual camera's target point will be fixed. For example, a match may be interrupted due to athlete injury or fouls, and such time gaps can be tedious for viewers. Upon detecting such an interruption, the following configuration can be implemented: automatically maneuver the virtual camera to a position and pose where the entire arena can be viewed, and automatically advance the timecode at twice the speed. Furthermore, it's possible to specify skipping all interruptions. In this case, the interruption can be input by staff or specified by the user using tags. By performing such automation, users can fast-forward or skip tedious time gaps when the match is interrupted.
[0238] As another example, a possible configuration is to set a playback speed that varies relative to a predetermined interval of timecode in scenes where the game progresses very rapidly. For instance, by setting the playback speed to a delay value such as half speed, users can easily perform manual operations on virtual cameras that are difficult to execute at full speed.
[0239] Note that the following configurations can be adopted: In the case of sports, image analysis can be used to detect when an ongoing sporting event is interrupted, specifically to determine if a referee has made a gesture or if the movement of the ball or player has stopped. Alternatively, the following configuration can be adopted: Image analysis can be used to detect when a sporting event is being played at a rapid pace, based on the speed of movement of the player or ball. Additionally, the following configuration can be adopted: Audio analysis can be used to detect when a sporting event is interrupted or when it is playing at a rapid pace.
[0240] Other embodiments
[0241] The embodiments of the present invention can also be implemented by providing software (programs) that perform the functions of the above embodiments to a system or device via a network or various storage media, and the computer or central processing unit (CPU) or microprocessor unit (MPU) of the system or device reads out and executes the program.
[0242] While the invention has been described with reference to exemplary embodiments, it should be understood that the invention is not limited to the disclosed exemplary embodiments. The scope of the claims should be interpreted as broadly as possible to encompass all such variations and equivalent structures and functions.
Claims
1. An image display device, the image display device comprising: An acquisition unit is configured to acquire first information and second information, the first information indicating multiple locations of a target point of a virtual camera, each of the multiple locations corresponding to a time code within a specific time period, and the second information indicating a pose viewed from the virtual camera. A receiving unit is configured to accept input corresponding to user actions for changing the posture viewed from the virtual camera; A display unit is configured to display a virtual viewpoint image generated based on multiple images, first information, second information, and the input. The virtual viewpoint image is generated by changing the viewing posture from a virtual camera indicated by the second information through the input without changing the position of the target point indicated by the first information.
2. The image display device according to claim 1, wherein, The position of the virtual camera is controlled on a circle centered on the target point.
3. The image display device according to claim 1, wherein, The receiving unit accepts user operations via at least one of a graphical user interface (GUI) and a touch panel displayed on the display unit.
4. The image display device according to claim 3, wherein, When the receiving unit receives a drag or slide operation at the end of the virtual viewpoint image, a first parameter is determined, which corresponds to zooming in or zooming out in the direction of the drag or slide operation.
5. The image display device according to claim 1, wherein, The display unit also displays a graphical user interface, which includes: a first slider through which the timecode of the virtual viewpoint image can be specified; and a second slider through which a portion of the timecode can be magnified to specify the timecode at a finer granularity.
6. The image display device according to claim 1, wherein, The display unit displays the virtual viewpoint image generated according to the update rate of the display unit.
7. The image display apparatus according to claim 1, further comprising a receiving unit configured to receive at least one of the virtual viewpoint image and data for generating the virtual viewpoint image from an external device.
8. A control method for an image display device, the control method comprising: Acquire first information and second information, wherein the first information indicates multiple locations of a target point of a virtual camera, each of the multiple locations corresponding to a time code within a specific time period, and the second information indicates the posture viewed from the virtual camera; Accepts input corresponding to user actions used to change the posture viewed from the virtual camera; A virtual viewpoint image generated based on multiple images, first information, second information, and the input is displayed on the display unit. This virtual viewpoint image is generated by changing the viewing posture from the virtual camera indicated by the second information through the input without changing the position of the target point indicated by the first information.
9. A non-transitory computer-readable storage medium storing a program for causing a computer having a display unit to execute a control method for an image display device, the control method comprising: Acquire first information and second information, wherein the first information indicates multiple locations of a target point of a virtual camera, each of the multiple locations corresponding to a time code within a specific time period, and the second information indicates the posture viewed from the virtual camera; Accepts input corresponding to user actions used to change the posture viewed from the virtual camera; A virtual viewpoint image generated based on multiple images, first information, second information, and the input is displayed on the display unit. This virtual viewpoint image is generated by changing the viewing posture from the virtual camera indicated by the second information through the input without changing the position of the target point indicated by the first information.
Citation Information
Patent Citations
Operation control for refrigerator
JP1989019278A
Graphic system displaying scroll bar
US20090282362A1
Information processing apparatus, control method therefor, and non-transitory computer-readable storage medium
US20180160049A1