Image processing device, image processing method, and program

The image processing device addresses the challenge of maintaining marker alignment in virtual viewpoint images by converting two-dimensional inputs into three-dimensional objects using a plane of interest, ensuring accurate display across different viewpoints.

JP7719811B2Active Publication Date: 2025-08-06CANON KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023002560
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-01-11
Publication Date
2025-08-06
Estimated Expiration
2043-01-11

AI Technical Summary

Technical Problem

Existing technologies face challenges in accurately positioning additional information, such as markers, in virtual viewpoint images when switching viewpoints, especially when mapping three-dimensional space onto two-dimensional space, leading to unintended display positions.

Method used

An image processing device that acquires additional information for a three-dimensional virtual space, sets a plane of interest within the virtual space, converts the information into a three-dimensional position using the plane, and displays the virtual viewpoint image with the object placed at the appropriate position, ensuring alignment with the subject even when the viewpoint changes.

Benefits of technology

Enables accurate and consistent display of additional information in virtual viewpoint images, maintaining alignment with the subject regardless of viewpoint changes, thus improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007719811000001
    Figure 0007719811000001
  • Figure 0007719811000002
    Figure 0007719811000002
  • Figure 0007719811000003
    Figure 0007719811000003
Patent Text Reader

Abstract

To appropriately process input and output of additional information to and from a virtual viewpoint image.SOLUTION: An image processing apparatus acquires input of additional information to a two-dimensional virtual viewpoint image based on a three-dimensional virtual space and a virtual viewpoint in the virtual space, sets an attention surface on which the additional information should be arranged in the virtual space within a range according to the virtual viewpoint in the virtual space, converts the additional information into an object arranged at a three-dimensional position in the virtual space by projecting the input additional information on the attention surface, and displays a virtual viewpoint image based on the virtual space in which the object was arranged at the three-dimensional position.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to advanced techniques for manipulating virtual viewpoint images. [Background technology]

[0002] In a presentation application, a function is known that accepts input of a marker in the shape of a circle or a line to indicate a point of interest within an image while the image is being displayed, and outputs the result by combining the page image and the marker. Patent Document 1 describes a technology in which such a function is applied to a remote conference system. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-151491 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, attention has been focused on a technology for generating an image (virtual viewpoint image) of a captured scene from a plurality of images captured using a plurality of imaging devices, viewed from an arbitrary viewpoint. It is also expected that, for example, a marker will be attached to an object of interest in the scene in such a virtual viewpoint image. When a marker is input into a virtual viewpoint image, the marker will be displayed at an appropriate position when viewed from the viewpoint at the time the marker was input. However, when switching to a different viewpoint, the marker may be displayed at an unintended position. Thus, when additional information such as a marker is drawn on a virtual viewpoint image, the drawn additional information may be displayed at an unintended position. Furthermore, when additional information such as a marker is drawn on an image in which a three-dimensional space is mapped onto a two-dimensional space, it is expected that it is not easy to input the additional information at an appropriate three-dimensional position.

[0005] The present disclosure provides a control technique for inputting and outputting additional information to a virtual viewpoint image. [Means for solving the problem]

[0006] An image processing device according to one aspect of the present disclosure includes: an acquisition unit that acquires input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; a setting means for setting a plane of interest on which the additional information is to be placed in the virtual space within a range corresponding to the virtual viewpoint in the virtual space; a conversion means for converting the input additional information into an object to be placed at a three-dimensional position in the virtual space by projecting the input additional information onto the plane of interest; and a display control means for displaying on a display means the virtual viewpoint image based on the virtual space in which the object is placed at the three-dimensional position. The plane of interest is a surface included in the range of a view frustum of the virtual viewpoint, and the setting means sets the plane of interest by accepting a user setting for the distance from the virtual viewpoint to a point where the plane intersects with the view frustum within the range where the plane intersects with the view frustum, and further accepts a change in the size of the plane of interest. . [Effects of the Invention]

[0007] According to the present disclosure, it is possible to appropriately process the input and output of additional information for a virtual viewpoint image. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 illustrates an example of the configuration of an image processing system. [Figure 2] FIG. 1 illustrates an example of the configuration of an image processing device. [Figure 3] 10A and 10B are diagrams illustrating the position and orientation of a virtual viewpoint and an operation method. [Figure 4] 1A and 1B are diagrams illustrating a virtual viewpoint, a plane of interest, and a marker object. [Figure 5] 10A and 10B are diagrams illustrating a method for manipulating the distance of a surface of interest from a virtual viewpoint. [Figure 6] 10A and 10B are diagrams illustrating a method for setting and manipulating multiple planes of interest. [Figure 7] FIG. 10 is a diagram illustrating an example of a flow of processing executed by an image processing apparatus. [Figure 8] FIG. 10 is a diagram illustrating a marker control process. [Figure 9] FIG. 10 is a diagram illustrating a marker control process. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the disclosure according to the claims. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the disclosure, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0010] [Embodiment 1] (System Configuration) 1(A) and 1(B), an example of the configuration of an image processing system 100 according to this embodiment will be described. The image processing system 100 includes a plurality of sensor systems (n sensor systems 101-1 to 101-n in the example of FIG. 1(A)). Each sensor system includes at least one imaging device (e.g., a camera). Note that, hereinafter, unless there is a particular need to distinguish between them, the sensor systems 101-1 to 101-n will be collectively referred to as "sensor systems 101." This image processing system 100 generates virtual viewpoint image data based on image data acquired by the plurality of sensor systems 101 and provides the generated data to a user.

[0011] FIG. 1B shows an example of the installation of these sensor systems 101. The multiple sensor systems 101 are installed to surround an area that is the subject of the photography (hereinafter referred to as subject area 120), and each sensor system 101 photographs the subject area 120 from a different direction. For example, if the subject area 120 is defined as the field of a stadium where soccer or rugby matches are played, n (a large number, such as 100) sensor systems 101 are installed to surround the field. The number of installed sensor systems 101 is not particularly limited, but should be at least plural. The sensor systems 101 do not have to be installed around the entire periphery of the subject area 120, and may be installed only in a portion of the periphery of the subject area 120 due to, for example, limitations on the installation location. Furthermore, the imaging devices included in each of the multiple sensor systems 101 may include imaging devices with different functions, such as a telephoto camera and a wide-angle camera.

[0012] Furthermore, the sensor system 101 may have a sound collection device (microphone) in addition to an imaging device (camera). The sound collection devices in each of the multiple sensor systems 101 collect sound synchronously. The image processing system 100 generates virtual listening point sound data to be played back together with a virtual viewpoint image based on the sound data collected by the multiple sound collection devices, and provides the generated data to the user. Note that, for the sake of simplicity, a description of sound will be omitted below, but it is assumed that images and sound are processed together.

[0013] Note that subject area 120 is not limited to the field of a stadium, and may be defined to include, for example, stadium seating. Furthermore, subject area 120 may be defined to be an indoor studio, stage, or the like. That is, the area of the subject for which a virtual viewpoint image is to be generated may be defined as subject area 120. Note that the "subject" here may be the area defined by subject area 120 itself, or may include all objects present within that area, such as people, players, referees, and the ball, in addition to or instead of that. Furthermore, throughout this embodiment, the virtual viewpoint image is defined as a moving image, but may also be a still image.

[0014] 1(B), multiple sensor systems 101 synchronize with each other and capture images of a common subject area 120 using the imaging devices of each sensor system 101. In this embodiment, an image included in a group of multiple images obtained by synchronously capturing images of the common subject area 120 from multiple viewpoints is referred to as a "multi-view image." Note that the multi-view image in this embodiment may be the captured image itself, but may also be, for example, an image after image processing such as a process of extracting a predetermined area from the captured image has been performed.

[0015] The image processing system 100 further includes an image recording device 102, a database 103, and an image processing device 104. The image recording device 102 collects multi-viewpoint images obtained by imaging in each of the multiple sensor systems 101, and stores the multi-viewpoint images in the database 103 together with the time code used for imaging. Here, the time code is information for uniquely identifying the time at which imaging was performed. For example, the time code may be information specifying the imaging time in a format such as day:hour:minute:second.frame number.

[0016] The image processing device 104 acquires multiple multi-viewpoint images corresponding to a common time code from the database 103 and generates a three-dimensional model of the subject from the acquired multi-viewpoint images. The three-dimensional model includes, for example, shape information such as a point cloud representing the shape of the subject, or shape information such as faces and vertices when the shape of the subject is represented as a collection of polygons, and texture information representing the color and texture of the surface of the shape. Note that this is just one example, and the three-dimensional model can be defined in any format that represents the subject in three dimensions. The image processing device 104 generates and outputs a virtual viewpoint image corresponding to a virtual viewpoint specified by, for example, a user using a three-dimensional model of the subject. For example, as shown in FIG. 1B, a virtual viewpoint 110 is specified by the position of the viewpoint and the line of sight direction in a virtual space associated with the subject region 120. By moving the virtual viewpoint within the virtual space and changing the direction of the line of sight, the user can view the subject generated based on the three-dimensional model of the subject existing in the virtual space from a viewpoint different from that of any of the imaging devices of the multiple sensor systems 101. Since the virtual viewpoint can move freely within the three-dimensional virtual space, the virtual viewpoint image may also be called a "free viewpoint image."

[0017] The image processing device 104 generates a virtual viewpoint image as an image showing a scene observed from the virtual viewpoint 110. The image generated here is a two-dimensional image. The image processing device 104 is, for example, a computer used by a user, and may be configured to include a display device such as a touch panel display or a liquid crystal display. The image processing device 104 may also have a display control function for displaying an image on an external display device. The image processing device 104 displays the virtual viewpoint image on the screen of the display device, for example. That is, the image processing device 104 generates an image of a scene within a range visible from the virtual viewpoint as a virtual viewpoint image and executes processing for displaying the image on the screen.

[0018] The image processing system 100 may have a configuration different from that shown in FIG. 1A. For example, a configuration including an operation / display device such as a touch panel display separate from the image processing device 104 may be used. For example, a configuration may be used in which a virtual viewpoint or the like is operated on a tablet or the like having a touch panel display, and the image processing device 104 generates a virtual viewpoint image in response to the operation and displays the image on the tablet. A configuration may be used in which multiple tablets are connected to the image processing device 104 via a server, and the image processing device 104 outputs a virtual viewpoint image to each of the multiple tablets. The database 103 and the image processing device 104 may be integrated. A configuration may be used in which the image recording device 102 performs processing from a multi-viewpoint image to generating a three-dimensional model of the subject, and the three-dimensional model of the subject is stored in the database 103. In this case, the image processing device 104 reads the three-dimensional model from the database 103 and generates a virtual viewpoint image. 1(A) shows an example in which a plurality of sensor systems 101 are daisy-chained, but for example, the sensor systems 101 may each be directly connected to the image recording device 102, or may be connected in another manner. Note that, for example, the image recording device 102 or another time synchronization device may be configured to notify each of the sensor systems 101 of reference time information so that the sensor systems 101 can capture images in synchronization.

[0019] In this embodiment, the image processing device 104 further accepts input of markers, such as circles or lines, from the user onto the virtual viewpoint image displayed on the screen and displays the markers superimposed on the virtual viewpoint image. When such a marker is input, the marker is displayed appropriately at the virtual viewpoint where the marker was input. However, changing the position or orientation of the virtual viewpoint may result in an unintended display, such as misalignment of the marker with the object to which the marker is attached. For this reason, in this embodiment, the image processing device 104 executes processing to display the accepted marker at an appropriate position on the displayed two-dimensional screen, regardless of movement of the virtual viewpoint. The image processing device 104 converts the two-dimensional marker into a three-dimensional marker object using at least one surface. The surface used for this conversion may be referred to hereinafter as a plane of interest. The image processing device 104 then aligns the subject represented by a three-dimensional model with the three-dimensional marker object to generate a virtual viewpoint image in which the position of the marker is appropriately adjusted in accordance with movement of the virtual viewpoint. Below, an example of the configuration of the image processing device 104 that executes such processing and a processing flow will be described.

[0020] (Configuration of image processing device) Next, the configuration of the image processing device 104 will be described using FIGS. 2(A) and 2(B). FIG. 2(A) shows an example of the functional configuration of the image processing device 104. The image processing device 104 includes, as its functional configuration, a virtual viewpoint control unit 201, a model generation unit 202, an image generation unit 203, a marker control unit 204, a focus plane control unit 205, and a marker management unit 206. Note that these are merely examples, and at least some of the functions shown may be omitted, or other functions may be added. Furthermore, as long as the functions described below can be executed, some or all of the functions shown in FIG. 2(A) may be replaced by other functional blocks. Furthermore, two or more functional blocks shown in FIG. 2(A) may be combined into one functional block, or one functional block may be divided into multiple functional blocks.

[0021] The virtual viewpoint control unit 201 accepts user operations related to the virtual viewpoint 110 and controls operations such as movement and rotation of the virtual viewpoint. A touch panel, a joystick, or the like is used for user operations of the virtual viewpoint, but the user operations are not limited to these and can be accepted by any device. The virtual viewpoint control unit 201 may also accept and control user operations related to a time code.

[0022] The model generation unit 202 acquires, from the database 103, multi-viewpoint images corresponding to a time code designated by a user operation or the like, and generates a three-dimensional model representing the three-dimensional shape of the subject included in the subject region 120. For example, the model generation unit 202 acquires, from the multi-viewpoint images, foreground images in which a foreground region corresponding to a subject, such as a person or a ball, is extracted, and background images in which a background region other than the foreground region is extracted. The model generation unit 202 then generates a three-dimensional model of the foreground based on the multiple foreground images. The three-dimensional model is formed, for example, by a point cloud generated by a shape estimation method such as Visual Hull. The format of the three-dimensional shape data representing the shape of the object is not limited to this, and meshes or three-dimensional data in a unique format may also be used. The model generation unit 202 can generate a three-dimensional model of the background in a similar manner, but the three-dimensional model of the background may be acquired in advance by an external device. Hereinafter, for convenience, the three-dimensional model of the foreground and the three-dimensional model of the background will be collectively referred to as the "three-dimensional model of the subject" or simply the "three-dimensional model."

[0023] The image generation unit 203 generates a virtual viewpoint image that reproduces a scene as viewed from a virtual viewpoint based on a three-dimensional model of the subject and a virtual viewpoint. For example, the image generation unit 203 acquires appropriate pixel values from the multi-viewpoint image for each point constituting the three-dimensional model and performs coloring processing. The image generation unit 203 then generates the virtual viewpoint image by placing the three-dimensional model in a three-dimensional virtual space and projecting or rendering the three-dimensional model together with the pixel values onto the virtual viewpoint. Note that the method for generating the virtual viewpoint image is not limited to this, and other methods may be used, such as a method for generating a virtual viewpoint image by projective transformation of a captured image without using a three-dimensional model. In addition to the three-dimensional model of the subject, the image generation unit 203 also renders a marker object generated based on a marker input (described later) in the virtual viewpoint image based on the virtual viewpoint and the position of the marker object in the three-dimensional virtual space.

[0024] The marker control unit 204 accepts marker input such as a line or a circle from the user onto the virtual viewpoint image. The marker control unit 204 converts the marker input made onto the two-dimensional virtual viewpoint image into a marker object of three-dimensional data in the same virtual space as the virtual space in which the subject is placed. The attention plane control unit 205, as described below, accepts user input onto the attention plane and controls the attention plane based on the user input. In this embodiment, a plane (attention plane) specified in the virtual space is used when converting the marker input into a marker object. The marker control unit 204 converts the marker input into a marker object using at least one attention plane based on the control of the attention plane control unit 205. These processes will be described later.

[0025] The marker control unit 204 transmits an instruction to the image generation unit 203 to generate a virtual viewpoint image that combines a three-dimensional model of the subject with a marker object according to the position and orientation of the virtual viewpoint. For example, the marker control unit 204 may provide the image generation unit 203 with a marker object as a three-dimensional model, and the image generation unit 203 may generate the virtual viewpoint image by treating the marker object in the same way as the subject. The image generation unit 203 may also execute a process for superimposing the marker object based on the marker object provided by the marker control unit 204, separately from the process for generating the virtual viewpoint image. The marker control unit 204 may also execute a process for superimposing a marker based on the marker object on the virtual viewpoint image provided by the image generation unit 203.

[0026] The marker management unit 206 performs storage control to store information capable of identifying the marker object and the surface of interest of the three-dimensional model obtained by conversion in the marker control unit 204, for example, in the storage unit 216 (described later). The marker management unit 206 performs storage control, for example, so that the information on the marker object and the surface of interest is stored in association with a time code. Note that the model generation unit 202 may calculate the coordinates of each object, such as a person or a ball in the foreground, and store the coordinates in the database 103, and the coordinates of each object may be used to specify the coordinates of the marker object.

[0027] 2(B) shows an example of the hardware configuration of the image processing device 104. The image processing device 104 includes, as its hardware configuration, for example, a CPU 211, RAM 212, ROM 213, an operation unit 214, a display unit 215, a storage unit 216, and an external interface 217. Note that CPU is an abbreviation for Central Processing Unit, RAM is an abbreviation for Random Access Memory, and ROM is an abbreviation for Read Only Memory.

[0028] The CPU 211 controls the entire image processing device 104 and executes the processes described below using programs and data stored in, for example, the RAM 212 and the ROM 213. The CPU 211 executes programs stored in the RAM 212 and the ROM 213, thereby realizing each functional block in FIG. 2(A). The image processing device 104 may have dedicated hardware such as one or more processors other than the CPU 211, and at least a part of the processes performed by the CPU 211 may be executed by the dedicated hardware. The dedicated hardware may be, for example, an MPU (Micro Processing Unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a DSP (Digital Signal Processor). The ROM 213 holds programs and data for executing processes on virtual viewpoint images and markers. The RAM 212 temporarily stores programs and data read from the ROM 213 and provides a work area for the CPU 211 to use when executing each process.

[0029] The operation unit 214 includes a device for receiving user operations, such as a touch panel or a button. The operation unit 214 acquires information indicating, for example, user operations on a virtual viewpoint, a marker, or a plane of interest. The operation unit 214 may be connected to an external controller and receive input information from the user regarding operations. The external controller is not particularly limited, and may be, for example, a three-axis controller such as a joystick, a keyboard, or a mouse. The display unit 215 includes a display device such as a display. The display unit 215 displays, for example, a virtual viewpoint image generated by the CPU 211 or the like. The display unit 215 may also include various output devices capable of presenting information to the user, such as a speaker for audio output or a device for vibration output. The operation unit 214 and the display unit 215 may be integrated into one unit, for example, by a touch panel display.

[0030] The storage unit 216 includes a large-capacity storage device such as an SSD (Solid State Drive) or an HDD (Hard Disk Drive). Note that these are merely examples, and the storage unit 216 may include any other storage device. The storage unit 216 stores data processed by a program. For example, the storage unit 216 stores a three-dimensional marker object obtained by converting a marker input received via the operation unit 214 by the CPU 211. The storage unit 216 may also store other information. The external interface 217 includes, for example, an interface device connected to a network such as a LAN (Local Area Network). Information is transmitted and received between the external interface 217 and an external device such as the database 103. The external interface 217 may also include an image output port such as an HDMI (High-Definition Multimedia Interface) (registered trademark) or an SDI (Serial Digital Interface). In this case, information can be transmitted to an external display device or projection device via the external interface 217. Furthermore, the external interface 217 may be used to connect to a network, and operation information for the virtual viewpoint and markers may be received and virtual viewpoint images may be transmitted via the network.

[0031] (Virtual viewpoint, gaze direction, and focus plane) Next, the virtual viewpoint 110 will be described using FIGS. 3(A) to 3(F). The virtual viewpoint 110 and its movement are specified using a coordinate system that defines a virtual space. In this embodiment, a typical three-dimensional Cartesian coordinate system consisting of an X axis, a Y axis, and a Z axis, as shown in FIG. 3(A), is used as the coordinate system. Note that this is just an example, and any coordinate system that can indicate a position in three-dimensional space may be used. The coordinates of the subject are set and used using this coordinate system. The subject includes, for example, a stadium field or a studio, as well as people and objects such as a ball that exist in the space of the field or studio. For example, in the example of FIG. 3(B), the subject includes the entire stadium field 391, people 392 such as players that exist thereon, and other objects (e.g., a ball 393). Note that the subject may also include spectator seats around the field. Note that a marker object 394 generated from a marker input by a user operation, which will be described later with reference to FIG. 5, is also included in this virtual space. 3(B), the coordinates of the center of field 391 are set as the origin (0,0,0), the X axis is set to the long side direction of field 391, the Y axis is set to the short side direction of field 391, and the Z axis is set to the vertical direction relative to field 391. By setting the coordinates of each subject in this way with the center of field 391 as the reference, a three-dimensional model generated from the subject and a marker object generated based on marker input are placed in a three-dimensional virtual space. Note that the method of setting coordinates is not limited to this.

[0032] Next, the virtual viewpoint will be described using Figures 3(C) and 3(D). The virtual viewpoint determines the viewpoint and line of sight direction for generating a virtual viewpoint image. In Figure 3(C), the apex of a quadrangular pyramid indicates the position 301 of the virtual viewpoint, and the vector extending from the apex indicates the line of sight direction 302. The position 301 of the virtual viewpoint is expressed by coordinates (x, y, z) in three-dimensional virtual space, and the line of sight direction 302 is expressed by a unit vector with the components of each axis as scalars. The line of sight direction 302 is also called the optical axis vector of the virtual viewpoint. The line of sight direction 302 passes through the center points of the front clip plane 303 and the back clip plane 304. Note that the clip planes are planes that define the area to be rendered. The space 305 between the front clip plane 303 and the back clip plane 304 is called the view frustum of the virtual viewpoint, and the virtual viewpoint image is generated within this range (or the virtual viewpoint image is projected and displayed within this range). In this embodiment, a rectangular clipping plane is used as a typical shape for a viewing frustum. However, the clipping plane may be configured with a polygon having more than one side, for example, a pentagon. Three-dimensional data of an object, a marker object, or the like within this range is projected onto a projection plane 306 to generate a virtual viewpoint image. The distance from the virtual viewpoint position 301 to the projection plane 306 is called the focal length. The focal length (not shown) can be set to any value, and changing the focal length changes the angle of view, similar to a typical camera. That is, shortening the focal length widens the angle of view and widens the viewing frustum. On the other hand, lengthening the focal length narrows the angle of view and narrows the viewing frustum, allowing the object to be captured in a larger size. The width and height of the projection plane 306 may be set to the width and height tangent to the viewing frustum 305 of the virtual camera.

[0033] The position of the virtual viewpoint and the line of sight from the virtual viewpoint can be moved and rotated within a virtual space expressed in three-dimensional coordinates. As shown in FIG. 3(D), the movement 307 of the virtual viewpoint is the movement of the virtual viewpoint position 301 and is expressed by the components of each axis (x, y, z). As shown in FIG. 3(A), the rotation 308 of the virtual viewpoint is expressed by yaw, which is a rotation around the Z axis, pitch, which is a rotation around the X axis, and roll, which is a rotation around the Y axis. As a result, the position of the virtual viewpoint and the line of sight from the virtual viewpoint can be freely moved and rotated in the three-dimensional virtual space, and the image processing device 104 can reproduce an image as if any region of the subject were observed from any angle as a virtual viewpoint image. Note that, hereinafter, unless there is a need to distinguish between them, the position of the virtual viewpoint and the line of sight from the virtual viewpoint will be collectively referred to as the "virtual viewpoint."

[0034] A method for operating a virtual viewpoint and a marker will be described using FIGS. 3(E) and 3(F). FIG. 3(E) is a diagram illustrating a screen displayed by the image processing device 104. Here, a case where a tablet-type terminal 320 having a touch panel display is used will be described. Note that the terminal 320 does not need to be a tablet-type terminal, and any other type of information processing device can be used as the terminal 320. When the terminal 320 is the image processing device 104, the terminal 320 is configured to generate and display a virtual viewpoint image and to accept operations such as specifying a virtual viewpoint or a time code and inputting a marker. On the other hand, when the terminal 320 is a device connected to the image processing device 104 via a communication network, for example, the terminal 320 transmits information specifying a virtual viewpoint or a time code to the image processing device 104 and receives the virtual viewpoint image. Furthermore, the terminal 320 accepts a marker input operation for the virtual viewpoint image and transmits information indicating the accepted marker input to the image processing device 104.

[0035] In FIG. 3(E), a display screen 321 on a terminal 320 is roughly divided into two areas: a virtual viewpoint operation area 322 and a time code operation area 323.

[0036] The virtual viewpoint operation area 322 accepts user operations related to the virtual viewpoint, and a virtual viewpoint image is displayed within the area. That is, the virtual viewpoint operation area 322 operates the virtual viewpoint, and a virtual viewpoint image that reproduces a scene as if observed from the virtual viewpoint after the operation is displayed. The virtual viewpoint operation area 322 also accepts marker input for the virtual viewpoint image. Note that, although marker operations and virtual viewpoint operations may be performed together, in this embodiment, the marker operations are accepted independently of virtual viewpoint operations. In one example, as shown in the example of FIG. 3(F), the virtual viewpoint can be operated by touch operations such as tapping and dragging with the user's finger on the terminal 320, and the marker can be operated by tapping and dragging with a drawing device such as a pencil 350. The user moves and rotates the virtual viewpoint, for example, by a drag operation 351 with the finger. The user also draws markers 352 and 353 on the virtual viewpoint image by a drag operation with the pencil 350. The terminal 320 draws a marker at successive coordinates of the drag operation by the pencil 350. Note that the operation by the finger may be assigned to the marker operation, and the operation by the pencil may be assigned to the virtual viewpoint operation. Alternatively, any other operation method may be used as long as the terminal 320 can distinguish between the virtual viewpoint operation and the marker operation. For example, a tap operation may be configured to move or rotate the virtual camera to the coordinates corresponding to the tapped position. This configuration allows the user to easily switch between the virtual viewpoint operation and the marker operation.

[0037] It should be noted that when the virtual viewpoint operation and the marker operation are performed independently, it is not necessary to use a drawing device such as the pencil 350. For example, an ON / OFF button (not shown) for the marker operation may be provided on the touch panel, and whether or not to perform the marker operation may be switched by operating the button. For example, when the marker operation is performed, the button may be turned ON, and while the button is ON, the virtual viewpoint operation may not be performed. Also, when the virtual viewpoint operation is performed, the button may be turned OFF, and while the button is OFF, the marker operation may not be performed.

[0038] The time code operation area 323 is used to specify the timing of the virtual viewpoint image to be viewed. The time code operation area 323 includes, for example, a main slider 331, a sub-slider 332, a speed specification slider 333, and a cancel button 334. The main slider 331 is used to accept any time code selected by the user by, for example, dragging the position of a knob 335. The range of the main slider 331 indicates the entire period during which the virtual viewpoint image can be played back. The sub-slider 332 enlarges a portion of the time code, allowing the user to perform detailed time code operations, such as frame-by-frame operations. The sub-slider 332 accepts the user's selection of any detailed time code by, for example, dragging the position of a knob 336.

[0039] On the terminal 320, the main slider 331 accepts a rough specification of the time code, and the sub-slider 332 accepts a detailed specification of the time code. For example, the main slider 331 and the sub-slider 332 can be set so that the main slider 331 corresponds to a range of three hours, which is the length of the entire game, and the sub-slider 332 corresponds to a portion of that range, about 30 seconds. For example, the sub-slider 332 can represent a section 15 seconds before or after the time code specified by the main slider 331, or a section 30 seconds from that time code. Alternatively, sections may be divided in 30-second increments in advance, and the section of those sections that includes the time code specified by the main slider 331 can be represented by the sub-slider 332. In this way, the main slider 331 and the sub-slider 332 have different time scales. Note that the above-mentioned time lengths are merely examples, and the sliders may be configured to correspond to other time lengths. Note that a user interface may be provided that allows the user to change the setting of the time length corresponding to the sub-slider 332, for example. While FIG. 3(E) shows an example in which the main slider 331 and the sub-slider 332 are displayed for the same length on the screen, their lengths may differ. That is, the main slider 331 may be longer, or the sub-slider 332 may be longer. The sub-slider 332 does not need to be displayed at all times. For example, the sub-slider 332 may be displayed after a display instruction is received, or may be displayed when a specific operation such as pausing is instructed. The time code may be specified and displayed without using the knob 335 of the main slider 421 and the knob 336 of the sub-slider 332. For example, the time code may be specified and displayed on the screen using numerical values, such as values in the format day:hour:minute:second.frame number.

[0040] The speed designation slider 333 is used to accept a user's designation of a playback speed such as normal playback or slow playback. For example, the count-up interval of the time code is controlled according to the playback speed selected using a knob 337 of the speed designation slider 333. The cancel button 334 can be used to cancel each operation related to the time code. The cancel button 334 may also be used to clear a pause and return to normal playback. Note that the button is not limited to the cancel button as long as it is a button for performing an operation related to the time code.

[0041] The plane of interest setting slider 340 and the plane of interest addition button 341 will be described later.

[0042] Using the above-described screen configuration, the user can operate the virtual viewpoint and time code to display a virtual viewpoint image of a three-dimensional model of a subject at an arbitrary time code as viewed from an arbitrary position and posture on the terminal 320. Then, the user can set a plane of interest and perform marker input for such a virtual viewpoint image, independently of operating the virtual viewpoint.

[0043] (Point of interest and marker object) In this embodiment, as described above, a marker input input to a two-dimensional virtual viewpoint image is converted into a marker object of three-dimensional data using a target plane. This will be explained with reference to Figs. 4(A) to 4(E).

[0044] FIG. 4A is a diagram illustrating the relationship between a virtual viewpoint and a plane of interest. In FIG. 4A, the plane of interest 401 is located within the view frustum 305 of the virtual viewpoint. The plane of interest 401 is always set as a plane perpendicular to the line of sight direction 302 (optical axis) of the virtual viewpoint and parallel to the projection plane 306. In this manner, the position of the plane of interest 401 changes depending on the position and orientation of the virtual viewpoint. The initial size of the plane of interest 401 may be set to the width and height at which the plane of interest 401 is in contact with the view frustum 305 of the virtual camera. That is, the size of the plane of interest 401 may be determined so that the four sides of the plane of interest 401 overlap with the four side surfaces of the view frustum 305, respectively. For example, the initial size of the plane of interest 401 may be set to a size that is narrower or wider by a predetermined width on each side than the size at which the plane of interest 401 is in contact with the view frustum 305. That is, the initial size of the plane of interest 401 may be set to correspond to the size of the view frustum 305. However, the initial size of the plane of interest 401 is not limited to this, and may be determined by user input, for example, or may be a predetermined size smaller than the size tangent to the view frustum 305. The distance from the virtual viewpoint position 301 to the plane of interest 401 can be changed within a range 402 in which the plane of interest 401 is included in the view frustum 305. A method for changing the distance from the virtual viewpoint position 301 to the plane of interest 401 will be described later.

[0045] As described above, the plane of interest 401 is located within the range included in the view frustum 305, i.e., between the front clipping plane 303 and the rear clipping plane 304. Therefore, like the three-dimensional data of the subject, the plane of interest 401 can be projected onto the projection plane 306 and rendered as a virtual viewpoint image. For example, as shown in FIG. 4(B), the plane of interest 401 can be rendered within the virtual viewpoint image as a semi-transparent plane 411. By rendering the plane of interest 401 semi-transparently, the superimposition relationship between the plane of interest and the subject in the two-dimensional virtual viewpoint image can be clearly indicated to the user. Note that, hereinafter, this semi-transparent plane 411 may be referred to as the plane of interest 401.

[0046] Here, a first range of the three-dimensional model of the subject that exists in front of a plane of interest 401 as viewed from the virtual viewpoint and a second range that exists behind the plane of interest 401 may be displayed in a distinguishable manner using different colors, for example. FIGS. 4A and 4B show an example in which a three-dimensional model representing a scene in which the subject is pitching a ball is reproduced within the view frustum 305 of the virtual viewpoint and rendered in a virtual viewpoint image. In this example, as shown in FIG. 4B, the plane of interest 401 is set near the subject's arm, and the right arm portion (area 412) located closer to the virtual viewpoint than the plane of interest 401 is represented in a dark color, while the majority of the remaining body (area 413) located behind the plane of interest 401 is represented in a light color. This allows the user to easily and intuitively recognize that the plane of interest 401 is located near the subject's right arm. However, this is just one example, and the portion of the three-dimensional model of the subject that is in front of the plane of interest 401 as viewed from the virtual viewpoint and the portion behind the plane of interest 401 may be distinguishably represented using any expression that is recognizable to the user.

[0047] Here, as shown in FIG. 4(B), assuming that marker input 414 is made to the virtual viewpoint image using pencil 350, a method for converting this marker input 414 into a three-dimensional marker object using a target plane 401 will be described.

[0048] FIG. 4(C) is a conceptual diagram illustrating an extracted display from the virtual camera position 301 to the projection plane 306 in FIG. 4(A). The virtual viewpoint image displayed as in FIG. 4(B) is an image in which a three-dimensional model in the view frustum 305 is projected onto the projection plane 306. Therefore, a marker input 414 made to this virtual viewpoint image can be considered as an input to the projection plane 306 of the virtual viewpoint. Here, the direction and magnitude from the virtual viewpoint position 301 to a point 421 on the trajectory of the marker input 414 input to the projection plane 306 are referred to as a marker input vector 422. Note that the trajectory of the marker input 414 can be considered as a set of points, and a marker input vector can be identified for each point included in the set. However, for simplicity, the following description focuses on one point 421.

[0049] In FIG. 4(C), when the position 301 of the virtual viewpoint is considered in camera coordinates with the origin (0,0,0), the center point of the projection surface, which is the intersection of the line of sight direction 302 (optical axis) of the virtual camera and the projection surface 306, can be expressed as (0,0,f). Here, f is the focal length of the virtual viewpoint. Meanwhile, the coordinates of point 421 on the trajectory of the marker input 414 are assumed to be a point moved from the center point on the projection surface 306 by a in the x direction and b in the y direction. In this case, the marker input vector 422 in the camera coordinates is expressed as [M c ]=(a, b, f). The marker input vector 422 in the camera coordinates is the marker input vector [M W ]=(m x ,m y ,m z ) is converted into a quaternion Q obtained from a rotation matrix that indicates the line of sight 302 of the virtual camera. t Using [M W ]=Q t [M c ] is calculated, and the marker input vector in world coordinates [M W ] is identified. In this way, the marker input vector [M W ] is determined based on the line-of-sight direction 302 of the virtual camera. Note that calculations using the rotation matrix and quaternions of the virtual camera are common, and therefore a description thereof will be omitted.

[0050] In this way, the marker input vector 422 from the virtual viewpoint position 301 to point 421 in the camera coordinate system of FIG. 4(C) is converted into a marker input vector 432 from the virtual viewpoint position 301 to point 431 in the world coordinate system of FIG. 4(D). Then, an intersection 433 between the marker input vector 432 and the target plane 401 is identified by a general method for calculating the intersection between a vector and a plane. Here, the coordinates of the intersection 433 in the world coordinate system are A W =(a x ,a y ,a z ) is expressed as follows. This coordinate A Ware identified as points constituting a marker object corresponding to point 431 on marker input 414. In this way, a marker input vector is identified for each of the points constituting marker input 414, and the intersections of the marker input vectors and the plane of interest 401 are identified as points constituting the marker object. The marker object is then identified by connecting these points. As a result, a marker object of three-dimensional data corresponding to marker input 414 is constructed on plane of interest 401. In this way, as shown in FIG. 4(E), a three-dimensional model 434 of the marker object is generated in the same virtual space as the three-dimensional model of the subject. As a result, marker object 434 is placed in the same three-dimensional space as the subject as a three-dimensional object that is in contact with plane of interest 401.

[0051] If such a marker object is not generated, a mismatch between the subject in the virtual viewpoint image and the intended marker input 414 occurs in response to a movement or rotation of the virtual viewpoint, as shown in FIG. 4(G), making it difficult to understand the content of the marker. On the other hand, by generating a three-dimensional marker object as described above, the positional relationship between the three-dimensional model of the subject and the marker object 434 does not change even if the virtual viewpoint moves or rotates after the marker is input. Therefore, it is possible to prevent (or reduce) the occurrence of a mismatch between the subject and the marker in the virtual viewpoint image projected on the projection surface 306. For example, even if the virtual viewpoint moves or rotates after a marker input is performed as shown in FIG. 4(B), the positional relationship between the subject and the marker 414 can be maintained as shown in FIG. 4(F).

[0052] (Manipulating the distance of the focal plane) Next, an example of a method for manipulating the plane of interest will be described with reference to FIGS. 5(A) to 5(C).

[0053] As explained using Fig. 4(A), the plane of interest 401 moves according to the position of the virtual viewpoint and the direction of the line of sight. Furthermore, the distance from the virtual viewpoint position 301 to the plane of interest 401 within the view frustum 305 can also be changed within the range 402 of the view frustum 305. This change in distance is performed by accepting a user operation using a slider 340 within the display screen 321, as shown in Fig. 3(E), for example.

[0054] For example, suppose that the knob of slider 340 is moved upward by user operation 501 from the situation shown in FIG. 5(A), resulting in the state shown in FIG. 5(B). This operation causes plane of interest 401 to move away from virtual viewpoint position 301 and toward rear clipping plane 304. As a result, as shown in FIG. 5(B), the right half of the subject's body is displayed in a dark color and the remaining left half is displayed in a light color. This means that, compared to FIG. 5(A), a larger portion of the subject is included closer to virtual viewpoint position 301 than plane of interest 401. Then, from the situation shown in FIG. 5(B), user operation 502 causes the knob of slider 340 to move downward, returning it to the same position as in FIG. 5(A), resulting in the state shown in FIG. 5(C). This operation causes plane of interest 401 to move toward virtual viewpoint position 301 and toward front clipping plane 303. As a result, in Figure 5(C), only the right arm of the subject is displayed in a dark color, and the remaining parts are displayed in a light color, so compared to Figure 5(B), less of the subject is included on the virtual viewpoint position 301 side of the plane of interest 401.

[0055] In this way, the distance from the virtual viewpoint position 301 to the plane of interest 401 can be arbitrarily set within the range 402 of the virtual viewpoint's viewing frustum 305 by a user operation. Note that the method of operating and setting the plane of interest is not limited to the above-described method, and other methods may be used to set the plane of interest based on the virtual viewpoint.

[0056] (Multiple focal plane settings) A plurality of planes of interest can be set as described above. A method for manipulating a plurality of planes of interest and a method for generating a marker object using a plurality of planes of interest will be described with reference to Figures 6(A) to 6(E).

[0057] Here, assume that an initial state is one in which one marker 414 has been input as described with reference to FIG. 4B. FIG. 6A shows a state in which, in this state, the user selects and drags the upper left corner of the plane of interest 401. This drag operation changes the size of the plane of interest 401. For example, in the state of FIG. 6A, if the user selects the upper left corner and drags it toward the lower right, the plane of interest 401 is narrowed toward the lower right, and areas that are not set as the plane of interest 401 are generated at the left and top edges. Similarly, if the user selects the lower right corner and drags it toward the upper left, the plane of interest 401 is narrowed toward the upper left, and areas that are not set as the plane of interest 401 are generated at the right and bottom edges. This operation allows the plane of interest 401 to be set only near the right arm of the subject, as shown in FIG. 6B. Note that the user operation can be performed, for example, by dragging from any of the four corners of the plane of interest 401, but it may also be performed, for example, by dragging from any of the four sides. Furthermore, when a corner is selected, an operation to change two sides by moving the corner (vertex) may be accepted, whereas when an edge is selected, only an operation to translate the edge may be accepted. For example, a user may use a pencil 350 or the like to draw an arbitrary closed area within the plane of interest 401, and a rectangular area adjacent to that area may be set as the plane of interest 401 after resizing. Furthermore, the plane of interest 401 does not have to be a rectangular area; its initial shape may be any shape, and changes to any shape may be accepted.

[0058] After the plane of interest 401 is set as shown in FIG. 6(B), for example, the user can set an additional plane of interest 611 as shown in FIG. 6(D) by pressing an add plane of interest button 601 as shown in FIG. 6(C). The same operations as those described with reference to FIGS. 5(A) to 5(C) can be applied to the added plane of interest 611. For example, as shown in FIG. 6(E), when the knob of the slider 340 is moved upward by a user operation 621, the plane of interest 611 moves away from the virtual viewpoint position 301 and toward the rear clipping plane 304 while maintaining the plane of interest 401. As a result, only the left leg of the subject is located behind the plane of interest 611 (toward the rear clipping plane 304), and the rest of the subject is located in front of the plane of interest 611 (toward the front clipping plane 303). In this case, as shown in FIG. 6(F), most of the subject is displayed in a dark color in the plane of interest 611, and the remaining left leg is displayed in a light color.

[0059] Then, as shown in FIG. 6(G), in the state of FIG. 6(F), a marker input 621 may be received by a user operation on the plane of interest 611, for example, near the subject's left leg. By converting this marker input 621 into a marker object as described above, the marker object 631 obtained by the conversion is positioned near the subject's left leg in three-dimensional space as shown in FIG. 6(H). That is, if only one plane of interest can be set, the marker object 434 near the subject's right arm and the marker object 631 may be positioned in the same plane, and the marker object 631 may be positioned away from the left leg, for example. In contrast, by allowing multiple planes of interest to be set, the marker object 434 and the marker object 631 are positioned on different planes. As a result, the marker objects are positioned at appropriate positions in three-dimensional space. Furthermore, even if the virtual viewpoint is moved or rotated after the marker input as shown in FIG. 6(G), the positional relationship between the subject and the markers 414 and 621 can be maintained.

[0060] In FIG. 6(F), since the plane of interest 401 has not moved, the screen display corresponding to the area of the plane of interest 401 does not change. The plane of interest 401 and the plane of interest 611 may be initially set to the same size, shape, and position. In this case, when an operation to add a plane of interest is performed after an operation to change the initial position and size of the plane of interest 401, a separate plane having the same initial position and initial size as the plane of interest 401 may be newly added as the plane of interest 611. In the example of FIG. 6(F), the plane of interest 611 exists behind the plane of interest 401, so the display of the plane of interest 401 does not change. However, for example, if the plane of interest 611 is positioned in front of the plane of interest 401, a semi-transparent plane representing the plane of interest 611 may be displayed superimposed on the plane of interest 401. In this way, different values independent of each other can be set for the plane of interest 401 and the plane of interest 630 with respect to the distance from the virtual viewpoint position 301.

[0061] Here, for example, after two planes of interest have been set, an operation for selecting which of the planes of interest to set may be accepted, so that, for example, the plane of interest 401 may be set again after the plane of interest 611 has been set. For example, in the example of FIG. 6(D), a user operation for the plane of interest 401 may be accepted by selecting an area of the plane of interest 401, and a user operation for the plane of interest 611 may be accepted by selecting an area other than the plane of interest 401. Furthermore, when a double tap with the pencil 350 or the like is accepted on a position where multiple planes of interest are set, the plane of interest to be processed may be sequentially changed. In this case, information that can identify the plane of interest to be processed may be presented. For example, the frame of the plane of interest to be processed may be drawn with a thick line, or a character string indicating the selected plane of interest may be displayed in a space, such as next to the plane of interest add button 601. Furthermore, for example, a pull-down plane of interest selection interface may be provided next to the plane of interest add button 601, allowing a user to select a plane of interest to be operated on from among the multiple set planes of interest.

[0062] For example, in FIG. 6(G), there are two planes of interest, but only one plane of interest 611 exists at the position of the marker input 621. Therefore, when this marker input 621 is accepted, the plane of interest 611 is identified as the plane of interest to be processed. Then, a marker object corresponding to the marker input 621 is generated using this plane of interest 611. On the other hand, when multiple planes of interest are set at the same position on the screen, the process of specifying the plane of interest to be valid may be performed as described above. Note that in the above example, when there are multiple planes of interest at the position where the marker input is made, the user selects which plane of interest to use. However, this is not limited to this. For example, the plane of interest closest to the position of the virtual viewpoint may be used to convert the marker input into a marker object.

[0063] In the above example, the size of the plane of interest 401 is changed before the additional plane of interest 611 is set, but this is not limiting. That is, the size of the plane of interest does not have to be changed. Also, the number of planes of interest may be more than two.

[0064] As described above, by setting and operating the plane of interest, when a marker input is made to a two-dimensional virtual viewpoint image, the marker object can be placed at an appropriate position in three-dimensional space. Even if the virtual viewpoint is operated after the marker input, the three-dimensional models of the subject and the marker are placed at appropriate positions in the virtual space, thereby maintaining the positional relationship between the input content of the subject and the marker. As a result, the marker can be displayed in a position in the generated virtual viewpoint image that does not cause discomfort to the user. Furthermore, by using multiple planes of interest, it is possible to place, for example, multiple marker objects at different distances from the virtual viewpoint position 301 in a single two-dimensional virtual viewpoint image. As a result, it is possible to place multiple marker objects with a high degree of freedom in the three-dimensional space in which the three-dimensional model of the subject is placed.

[0065] (Processing flow) Next, an example of the flow of processing executed by the image processing device 104 will be described with reference to FIGS. 7A and 7B. This processing is configured by a loop process that repeats the processing between S701 and S713, and this loop is executed at a predetermined frame rate. That is, the processing from S702 onward is executed repeatedly at the frame rate period. For example, when the frame rate is 60 FPS, one loop (one frame) of processing is executed at an interval of approximately 16.6 ms. As a result, in S713 described later, a virtual viewpoint image is output at that frame rate. The frame rate can be set to be synchronized with the update rate of the screen display of the image processing device 104 or the like, but may also be set to match the frame rate of the imaging device that captured the multi-viewpoint image or the frame rate of the three-dimensional model stored in the database 103. In the following description, it is assumed that the time code is counted up by one frame each time the loop processing is executed, but the interval at which the time code is counted up may be changed in response to a user operation or the like. For example, if a 1 / 2 playback speed is specified, the time code frame may be counted up once for every two loop processes. Also, if a pause is specified, the time code may stop counting up.

[0066] In the loop processing, the image processing device 104 updates the time code of the processing target (S702). Here, the time code is expressed in the format of day:hour:minute:second:frame, as described above, and may be updated by counting up or the like in units of frames. Then, the image processing device 104 determines whether the received user operation is a virtual viewpoint operation, a marker input operation, or an operation on a plane of interest (S703). Note that the type of operation is not limited to these. For example, if the image processing device 104 receives an operation on the time code, it may return the processing to S702 and update the time code. Furthermore, if the image processing device 104 does not receive a user operation, it may proceed with the processing assuming that a virtual viewpoint designation operation was performed during the immediately preceding virtual viewpoint image generation processing. Furthermore, the image processing device 104 may determine that another operation has been received.

[0067] If the image processing device 104 determines in S703 that it has received a virtual viewpoint operation, it acquires two-dimensional operation coordinates for the virtual viewpoint (S704). The two-dimensional operation coordinates here are, for example, coordinates indicating the position where a tap operation on a touch panel was received. Then, based on the operation coordinates acquired in S704, the image processing device 104 moves and / or rotates the virtual viewpoint in a three-dimensional virtual space (S705). The movement and rotation of the virtual viewpoint have been described with reference to FIG. 3(D) and will not be repeated here. Furthermore, the process of determining the amount of movement and rotation of the virtual viewpoint in three-dimensional space from the two-dimensional coordinates obtained by a touch operation on the touch panel can be performed using well-known technology and will not be described in detail here. In response to the movement or rotation of the virtual viewpoint performed in the process of S705, the image processing device 104 moves the plane of interest within a viewing frustum defined by the virtual viewpoint after the movement or rotation (S706). Then, after or in parallel with the processes of S704 to S706, the image processing device 104 determines whether a marker object exists within the field of view determined by the virtual viewpoint after the movement or rotation (S707). If a marker object exists within the field of view determined by the virtual viewpoint after the movement or rotation (YES in S707), the image processing device 104 reads out the marker object and places it in a three-dimensional virtual space (S708). After placing the marker object, the image processing device 104 generates and outputs a virtual viewpoint image including the marker object (S712). That is, the image processing device 104 generates a virtual viewpoint image using a three-dimensional model of the subject corresponding to the time code updated in S702 and the marker object placed in the virtual space, depending on the virtual viewpoint. Note that if a marker object does not exist within the field of view determined by the virtual viewpoint after the movement or rotation (NO in S707), the image processing device 104 generates and outputs a virtual viewpoint image without placing the marker object (S712).

[0068] When it is determined in S703 that a marker input operation has been received, the image processing device 104 acquires two-dimensional coordinates of the marker operation in the virtual viewpoint image (S709). That is, as described with reference to FIGS. 3(E) and 3(F), the image processing device 104 acquires two-dimensional coordinates of the marker operation input with respect to the virtual viewpoint operation area 322. Then, the image processing device 104 converts the two-dimensional coordinates of the marker input acquired in S709 into a marker object, which is three-dimensional data in a plane of interest in the virtual space (S710). The method of converting the marker input into a marker object has been described with reference to FIGS. 4(A) to 4(E), and therefore will not be described again here. When the image processing device 104 acquires the marker object in S710, it stores the marker object together with a time code (S711). Then, the image processing device 104 generates a virtual viewpoint image using the three-dimensional model of the subject corresponding to the time code updated in S702 and the marker object placed in the virtual space in accordance with the virtual viewpoint (S712). As a result, when a virtual viewpoint image corresponding to the time code at the time of marker input is generated and displayed, not only the subject is drawn but also the marker object placed in the virtual space in response to the marker input can be drawn and displayed.

[0069] When it is determined in S703 that a plane-of-interest operation has been received, the image processing device 104 determines the content of the operation (S751). For example, the image processing device 104 changes the distance from the position of the virtual viewpoint to the plane of interest as shown in Figures 5(A) to 5(C) (S752), adds a plane of interest as shown in Figures 6(C) or 6(D) (S753), or changes the size of the plane of interest as shown in Figure 6(A) (S754).

[0070] According to the above processing, a marker input to a two-dimensional virtual viewpoint image can be converted into three-dimensional data using a plane of interest in the virtual space, and a three-dimensional model of the subject and a marker object can be placed in the same virtual space. Therefore, regardless of the position and orientation of the virtual viewpoint, it is possible to generate a virtual viewpoint image in which the positional relationship between the subject and the marker is maintained. Furthermore, by setting an appropriate plane of interest in the three-dimensional space when inputting the marker, it is possible to place the marker object at an appropriate position in the virtual space. Furthermore, by setting multiple planes of interest, it is possible to place multiple marker objects at appropriate positions in the three-dimensional space.

[0071] [Embodiment 2] In the first embodiment, an example of a procedure for generating a marker object in one frame was described. In the present embodiment, an example of a procedure for generating marker objects across multiple consecutive frames will be described with reference to Figs. 8(A) to 8(F). The system configuration and device configuration of this embodiment are the same as those of the first embodiment, and therefore the description will not be repeated here.

[0072] FIG. 8(A) shows a virtual viewpoint image in the same state as that shown in FIG. 4(B) in the first embodiment. However, the screen in FIG. 8(A) provides a continuation instruction button 801 indicating that marker input is to be continued across multiple frames. By operating this continuation instruction button 801, the user can enable or disable the function of continuing marker input across multiple frames. Note that an interface format other than a button format may be used as long as it is possible to specify that marker input is to be continued across multiple frames. In this embodiment, it is assumed that in FIG. 8(A), first, a user operation 802 of pressing the continuation instruction button 801 is accepted, and then a marker input 803 is accepted. For example, it is assumed that a marker input 803 for the vicinity of a subject's arm is accepted at the timing when the subject starts to swing the arm. When the marker input 803 is accepted, a marker object is generated based on the marker input and the plane of interest, as described in the first embodiment. It is assumed that a marker object 811 shown in Fig. 8(B) is generated from the marker input 803 in Fig. 8(A) and placed in the same three-dimensional space as the subject. For convenience, the time code at this point is assumed to be ww:xx:yy:zz.000.

[0073] Thereafter, as shown in FIG. 8(C), an operation 821 on the time code main slider 331 or sub-slider 332 is accepted, and a time code different from that shown in FIG. 8(A) is specified. Here, it is assumed that the time code at this point is ww:xx:yy:zz.010. FIG. 8(C) shows a scene in which the subject has finished waving their arm, and shows a state in which a marker input 822 extending to the vicinity of the arm where the waving has finished is accepted. When this marker input 822 is accepted, a corresponding marker object 831 is generated, as shown in FIG. 8(D), and the marker object 831 is positioned so as to extend to the vicinity of the arm where the subject has finished waving. Finally, a user input 823 of pressing the continuation instruction button 801 again is accepted, and the instruction to continue the marker input is released.

[0074] As a result, a marker object 831 shown in FIG. 8(D) is generated between time codes ww:xx:yy:zz.000 and ww:xx:yy:zz.010. After the marker object 831 is generated, the generated marker object is continuously placed across multiple frames corresponding to the time code period for which the continuation instruction was issued. For example, after the marker object 831 is generated as shown in FIGS. 8(C) and 8(D), a user operation 841 may be received on the time code main slider 331 or sub-slider 332, as shown in FIG. 8(E). Here, it is assumed that the time code ww:xx:yy:zz.000 is specified by the user operation 841. This time code corresponds to the scene where the subject begins to swing their arms, and is the time code for which the marker input 822 has not yet been completed. However, by specifying this time code after the marker object 831 is completed, a display is produced that appears to indicate that the marker input 822 has been completed, as shown in FIG. 8(E). 8(F), the generated marker object 831 is placed in the same three-dimensional space as the subject at time code ww:xx:yy:zz.000. In this way, the marker object 831 continues to be placed near the subject during the period from consecutive time codes ww:xx:yy:zz.000 to ww:xx:yy:zz.010. Note that the marker object 831 may not be placed in the virtual space outside of this period.

[0075] As described above, according to this embodiment, when a marker input is received for one of multiple frames of a virtual viewpoint image, a marker object can be generated that is placed across the multiple frames while maintaining the positional relationship between the subject and the marker. This makes it possible to display a common marker at an appropriate position in a virtual viewpoint image that is drawn across multiple frames over a certain period of time. Furthermore, since marker input can be performed while referring to the movement of the virtual viewpoint image across multiple frames, the user can perform marker input that is more suited to the movement of the subject.

[0076] [Embodiment 3] In this embodiment, a process for editing a marker object generated by the procedure described in Embodiment 1 or Embodiment 2 will be described. In this embodiment, the system configuration and device configuration are the same as those in Embodiment 1, so the description will not be repeated here.

[0077] First, Fig. 9(A) shows a state in which a subject is captured from the side with the same virtual viewpoint position and line of sight direction as Fig. 4(A). Fig. 9(B) shows a state in which a marker input 901 is received by the virtual camera of Fig. 9(A) on the plane 411 corresponding to the plane of interest 401. This marker input 901 is converted into a marker object as described above.

[0078] Thereafter, in order to edit the marker object corresponding to this marker input 901, it is assumed that the position of the virtual viewpoint and the line of sight direction are changed as shown in Fig. 9(C). Fig. 9(C) shows a state in which a virtual viewpoint position 911 and a line of sight direction 912 of the virtual viewpoint are set so that the subject is viewed from above. Even in this state, a 3D model of the subject and a marker object corresponding to the marker input 901 are placed within the range of a view frustum 915 between a front clip plane 913 and a rear clip plane 914. A plane of interest 916 is also set within the range of the view frustum 915. The subject and markers are projected onto a projection plane 917, and the projected image is shown as shown in Fig. 9(D).

[0079] A marker 922 onto which a marker object is projected is displayed on the screen in FIG. 9(D). The user can change the shape of the marker 922 by dragging a point on the marker 922 using, for example, the pencil 350. The shape of the marker 922 can be edited, for example, within a range that contacts the plane of interest 916. As a result, for example, the shape of the marker object viewed from above can be changed to a wavy shape, as shown by the marker 931 in FIG. 9(E). Note that this is just one example, and editing to any shape is naturally possible. Note that the user can also input a new marker by, for example, selecting a location where the marker 922 does not exist using the pencil 350. That is, an operation on a location where a marker already exists on the screen of the virtual viewpoint image may be accepted as a marker editing operation, and an operation on a location where no marker exists may be accepted as a new marker input operation. Alternatively, a separate operation mode such as a marker transformation mode may be defined, and while that mode is active, new marker input may not be accepted and only editing of existing markers may be accepted. Also, while that mode is inactive, tapping a marker may not result in editing of the marker.

[0080] In this embodiment, the shape of a marker object generated by marker input can be edited in this way, making it possible to generate a marker object with a more appropriate shape.

[0081] Although the above-described embodiment describes processing when a marker is added to a virtual viewpoint image based on multi-viewpoint images captured by multiple imaging devices, this is not limiting. For example, even when a marker is added to a virtual viewpoint image generated based on a three-dimensional virtual space entirely artificially created on a computer, the marker may be converted into a three-dimensional object within the virtual space. Furthermore, the above-described embodiment describes an example in which a marker object associated with a time code corresponding to a virtual viewpoint image to which a marker is added is generated and stored. However, the time code does not have to be associated with the marker object. For example, when the virtual viewpoint image is a still image or when the virtual viewpoint image is used only for temporarily adding a marker in a conference, the marker object may be displayed or erased, for example, by a user operation, regardless of the time code. Furthermore, at least some of the image processing devices used in, for example, a conference system may not have the ability to specify a virtual viewpoint. That is, after a marker is added to a virtual viewpoint image, it is sufficient that only a specific user, such as a person in charge of the conference, can specify the virtual viewpoint. Image processing devices owned by other users do not need to accept virtual viewpoint operations. In this case too, the marker object is drawn according to the virtual viewpoint and plane of interest specified by a specific user, so that it is possible to prevent the relationship between the marker and the subject of the virtual viewpoint image from becoming inconsistent.

[0082] In the above-described embodiment, an example has been described in which a marker object is displayed as additional information displayed on a virtual viewpoint image. However, the additional information displayed on a virtual viewpoint image is not limited to this. For example, at least one of a marker, an icon, an avatar, an illustration, etc. designated by a user may be displayed on the virtual viewpoint image. Alternatively, a plurality of pieces of additional information may be prepared in advance, and a user may select and place any one of them on the virtual viewpoint image. Alternatively, a user may be allowed to drag an icon or the like on a touch panel display to place it at any position. The placed additional information, such as an icon, is converted into three-dimensional data using the position of the virtual viewpoint and the plane of interest in a manner similar to that of the above-described embodiment. However, this is not a limitation. Alternatively, two-dimensional data additional information and three-dimensional data additional information may be associated in advance, and the two-dimensional data additional information may be replaced with the corresponding three-dimensional data additional information when the two-dimensional data additional information is placed. In this case, the three-dimensional data may be prepared in advance for each distance between the position of the virtual viewpoint and the plane of interest, for example, so that the data corresponds to the distance between the position of the virtual viewpoint and the plane of interest. In this case, corresponding three-dimensional data may be selected and placed in three-dimensional space depending on the setting of the position of the virtual viewpoint and the plane of interest at the time the two-dimensional data additional information is placed. Furthermore, by transforming this three-dimensional data as in the above-described embodiment 3, three-dimensional data in a more appropriate position or shape can be used as additional information. In this way, this embodiment is applicable to cases where various types of additional information are displayed on a virtual viewpoint image.

[0083] In the above embodiment, the additional information (marker input) is converted into three-dimensional data (marker object), but the three-dimensional data does not necessarily have to represent a three-dimensional shape. That is, the three-dimensional data is data that at least has a three-dimensional position in virtual space, and the shape of the additional information may be a plane, a line, or a point.

[0084] In the above-described embodiment, an example in which the plane of interest is a plane within the range of the view frustum has been described, but this is not limiting. For example, the plane of interest may be a curved surface. As an example in which the plane of interest is a curved surface, a hemispherical (or full-sphere) surface with a virtual viewpoint as its center and a radius a certain distance from the virtual viewpoint may be used as the plane of interest. In this case, the plane of interest may be set by specifying only the distance from the virtual viewpoint. Furthermore, a curved surface obtained by cutting such a hemispherical surface with a view frustum may be used as the plane of interest. Alternatively, a hemispherical (or full-sphere) surface with a center at the position of a subject in three-dimensional space and a radius a certain distance from that position may be used as the plane of interest. In this case, the plane of interest may be set by specifying only the distance from the subject. Furthermore, when there are multiple subjects, the subjects may be specified to set the plane of interest, or a plane of interest may be set for each of the multiple subjects. Furthermore, the plane of interest set based on the subject may not be a curved surface but may be a plane. Furthermore, instead of a plane with a size limited to the range of the view frustum as described above, the plane may be a plane of any size (e.g., spanning the entire range of the virtual space) that includes the plane. In this case, the distance from the virtual viewpoint to the plane of interest can be set within the range where the plane intersects with the view frustum. The case where this plane is clipped within the range where it intersects with the view frustum corresponds to the case described with reference to FIG. 4A, for example. This restriction is intended to prevent objects located closer to the virtual viewpoint than the front clipping plane and farther from the virtual viewpoint than the rear clipping plane from being rendered in the virtual viewpoint image. However, this is merely an example, and the plane of interest may be set at a position closer to the virtual viewpoint than the front clipping plane or farther from the virtual viewpoint than the rear clipping plane. In this case, additional information is not rendered together with the subject in the virtual viewpoint image. However, by changing the position or focal length of the virtual viewpoint, for example, it becomes possible to generate a virtual viewpoint image showing the additional information. This allows additional information, such as the position of the subject or a description of the subject, to be placed in the virtual space.

[0085] Furthermore, it is not necessary to have all of the functions described in the above embodiments, and it is possible to implement any combination of functions.

[0086] [Other embodiments] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0087] The present disclosure is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the present disclosure.

[0088] (Summary of the embodiment) At least some of the above-described embodiments can be summarized as follows. (Item 1) an acquisition means for acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; a setting means for setting a plane of interest in the virtual space on which the additional information is to be arranged within a range corresponding to the virtual viewpoint in the virtual space; a conversion means for converting the input additional information into an object to be placed at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; a display control means for causing a display means to display the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; 1. An image processing device comprising: (Item 2) the plane of interest is a plane perpendicular to a vector indicating a line of sight from the position of the virtual viewpoint, the setting means sets the plane of interest by accepting a user setting of a distance from the virtual viewpoint to a point where the plane and the vector intersect. 2. The image processing device according to item 1, (Item 3) the plane of interest is a plane that is included in the range of a viewing frustum at the virtual viewpoint, among planes that perpendicularly intersect with a vector that indicates a line of sight from the position of the virtual viewpoint, the setting means sets the plane of interest by accepting a setting by a user regarding a distance from the virtual viewpoint to a point where the plane intersects with the vector within a range where the plane intersects with the view frustum. 2. The image processing device according to item 1, (Item 4) 4. The image processing device according to item 3, wherein the setting means further accepts a change in the size of the plane of interest. (Item 5) 5. The image processing device according to item 3 or 4, wherein the plane of interest has an initial size corresponding to the size of the view frustum. (Item 6) the plane of interest is a hemispherical surface centered on the position of the virtual viewpoint, the setting means sets the plane of interest by accepting a user setting of a distance from the virtual viewpoint to a point where the plane intersects with a vector indicating a line of sight from the position of the virtual viewpoint. 2. The image processing device according to item 1, (Item 7) 7. The image processing device according to any one of items 1 to 6, wherein the setting means sets a plurality of the planes of interest when an operation to add the plane of interest is received. (Item 8) 8. The image processing device according to item 7, wherein, when a plurality of planes of interest are set, the conversion means converts the additional information into the object using the plane of interest corresponding to the position where the additional information is input in the virtual viewpoint image. (Item 9) Item 9. The image processing device according to item 7 or 8, wherein, when there are a plurality of planes of interest corresponding to the position at which the additional information is input in the virtual viewpoint image, the conversion means converts the additional information into the object using the plane of interest that is closest in distance from the position of the virtual viewpoint. (Item 10) Item 9. The image processing device according to item 7 or 8, wherein, when there are a plurality of planes of interest corresponding to positions where the additional information is input in the virtual viewpoint image, the conversion means accepts a user selection of which of the plurality of planes of interest to use to convert the additional information into the object. (Item 11) a time code corresponding to the virtual viewpoint image at the time when the additional information is input is associated with the object; when the virtual viewpoint image corresponding to the time code associated with the object is displayed, the display control means displays the virtual viewpoint image based on the virtual space in which the object is placed. 11. The image processing device according to any one of items 1 to 10, characterized in that: (Item 12) the acquisition means has a function of receiving input of the additional information for a plurality of consecutive frames of the virtual viewpoint image, the conversion means converts the additional information input in each of the plurality of frames for which the function is enabled into an object to be placed at a three-dimensional position in the virtual space during a period corresponding to the plurality of frames. 12. The image processing device according to any one of items 1 to 11, characterized in that: (Item 13) 13. The image processing device according to any one of items 1 to 12, further comprising means for accepting edits to the additional information included in the virtual viewpoint image displayed on the display means. (Item 14) Item 14. The image processing device according to item 13, characterized in that when a position where the additional information is displayed is selected in the virtual viewpoint image displayed on the display means, editing of the additional information is performed, and when a position where the additional information is not displayed is selected, new additional information is input. (Item 15) 15. The image processing device according to any one of items 1 to 14, wherein the additional information includes at least one of a marker, an icon, an avatar, and an illustration. (Item 16) 16. The image processing device according to any one of items 1 to 15, wherein the virtual viewpoint image is generated based on multi-viewpoint images acquired by capturing images using a plurality of imaging devices. (Item 17) An image processing method executed by an image processing device, comprising: acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; setting a plane of interest in which the additional information is to be arranged in the virtual space within a range corresponding to the virtual viewpoint in the virtual space; converting the input additional information into an object that is arranged at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; displaying, on a display means, the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; An image processing method comprising: (Item 18) 17. A program for causing a computer to function as the image processing device according to any one of items 1 to 16. [Explanation of symbols]

[0089] 201: Virtual viewpoint control unit, 202: Model generation unit, 203: Image generation unit, 204: Marker control unit, 205: Attention plane control unit, 206: Marker management unit

Claims

1. an acquisition means for acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; a setting means for setting a plane of interest in the virtual space on which the additional information is to be arranged within a range corresponding to the virtual viewpoint in the virtual space; a conversion means for converting the input additional information into an object to be placed at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; a display control means for causing a display means to display the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; and the plane of interest is a plane included in the range of a viewing frustum of the virtual viewpoint, the setting means sets the plane of interest by accepting a setting by a user regarding a distance from the virtual viewpoint to a point where the plane intersects with a vector indicating a line of sight from the position of the virtual viewpoint, within a range where the plane intersects with the view frustum, and further accepts a change in the size of the plane of interest.

1. An image processing device comprising:

2. The plane of interest is a plane perpendicular to the vector.

2. The image processing device according to claim 1, wherein:

3. An acquisition means for acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; a setting means for setting a plane of interest in the virtual space on which the additional information is to be arranged within a range corresponding to the virtual viewpoint in the virtual space; a conversion means for converting the input additional information into an object to be placed at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; a display control means for causing a display means to display the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; and the plane of interest is a plane included in the range of a viewing frustum of the virtual viewpoint, 10. An image processing device, wherein the plane of interest has an initial size corresponding to the size of the view frustum.

4. An acquisition means for acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; a setting means for setting a plane of interest in the virtual space on which the additional information is to be arranged within a range corresponding to the virtual viewpoint in the virtual space; a conversion means for converting the input additional information into an object to be placed at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; a display control means for causing a display means to display the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; and the plane of interest is a hemispherical surface centered on the position of the virtual viewpoint, the setting means sets the plane of interest by accepting a user setting of a distance from the virtual viewpoint to a point where the plane intersects with a vector indicating a line of sight from the position of the virtual viewpoint.

1. An image processing device comprising:

5. An acquisition means for acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; a setting means for setting a plane of interest in the virtual space on which the additional information is to be arranged within a range corresponding to the virtual viewpoint in the virtual space; a conversion means for converting the input additional information into an object to be placed at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; a display control means for causing a display means to display the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; and The setting means sets a plurality of the target planes, The image processing device is characterized in that, when multiple planes of interest are set, the conversion means converts the additional information into the object using the plane of interest that corresponds to the position where the additional information is input in the virtual viewpoint image.

6. An acquisition means for acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; a setting means for setting a plane of interest in the virtual space on which the additional information is to be arranged within a range corresponding to the virtual viewpoint in the virtual space; a conversion means for converting the input additional information into an object to be placed at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; a display control means for causing a display means to display the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; and The setting means sets a plurality of the target planes, The image processing device is characterized in that, when there are multiple planes of interest corresponding to the position where the additional information is input in the virtual viewpoint image, the conversion means converts the additional information into the object using the plane of interest that is closest in distance from the position of the virtual viewpoint.

7. An acquisition means for acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; a setting means for setting a plane of interest in the virtual space on which the additional information is to be arranged within a range corresponding to the virtual viewpoint in the virtual space; a conversion means for converting the input additional information into an object to be placed at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; a display control means for causing a display means to display the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; and The setting means sets a plurality of the target planes, The image processing device is characterized in that, when there are multiple planes of interest corresponding to the position where the additional information is input in the virtual viewpoint image, the conversion means accepts a user selection of which of the multiple planes of interest to use to convert the additional information into the object.

8. An acquisition means for acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; a setting means for setting a plane of interest in the virtual space on which the additional information is to be arranged within a range corresponding to the virtual viewpoint in the virtual space; a conversion means for converting the input additional information into an object to be placed at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; a display control means for causing a display means to display the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; and a time code corresponding to the virtual viewpoint image at the time when the additional information is input is associated with the object; when the virtual viewpoint image corresponding to the time code associated with the object is displayed, the display control means displays the virtual viewpoint image based on the virtual space in which the object is placed.

1. An image processing device comprising:

9. An acquisition means for acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; a setting means for setting a plane of interest in the virtual space on which the additional information is to be arranged within a range corresponding to the virtual viewpoint in the virtual space; a conversion means for converting the input additional information into an object to be placed at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; a display control means for causing a display means to display the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; and the acquisition means has a function of receiving input of the additional information for a plurality of consecutive frames of the virtual viewpoint image, the conversion means converts the additional information input in each of the plurality of frames for which the function is enabled into an object to be placed at a three-dimensional position in the virtual space during a period corresponding to the plurality of frames.

1. An image processing device comprising:

10. 2. The image processing apparatus according to claim 1, further comprising means for accepting edits to the additional information included in the virtual viewpoint image displayed on the display means.

11. 11. The image processing device according to claim 10, wherein editing of the additional information is performed when a position where the additional information is displayed is selected in the virtual viewpoint image displayed on the display means, and new additional information is input when a position where the additional information is not displayed is selected.

12. 2. The image processing device according to claim 1, wherein the additional information includes at least one of a marker, an icon, an avatar, and an illustration.

13. An acquisition means for acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; a setting means for setting a plane of interest in the virtual space on which the additional information is to be arranged within a range corresponding to the virtual viewpoint in the virtual space; a conversion means for converting the input additional information into an object to be placed at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; a display control means for causing a display means to display the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; and 10. An image processing device, wherein the virtual viewpoint image is generated based on multi-viewpoint images acquired by capturing images using a plurality of imaging devices.

14. An image processing method executed by an image processing device, comprising: acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; setting a plane of interest in which the additional information is to be arranged in the virtual space within a range corresponding to the virtual viewpoint in the virtual space; converting the input additional information into an object that is arranged at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; displaying, on a display means, the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; Including, the plane of interest is a plane included in the range of a viewing frustum of the virtual viewpoint, In the setting, the setting by the user regarding the distance from the virtual viewpoint to the point where the plane intersects with a vector indicating the line of sight from the position of the virtual viewpoint is accepted within a range where the plane intersects with the viewing frustum, thereby setting the plane of interest, and further accepting a change in the size of the plane of interest. An image processing method comprising:

15. An image processing method executed by an image processing device, comprising: acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; setting a plane of interest in which the additional information is to be arranged in the virtual space within a range corresponding to the virtual viewpoint in the virtual space; converting the input additional information into an object that is arranged at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; displaying, on a display means, the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; Including, the plane of interest is a plane included in the range of a viewing frustum of the virtual viewpoint, 10. An image processing method, wherein the surface of interest has an initial size corresponding to the size of the view frustum.

16. An image processing method executed by an image processing device, comprising: acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; setting a plane of interest in which the additional information is to be arranged in the virtual space within a range corresponding to the virtual viewpoint in the virtual space; converting the input additional information into an object that is arranged at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; displaying, on a display means, the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; Including, the plane of interest is a hemispherical surface centered on the position of the virtual viewpoint, In the setting, the plane of interest is set by accepting a user setting of a distance from the virtual viewpoint to a point where the plane intersects with a vector indicating the direction of the line of sight from the position of the virtual viewpoint.

17. An image processing method executed by an image processing device, comprising: acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; setting a plane of interest in which the additional information is to be arranged in the virtual space within a range corresponding to the virtual viewpoint in the virtual space; converting the input additional information into an object that is arranged at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; displaying, on a display means, the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; Including, In the setting, a plurality of planes of interest are set; In the conversion, when a plurality of planes of interest are set, the plane of interest corresponding to the position where the additional information is input in the virtual viewpoint image is used to convert the additional information into the object.

18. An image processing method executed by an image processing device, comprising: acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; setting a plane of interest in which the additional information is to be arranged in the virtual space within a range corresponding to the virtual viewpoint in the virtual space; converting the input additional information into an object that is arranged at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; displaying, on a display means, the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; Including, In the setting, a plurality of planes of interest are set; In the conversion, when there are a plurality of planes of interest corresponding to the position where the additional information is input in the virtual viewpoint image, the plane of interest that is closest in distance from the position of the virtual viewpoint is used to convert the additional information into the object.

19. An image processing method executed by an image processing device, comprising: acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; setting a plane of interest in which the additional information is to be arranged in the virtual space within a range corresponding to the virtual viewpoint in the virtual space; converting the input additional information into an object that is arranged at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; displaying, on a display means, the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; Including, In the setting, a plurality of planes of interest are set; an image processing method characterized in that, in the conversion, when there are multiple planes of interest corresponding to the position where the additional information is input in the virtual viewpoint image, a user's selection is accepted as to which of the multiple planes of interest to use to convert the additional information into the object.

20. An image processing method executed by an image processing device, comprising: acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; setting a plane of interest in which the additional information is to be arranged in the virtual space within a range corresponding to the virtual viewpoint in the virtual space; converting the input additional information into an object that is arranged at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; displaying, on a display means, the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; Including, a time code corresponding to the virtual viewpoint image at the time when the additional information is input is associated with the object; An image processing method characterized by displaying a virtual viewpoint image based on the virtual space in which the object is placed when the virtual viewpoint image corresponding to the time code associated with the object is displayed.

21. An image processing method executed by an image processing device, comprising: acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; setting a plane of interest in which the additional information is to be arranged in the virtual space within a range corresponding to the virtual viewpoint in the virtual space; converting the input additional information into an object that is arranged at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; displaying, on a display means, the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; Including, the image processing device has a function of receiving input of the additional information for a plurality of consecutive frames of the virtual viewpoint image, an image processing method characterized in that, in the conversion, the additional information input in each of the plurality of frames for which the function is enabled is converted into an object that is placed at a three-dimensional position in the virtual space during a period corresponding to the plurality of frames.

22. An image processing method executed by an image processing device, comprising: acquiring input of additional information for a three-dimensional virtual space and a two-dimensional virtual viewpoint image based on a virtual viewpoint in the virtual space; setting a plane of interest in which the additional information is to be arranged in the virtual space within a range corresponding to the virtual viewpoint in the virtual space; converting the input additional information into an object that is arranged at a three-dimensional position in the virtual space by projecting the input additional information onto the target surface; displaying, on a display means, the virtual viewpoint image based on the virtual space in which the object is arranged at the three-dimensional position; Including, An image processing method, characterized in that the virtual viewpoint image is generated based on multi-viewpoint images acquired by capturing images using a plurality of imaging devices.

23. A program for causing a computer to function as the image processing device according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Correction system for graphic in stroke input

    JP1993159038A

  • Curve correction method and apparatus and program

    JP2008262253A

  • Image display device, image processing system, image processing method, and image processing program

    JP2017151491A

  • Method for defining drawing planes for design of 3D object

    JP2019121387A