Information Processing Apparatus, Information Processing Method, and Program

By generating additional objects with a three-dimensional shape and ensuring they are displayed on a plane parallel to the ground plane in virtual viewpoint images, the system addresses the issue of inconsistent marker placement across different viewpoints, achieving accurate and reliable display of additional information.

JP7682251B2Active Publication Date: 2025-05-23CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023210281
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-05-23
Estimated Expiration
2041-07-28

AI Technical Summary

Technical Problem

In virtual viewpoint images, additional information such as markers may be displayed at unintended positions when switching between viewpoints, leading to inconsistencies in their placement.

Method used

The system generates an additional object with a three-dimensional shape based on the additional information and ensures it is displayed on a plane parallel to the ground plane of background objects in the virtual space, maintaining its appropriate position across different viewpoints.

Benefits of technology

This approach allows additional information to be consistently displayed at appropriate positions regardless of the viewpoint, enhancing the accuracy and reliability of virtual viewpoint image manipulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682251000001
    Figure 0007682251000001
  • Figure 0007682251000002
    Figure 0007682251000002
  • Figure 0007682251000003
    Figure 0007682251000003
Patent Text Reader

Abstract

To display additional information drawn with respect to a virtual viewpoint image to a proper position regardless of a viewpoint.SOLUTION: An image processing apparatus according to the present invention acquires an input of additional information for a two-dimensional virtual viewpoint image based on a virtual viewpoint and a three-dimensional virtual space, converts the input additional information into an object located in a three-dimensional position in the virtual space, and displays the virtual viewpoint image based on the virtual space where the object is located in the three-dimensional position.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to techniques for enhancing the manipulation of virtual viewpoint images. [Background technology]

[0002] In a presentation application, a function is known that accepts input of a marker in the shape of a circle or a line to indicate a point of interest in an image while the image is being displayed, and outputs the page image by combining the marker with the image. Patent Document 1 describes a technology in which such a function is applied to a remote conference system. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2017-151491 A Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, a technology that generates an image (virtual viewpoint image) of a captured scene from an arbitrary viewpoint from multiple images obtained by capturing images using multiple imaging devices has been attracting attention. In such a virtual viewpoint image, for example, it is assumed that a marker is attached to an object to be focused on in the scene. When a marker is input to a virtual viewpoint image, the marker is displayed at an appropriate position when viewed from the viewpoint when the marker is input, but when switching to another viewpoint, the marker may be displayed at an unintended position. In this way, when drawing additional information such as a marker on a virtual viewpoint image, the drawn additional information may be displayed at an unintended position.

[0005] The present disclosure provides a technique for allowing additional information rendered on a virtual viewpoint image to be displayed at an appropriate position regardless of the viewpoint. [Means for solving the problem]

[0006] According to one aspect of the present disclosure information The processing device The present invention has: first viewpoint information indicating a position of a first virtual viewpoint and a line of sight direction from the first virtual viewpoint; second viewpoint information indicating a position of a second virtual viewpoint different from the first virtual viewpoint and a line of sight direction from the second virtual viewpoint; an acquisition means for acquiring input of additional information for a first virtual viewpoint image generated based on the first viewpoint information and a virtual space; a generation means for generating an additional object having a three-dimensional shape to be placed in the virtual space based on the additional information; and a display control means for controlling the display of the virtual space in which the additional object is placed and a second virtual viewpoint image generated based on the second viewpoint information on a display means, wherein the generation means generates the additional object on a plane parallel to a plane equivalent to the ground of a background object arranged in the virtual space, and the second virtual viewpoint image includes the additional object. 。

Advantages of the Invention

[0007] According to the present disclosure, it is possible to display the additional information drawn on the virtual viewpoint image at an appropriate position regardless of the viewpoint.

Brief Description of the Drawings

[0008] [Figure 1] It is a diagram showing a configuration example of an image processing system. [Diagram 2] It is a diagram showing a configuration example of an image processing apparatus. [Diagram 3] It is a diagram for explaining a virtual viewpoint. [Figure 4] It is a diagram showing a configuration example of an operation screen. [Diagram 5] It is a diagram for explaining a marker object and a target plane. [Figure 6] It is a diagram showing an example of a processing flow executed by an image processing apparatus. [Figure 7] It is a diagram showing an example of a displayed screen. [Figure 8] It is a diagram showing a configuration example of an image processing system. [Figure 9] It is a diagram showing a data configuration example of a marker object. [Figure 10] It is a diagram showing an example of a displayed screen. [Figure 11] It is a diagram showing an example of a displayed screen.

Modes for Carrying Out the Invention

[0009] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the disclosure according to the claims. Although the embodiments describe a plurality of features, not all of these features are essential to the disclosure, and the plurality of features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.

[0010] (System Configuration) An example of the configuration of an image processing system 100 according to this embodiment will be described with reference to Fig. 1(A) and Fig. 1(B). The image processing system 100 includes a plurality of sensor systems (n sensor systems 101-1 to 101-n in the example of Fig. 1(A)). Each sensor system includes at least one imaging device (e.g., a camera). In the following, the sensor systems 101-1 to 101-n will be collectively referred to as "sensor systems 101" unless there is a particular need to distinguish them. In this image processing system 100, virtual viewpoint image data is generated based on image data acquired by the plurality of sensor systems 101 and provided to a user.

[0011] FIG. 1B shows an example of the installation of these sensor systems 101. The multiple sensor systems 101 are installed to surround an area that is the subject of the shooting (hereinafter, referred to as subject area 120), and each of them shoots the subject area 120 from a different direction. For example, if the subject area 120 is defined as a field of a stadium where soccer or rugby matches are played, n (for example, a large number such as 100) sensor systems 101 are installed to surround the field. The number of installed sensor systems 101 is not particularly limited, but is at least a plurality. The sensor systems 101 do not have to be installed all around the subject area 120, and may be installed only in a part of the periphery of the subject area 120 due to, for example, restrictions on the installation location. Furthermore, the imaging devices possessed by each of the multiple sensor systems 101 may include imaging devices with different functions, such as a telephoto camera and a wide-angle camera.

[0012] Furthermore, the sensor system 101 may have a sound collecting device (microphone) in addition to an imaging device (camera). The sound collecting devices in each of the multiple sensor systems 101 collect sound in synchronization. The image processing system 100 generates virtual listening point sound data to be played back together with a virtual viewpoint image based on the sound data collected by each of the multiple sound collecting devices, and provides the generated data to a user. Note that, in the following, a description of sound will be omitted for the sake of simplicity, but it is assumed that images and sound are processed together.

[0013] Note that the subject area 120 is not limited to the field of the stadium, and may be determined to include, for example, the seats of the stadium. Also, the subject area 120 may be determined to be an indoor studio, stage, or the like. That is, the area of ​​the subject for which a virtual viewpoint image is to be generated may be determined as the subject area 120. Note that the "subject" here may be the area itself determined by the subject area 120, and may also include all objects, such as people such as players and referees and balls, that exist within the area. Also, throughout this embodiment, the virtual viewpoint image is considered to be a moving image, but may also be a still image.

[0014] The multiple sensor systems 101 arranged as shown in FIG. 1B synchronously capture images of a common subject area 120 using the imaging devices of the respective sensor systems 101. In this embodiment, an image included in a group of multiple images obtained by synchronously capturing images of the common subject area 120 from multiple viewpoints is called a "multi-view image." Note that the multi-view image in this embodiment may be the captured image itself, but may also be, for example, an image after image processing such as processing to extract a predetermined area from the captured image has been performed.

[0015] The image processing system 100 further includes an image recording device 102, a database 103, and an image processing device 104. The image recording device 102 collects multi-viewpoint images obtained by imaging in each of the multiple sensor systems 101, and stores the other viewpoint images in the database 103 together with the time code used for imaging. Here, the time code is information for uniquely identifying the time when imaging was performed. For example, the time code can be information specifying the imaging time in a format such as day:hour:minute:second.frame number.

[0016] The image processing device 104 acquires a plurality of multi-viewpoint images corresponding to a common time code from the database 103, and generates a three-dimensional model of the subject from the acquired multi-viewpoint images. The three-dimensional model includes, for example, shape information such as a point cloud expressing the shape of the subject, faces and vertices when the shape of the subject is expressed as a collection of polygons, and texture information expressing the color and texture on the surface of the shape. Note that this is only an example, and the three-dimensional model can be defined in any format that expresses the subject in three dimensions. The image processing device 104 generates and outputs a virtual viewpoint image corresponding to a virtual viewpoint specified by, for example, a user using a three-dimensional model of the subject. For example, as shown in FIG. 1B, the virtual viewpoint 110 is specified by the position of the viewpoint and the direction of the line of sight in a virtual space associated with the subject area 120. The user can view the subject generated based on the three-dimensional model of the subject existing in the virtual space from a viewpoint different from any of the imaging devices of the multiple sensor systems 101, for example, by moving the virtual viewpoint in the virtual space and changing the direction of the line of sight. In addition, since the virtual viewpoint can move freely within the three-dimensional virtual space, the virtual viewpoint image is also called a "free viewpoint image."

[0017] The image processing device 104 generates a virtual viewpoint image as an image showing a scene observed from the virtual viewpoint 110. The image generated here is a two-dimensional image. The image processing device 104 is, for example, a computer used by a user, and can be configured to include a display device such as a touch panel display or a liquid crystal display. The image processing device 104 may also have a display control function for displaying an image on an external display device. The image processing device 104 displays the virtual viewpoint image on the screen of the display device, for example. That is, the image processing device 104 executes processing for generating an image of a scene within a range visible from the virtual viewpoint as a virtual viewpoint image and displaying it on the screen.

[0018] The image processing system 100 may have a configuration different from that shown in FIG. 1(A). For example, a configuration including an operation / display device such as a touch panel display may be used separately from the image processing device 104. For example, a configuration may be used in which a virtual viewpoint or the like is operated on a tablet or the like having a touch panel display, a virtual viewpoint image is generated in the image processing device 104 in response to the operation, and the image is displayed on the tablet. A configuration may be used in which a plurality of tablets are connected to the image processing device 104 via a server, and the image processing device 104 outputs a virtual viewpoint image to each of the plurality of tablets. The database 103 and the image processing device 104 may be integrally configured. A configuration may be used in which the image recording device 102 performs processing from a multi-viewpoint image to the generation of a three-dimensional model of the subject, and the three-dimensional model of the subject is stored in the database 103. In that case, the image processing device 104 reads out the three-dimensional model from the database 103 to generate a virtual viewpoint image. 1A shows an example in which a plurality of sensor systems 101 are daisy-chained, but for example, the sensor systems 101 may each be directly connected to the image recording device 102, or may be connected in another manner. Note that, for example, the image recording device 102 or another time synchronization device may be configured to notify each of the sensor systems 101 of reference time information so that the sensor systems 101 can capture images in synchronization.

[0019] In this embodiment, the image processing device 104 further accepts input of a marker such as a circle or a line from the user for the virtual viewpoint image displayed on the screen, and displays the marker on the screen by superimposing it on the virtual viewpoint image. When such a marker is input, the marker is appropriately displayed at the virtual viewpoint to which the marker is input, but by changing the position or orientation of the virtual viewpoint, an unintended display may be performed, such as a misalignment between the object to which the marker is attached and the marker. For this reason, in this embodiment, the image processing device 104 executes a process for displaying the marker received for the displayed two-dimensional screen at an appropriate position without moving the virtual viewpoint. The image processing device 104 converts the two-dimensional marker into a three-dimensional marker object. Then, the image processing device 104 combines the three-dimensional model of the subject with the three-dimensional marker object to generate a virtual viewpoint image in which the position of the marker is appropriately adjusted with the movement of the virtual viewpoint. Below, an example of the configuration of the image processing device 104 that executes such processing and a flow of the processing will be described.

[0020] (Configuration of image processing device) Next, the configuration of the image processing device 104 will be described with reference to Fig. 2(A) and Fig. 2(B). Fig. 2(A) shows an example of the functional configuration of the image processing device 104. The image processing device 104 includes, for example, a virtual viewpoint control unit 201, a model generation unit 202, an image generation unit 203, a marker control unit 204, and a marker management unit 205 as its functional configuration. Note that these are only examples, and at least some of the functions shown may be omitted, or other functions may be added. In addition, as long as the functions described below can be executed, all of the functions shown in Fig. 2(A) may be replaced by other functional blocks. In addition, two or more functional blocks shown in Fig. 2(A) may be combined into one functional block, or one functional block may be divided into multiple functional blocks.

[0021] The virtual viewpoint control unit 201 receives a user operation related to the virtual viewpoint 110 or the time code, and controls the operation of the virtual viewpoint. For the user operation of the virtual viewpoint, a touch panel, a joystick, or the like is used, but is not limited to these, and the user operation can be received by any device. The model generation unit 202 acquires a multi-view image corresponding to a time code designated by a user operation or the like from the database 103, and generates a three-dimensional model showing the three-dimensional shape of the subject included in the subject area 120. For example, the model generation unit 202 acquires a foreground image in which a foreground area corresponding to a subject such as a person or a ball is extracted from the multi-view image, and a background image in which a background area other than the foreground area is extracted. Then, the model generation unit 202 generates a three-dimensional model of the foreground based on the multiple foreground images. The three-dimensional model is composed of a point cloud generated by a shape estimation method such as Visual Hull. Note that the format of the three-dimensional shape data representing the shape of the object is not limited to this, and a mesh or three-dimensional data in a unique format may be used. The model generation unit 202 can generate a three-dimensional model of the background in a similar manner, but the three-dimensional model of the background may be acquired from a model generated in advance by an external device. For convenience, the three-dimensional model of the foreground and the three-dimensional model of the background will hereinafter be collectively referred to as the "three-dimensional model of the subject" or simply as the "three-dimensional model."

[0022] The image generating unit 203 generates a virtual viewpoint image that reproduces a scene as seen from a virtual viewpoint based on a three-dimensional model of a subject and a virtual viewpoint. For example, the image generating unit 203 acquires appropriate pixel values ​​from the multi-viewpoint image for each point constituting the three-dimensional model and performs coloring processing. Then, the image generating unit 203 generates a virtual viewpoint image by arranging the three-dimensional model in a three-dimensional virtual space and projecting and rendering the three-dimensional model together with the pixel values ​​to the virtual viewpoint. Note that the method of generating the virtual viewpoint image is not limited to this, and other methods may be used, such as a method of generating a virtual viewpoint image by projective transformation of a captured image without using a three-dimensional model.

[0023] The marker control unit 204 accepts a marker input such as a circle or a line from the user to the virtual viewpoint image. The marker control unit 204 converts the marker input made to the two-dimensional virtual viewpoint image into a marker object, which is three-dimensional data in a virtual space. The marker control unit 204 transmits an instruction to the image generation unit 203 to generate a virtual viewpoint image combining a three-dimensional model of a subject and a marker object according to the position and posture of the virtual viewpoint. Note that the marker control unit 204 may provide the image generation unit 203 with a marker object as a three-dimensional model, for example, and the image generation unit 203 may generate a virtual viewpoint image by treating the marker object in the same way as a subject. In addition, the image generation unit 203 may execute a process for superimposing the marker object based on the marker object provided by the marker control unit 204, separately from the process for generating the virtual viewpoint image. In addition, the marker control unit 204 may execute a process for superimposing a marker based on the marker object on the virtual viewpoint image provided by the image generation unit 203. The marker management unit 205 performs storage control to store the marker objects of the three-dimensional model converted by the marker control unit 204, for example, in the storage unit 216 described below. The marker management unit 205 performs storage control so that the marker objects are stored, for example, in association with a time code. Note that the model generation unit 202 may calculate the coordinates for each object, such as a person or a ball in the foreground, store the coordinates in the database 103, and use the coordinates for each object to specify the coordinates of the marker object.

[0024] 2B shows an example of the hardware configuration of the image processing device 104. The image processing device 104 includes, as its hardware configuration, for example, a CPU 211, a RAM 212, a ROM 213, an operation unit 214, a display unit 215, a storage unit 216, and an external interface 217. Note that CPU is an abbreviation for Central Processing Unit, RAM is an abbreviation for Random Access Memory, and ROM is an abbreviation for Read Only Memory.

[0025] The CPU 211 controls the entire image processing device 104 and executes the processes described below using programs and data stored in, for example, the RAM 212 and the ROM 213. The CPU 211 executes the programs stored in the RAM 212 and the ROM 213, thereby realizing each functional block in FIG. 2(A). The image processing device 104 may have dedicated hardware such as one or more processors other than the CPU 211, and at least a part of the processes performed by the CPU 211 may be executed by the dedicated hardware. The dedicated hardware may be, for example, an MPU (Micro Processing Unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), and a DSP (Digital Signal Processor). The ROM 213 holds programs and data for executing processes on virtual viewpoint images and markers. The RAM 212 temporarily stores programs and data read from the ROM 213, and also provides a work area for use when the CPU 211 executes each process.

[0026] The operation unit 214 includes a device for accepting an operation by a user, such as a touch panel or a button. The operation unit 214 acquires information indicating an operation by a user on a virtual viewpoint or a marker, for example. The operation unit 214 may be connected to an external controller and accept input information from the user regarding the operation. The external controller is not particularly limited, but may be, for example, a three-axis controller such as a joystick, a keyboard, a mouse, or the like. The display unit 215 includes a display device such as a display. The display unit 215 displays, for example, a virtual viewpoint image generated by the CPU 211 or the like. The display unit 215 may include various output devices capable of presenting information to a user, such as a speaker for audio output and a device for vibration output. The operation unit 214 and the display unit 215 may be integrally configured, for example, by a touch panel display.

[0027] The storage unit 216 includes a large-capacity storage device such as a solid state drive (SSD) or a hard disk drive (HDD). Note that these are only examples, and the storage unit 216 may include any other storage device. and records data processed by a program. The storage unit 216 stores, for example, a three-dimensional marker object obtained by converting a marker input accepted via the operation unit 214 by the CPU 211. The storage unit 216 may further store other information. The external interface 217 includes, for example, an interface device that connects to a network such as a local area network (LAN). Information is transmitted and received between the external device such as the database 103 and the like via the external interface 217. The external interface 217 may also include an image output port such as HDMI (registered trademark) or SDI. Note that HDMI is an abbreviation for High-Definition Multimedia Interface, and SDI is an abbreviation for Serial Digital Interface. In this case, information can be transmitted to an external display device or projection device via the external interface 217. Also, the external interface 217 may be used to connect to a network, and operation information for the virtual viewpoint and markers may be received and virtual viewpoint images may be transmitted via the network.

[0028] (Virtual viewpoint and line of sight) Next, the virtual viewpoint 110 will be described with reference to FIG. 3(A) to FIG. 3(D). The virtual viewpoint 110 and its movement are specified using one coordinate system. In this embodiment, a typical three-dimensional Cartesian coordinate system consisting of an X axis, a Y axis, and a Z axis as shown in FIG. 3(A) is used as the coordinate system. Note that this is only an example, and any coordinate system capable of indicating a position in a three-dimensional space may be used. The coordinates of the subject are set and used using this coordinate system. The subject includes, for example, a field or studio of a stadium, and people and objects such as a ball that exist in the space of the field or studio. For example, in the example of FIG. 3(B), the subject includes the entire field 391 of the stadium, a ball 392 and players 393 that exist thereon. Note that the subject may include spectator seats around the field. 3(B), the coordinates of the center of field 391 are set as the origin (0,0,0), the X axis is the long side direction of field 391, the Y axis is the short side direction of field 391, and the Z axis is the vertical direction to the field. In this way, by setting the coordinates of each subject with reference to the center of field 391, a three-dimensional model generated from the subject can be placed in a three-dimensional virtual space. Note that the method of setting the coordinates is not limited to this.

[0029] Next, the virtual viewpoint will be described with reference to FIG. 3(C) and FIG. 3(D). The virtual viewpoint determines the viewpoint and line of sight direction for generating a virtual viewpoint image. In FIG. 3(C), the apex of a pyramid indicates the position 301 of the virtual viewpoint, and the vector extending from the apex indicates the line of sight direction 302. The position 301 of the virtual viewpoint is expressed by coordinates (x, y, z) in a three-dimensional virtual space, and the line of sight direction 302 is expressed by a unit vector with the components of each axis as scalars, and is also called the optical axis vector of the virtual viewpoint. The line of sight direction 302 passes through the center points of the front clip plane 303 and the rear clip plane 304. The clip plane is a plane that defines the area to be drawn. The space 305 sandwiched between the front clip plane 303 and the rear clip plane 304 is called the view frustum of the virtual viewpoint, and the virtual viewpoint image is generated within this range (or the virtual viewpoint image is projected and displayed within this range). The focal length (not shown) can be set to any value, and the angle of view can be changed by changing the focal length, as in a typical camera. That is, by shortening the focal length, the angle of view can be widened and the viewing frustum can be widened. On the other hand, by lengthening the focal length, the angle of view can be narrowed and the viewing frustum can be narrowed, allowing the subject to be captured larger. Note that the focal length is just one example, and any parameter capable of setting the position and size of the viewing frustum can be used.

[0030] The position of the virtual viewpoint and the line of sight from the virtual viewpoint can be moved and rotated in a virtual space expressed by three-dimensional coordinates. As shown in FIG. 3(D), the movement 306 of the virtual viewpoint is the movement of the position 301 of the virtual viewpoint, and is expressed by the components of each axis (x, y, z). As shown in FIG. 3(A), the rotation 307 of the virtual viewpoint is expressed by Yaw, which is a rotation around the Z axis, Pitch, which is a rotation around the X axis, and Roll, which is a rotation around the Y axis. With these, the position of the virtual viewpoint and the line of sight from the virtual viewpoint can be freely moved and rotated in the three-dimensional virtual space, and the image processing device 104 can reproduce an image in which an arbitrary region of the subject is assumed to be observed from an arbitrary angle as a virtual viewpoint image. In the following, when there is no particular need to distinguish between them, the position of the virtual viewpoint and the line of sight from the virtual viewpoint are collectively referred to as the "virtual viewpoint".

[0031] (How to operate the virtual viewpoint and markers) A method of operating a virtual viewpoint and a marker will be described with reference to FIG. 4(A) and FIG. 4(B). FIG. 4(A) is a diagram for explaining a screen displayed by the image processing device 104. Here, a case where a tablet-type terminal 400 having a touch panel display is used will be described. The terminal 400 does not need to be a tablet-type terminal, and any other type of information processing device can be used as the terminal 400. When the terminal 400 is the image processing device 104, the terminal 400 is configured to generate and display a virtual viewpoint image, and to accept operations such as designation of a virtual viewpoint or a time code and marker input. On the other hand, when the terminal 400 is a device connected to the image processing device 104 via a communication network, for example, the terminal 400 transmits information designating a virtual viewpoint or a time code to the image processing device 104 and receives the provision of a virtual viewpoint image. In addition, the terminal 400 accepts a marker input operation for the virtual viewpoint image, and transmits information indicating the accepted marker input to the image processing device 104.

[0032] In FIG. 4A, a display screen 401 of a terminal 400 is roughly divided into two areas: a virtual viewpoint operation area 402 and a time code operation area 403.

[0033] In the virtual viewpoint operation area 402, a user operation regarding the virtual viewpoint is accepted, and a virtual viewpoint image is displayed within the range of the area. That is, in the virtual viewpoint operation area 402, a virtual viewpoint image that reproduces a scene when it is assumed that the virtual viewpoint is operated and observed from the virtual viewpoint after the operation is displayed. In addition, in the virtual viewpoint operation area 402, a marker input to the virtual viewpoint image is accepted. Note that the marker operation and the virtual viewpoint operation may be performed together, but in this embodiment, the marker operation is accepted independently of the virtual viewpoint operation. In one example, as in the example of FIG. 4(B), the virtual viewpoint is operated by a touch operation such as tapping and dragging with the user's finger on the terminal 400, and the marker operation may be performed by tapping and dragging with a drawing device such as a pencil 450. For example, the user moves or rotates the virtual viewpoint by a drag operation 431 with the finger. In addition, the user draws a marker 451 or a marker 452 on the virtual viewpoint image by a drag operation with the pencil 450. The terminal 400 draws a marker at the successive coordinates of the drag operation by the pencil 450. Note that the operation by the finger may be assigned to the marker operation, and the operation by the pencil may be assigned to the virtual viewpoint operation. In addition, any other operation method may be used as long as the terminal 400 can distinguish between the virtual viewpoint operation and the marker operation. With such a configuration, the user can easily use the virtual viewpoint operation and the marker operation separately.

[0034] It should be noted that when the virtual viewpoint operation and the marker operation are performed independently, a drawing device such as the pencil 450 does not necessarily have to be used. For example, an ON / OFF button (not shown) for the marker operation may be provided on the touch panel, and whether or not to perform the marker operation may be switched by operating the button. For example, when the marker operation is performed, the button may be turned ON, and while the button is turned ON, the virtual viewpoint operation may not be performed. Also, when the virtual viewpoint operation is performed, the button may be turned OFF, and while the button is turned OFF, the marker operation may not be performed.

[0035] The time code operation area 403 is used to specify the timing of the virtual viewpoint image to be viewed. The time code operation area 403 includes, for example, a main slider 412, a sub-slider 413, a speed designation slider 414, and a cancel button 415. The main slider 412 is used to accept any time code selected by the user by dragging the position of a knob 422, etc. The range of the main slider 412 indicates the entire period during which the virtual viewpoint image can be played back. The sub-slider 413 enlarges and displays a part of the time code, enabling the user to perform detailed operation of the time code, such as on a frame-by-frame basis. The sub-slider 413 accepts the user's selection of any detailed time code by dragging the position of a knob 423, etc.

[0036] In the terminal 400, the main slider 412 accepts a rough designation of the time code, and the sub-slider 413 accepts a detailed designation of the time code. For example, the main slider 412 and the sub-slider 413 can be set so that the main slider 412 corresponds to a range of three hours that corresponds to the length of the entire game, and the sub-slider 413 corresponds to a time range of about 30 seconds that is a part of that range. For example, a section of 15 seconds before and after the time code designated by the main slider 412 or a section of 30 seconds from that time code can be expressed by the sub-slider 413. Also, sections may be divided in units of 30 seconds in advance, and the section including the time code designated by the main slider 412 among those sections may be expressed by the sub-slider 413. In this way, the time scales of the main slider 412 and the sub-slider 413 are different. Note that the above-mentioned time length is one example, and the sub-slider 413 may be configured to correspond to another time length. Note that, for example, a user interface may be provided that allows the setting of the time length corresponding to the sub-slider 413 to be changed. Also, while FIG. 4A shows an example in which the main slider 412 and the sub-slider 413 are displayed on the screen with the same length, their lengths may be different from each other. That is, the main slider 412 may be longer, or the sub-slider 413 may be longer. Also, the sub-slider 413 does not have to be displayed all the time. For example, the sub-slider 413 may be displayed after a display instruction is accepted, or the sub-slider 413 may be displayed when a specific operation such as pause is instructed. Also, the time code may be specified and displayed without using the knob 422 of the main slider 421 and the knob 423 of the sub-slider 413. For example, the time code may be specified and displayed by numerical values, such as numerical values ​​in the format of day:hour:minute:second.frame number.

[0037] The speed designation slider 414 is used to accept a user's designation of a playback speed such as normal playback or slow playback. For example, the count-up interval of the time code is controlled according to the playback speed selected using the knob 424 of the speed designation slider 414. The cancel button 415 can be used to cancel each operation related to the time code. The cancel button 415 may also be used to clear a pause and return to normal playback. Note that the button is not limited to cancel as long as it is a button for performing an operation related to the time code.

[0038] Using the above-mentioned screen configuration, the user can operate the virtual viewpoint and the time code to display a virtual viewpoint image of a three-dimensional model of a subject at an arbitrary time code as viewed from an arbitrary position and posture on the terminal 400. The user can then perform marker input to such a virtual viewpoint image, independently of the operation of the virtual viewpoint.

[0039] (marker object and surface of interest) In this embodiment, a two-dimensional marker input by a user to a virtual viewpoint image displayed in two dimensions is converted into a three-dimensional marker object. The marker object converted into three dimensions is placed in the same three-dimensional virtual space as the virtual viewpoint. First, a method of converting the marker object will be described with reference to Figs. 5(A) to 5(E).

[0040] FIG. 5(A) shows a state in which a marker is input to a virtual viewpoint image 500 using a pencil 450. Here, a vector representing one point on the marker is called a marker input vector 501. FIG. 5(B) is a diagram showing a schematic diagram of a virtual viewpoint specified when the virtual viewpoint image 500 displayed in FIG. 5(A) is generated. If the marker input vector 501 in FIG. 5(A) is expressed as [M c ], then this marker input vector 501 can be expressed as [M c]=(a,b,f), where “f” is the focal length of the virtual viewpoint. In FIG. 5B, the intersection point between the optical axis of the virtual camera and the focal plane of the virtual viewpoint is (0,0,f), and the marker input vector [M c ]=(a,b,f) refers to a point that is a distance a in the x direction and b in the y direction from the intersection point. For the marker input vector, [M c ] and in world coordinates, [M w ]=(m x ,m y ,m z ) The world coordinates are coordinates in the virtual space. The marker input vector 502 in the world coordinates ([M w ]) is calculated by using the quaternion Qt obtained from the rotation matrix that indicates the orientation of the virtual camera, w ]=Qt·[M c Since the quaternion of the virtual camera is a general technical term, it will not be described in detail here.

[0041] In this embodiment, in order to convert a marker input to a two-dimensional virtual viewpoint image into a marker object of three-dimensional data, a plane of interest 510 as shown in Fig. 5(C) and Fig. 5(D) is used. Fig. 5(C) is a diagram overlooking the entire plane of interest 510, and Fig. 5(D) is a diagram viewed from a direction parallel to the plane of interest 510. In other words, Fig. 5(C) and Fig. 5(D) show the same subject viewed from different directions. The plane of interest 510 is, for example, a plane parallel to the field 391 (Z=0) and at a height (e.g. Z=1.5 m) that is easy for a viewer to notice. The plane of interest 510 is located at a position where z=z fix It can be expressed by the following formula:

[0042] The marker object is generated as three-dimensional data tangent to the plane of interest 510. For example, intersection 503 of marker input vector 502 and plane of interest 510 is generated as a marker object corresponding to a point on the marker corresponding to marker input vector 502. In other words, conversion from the marker to the marker object is performed so that the intersection of a straight line passing through the virtual viewpoint and a point on the marker, and a predetermined surface prepared as plane of interest 510, becomes a point in the marker object corresponding to the point on the marker. Note that marker input vector 502 ([M w ]=(m x ,m y ,m z )) and plane 510 (z=z fix ) can be calculated by a general mathematical solution, so a detailed explanation will be omitted here. Here, this intersection is expressed by the intersection coordinate A w =(a x ,a y ,a z )). Such intersection coordinates are calculated for each of the consecutive points obtained as the marker input, and the three-dimensional data obtained by linking these points becomes the marker object 520. That is, when the marker is input as a continuous straight line or curve, the marker object 520 is generated as a straight line or curve on the attention plane 510 corresponding to the continuous straight line or curve. As a result, the marker object 520 is generated as a three-dimensional object that is in contact with the attention plane that is easy for the viewer to notice. The attention plane can be set based on the height of the object that the viewer notices, for example, based on the height where the ball mainly exists or the height of the center part of the player. Note that the conversion from the marker to the marker object 520 may be performed by calculating the intersection between the vector and the attention plane as described above, or the conversion from the marker to the marker object may be performed by a predetermined matrix operation or the like. Also, the marker object may be generated based on a table corresponding to the virtual viewpoint and the position where the marker is input. Note that regardless of which method is used, processing may be performed so that the marker object is generated on the attention plane.

[0043] In one example, the above-mentioned conversion is performed on a marker input as a two-dimensional circle on the tablet shown in Fig. 5(A), thereby generating a three-dimensional doughnut-shaped marker object 520 shown in Fig. 5(E). The marker object is created from continuous points, but is generated as three-dimensional data with a predetermined height, as shown in Fig. 5(E).

[0044] In the example described with reference to FIG. 5(A) to FIG. 5(E), the attention plane 510 is a plane parallel to the XY plane, which is the field, but is not limited thereto. For example, in the case where the subject is a sport performed in a vertical direction, such as bouldering, the attention plane may be set as a plane parallel to the XZ plane or the YZ plane. In other words, any plane that can be defined in a three-dimensional space can be used as the attention plane. The attention plane is not limited to a plane, and may be a curved surface. In other words, even if it is a curved surface, the shape is not particularly limited as long as the intersection of the marker input vector with the curved surface can be uniquely calculated.

[0045] (Processing flow) An example of the flow of processing executed by the image processing device 104 will be described with reference to FIG. 6. This processing is configured by loop processing (S601, S612) that repeats the processing of S602 to S611, and this loop is executed at a predetermined frame rate. For example, when the frame rate is 60 FPS, one loop (one frame) of processing is executed at intervals of about 16.6 [ms]. As a result, in S611 described later, a virtual viewpoint image is output at that frame rate. The frame rate can be set so as to be synchronized with the update rate of the screen display of the image processing device 104, but may also be set to match the frame rate of the imaging device that captured the multi-view image or the frame rate of the three-dimensional model stored in the database 103. In the following description, it is assumed that the time code is counted up by one frame each time the loop processing is executed, but the interval of the count-up of the time code may be changed according to a user operation or the like. For example, when a 1 / 2 playback speed is specified, one frame of the time code may be counted up for two loop processings. Also, for example, when a pause is specified, the count-up of the time code can be stopped.

[0046] In the loop process, the image processing device 104 updates the time code of the processing target (S602). The time code here is expressed in the format of day:hour:minute:second.frame as described above, and is updated by counting up in units of frames. Then, the image processing device 104 determines whether the accepted user operation is a virtual viewpoint operation or a marker input operation (S603). Note that the type of operation is not limited to these. For example, when the image processing device 104 accepts an operation on the time code, the image processing device 104 may return the process to S602 and update the time code. Also, when the image processing device 104 does not accept a user operation, the image processing device 104 may proceed with the process assuming that a virtual viewpoint designation operation has been performed during the immediately preceding virtual viewpoint image generation process. Also, the image processing device 104 may determine that another operation has been accepted.

[0047] When the image processing device 104 determines in S603 that the virtual viewpoint operation has been received, it acquires two-dimensional operation coordinates for the virtual viewpoint (S604). The two-dimensional operation coordinates here are, for example, coordinates indicating the position where the tap operation on the touch panel has been received. Then, the image processing device 104 performs at least one of moving and rotating the virtual viewpoint in the three-dimensional virtual space based on the operation coordinates acquired in S604 (S605). The movement and rotation of the virtual viewpoint have been described with reference to FIG. 3(D), and will not be repeated here. In addition, the process of determining the amount of movement and rotation of the virtual viewpoint in the three-dimensional space from the two-dimensional coordinates obtained by the touch operation on the touch panel can be performed using a well-known technique, and will not be described in detail here. After or in parallel with the processes of S604 and S605, the image processing device 104 determines whether or not a marker object exists within the range of the field of view determined by the virtual viewpoint after the movement and rotation (S606). Then, when a marker object exists within the range of the field of view determined by the virtual viewpoint after the movement and rotation (YES in S606), the image processing device 104 reads out the marker object and places it in the three-dimensional virtual space (S607). After placing the marker object, the image processing device 104 generates and outputs a virtual viewpoint image including the marker object (S611). That is, the image processing device 104 generates a virtual viewpoint image using the three-dimensional model of the subject corresponding to the time code updated in S602 and the marker object placed in the virtual space according to the virtual viewpoint. Note that, when a marker object does not exist within the range of the field of view determined by the virtual viewpoint after the movement and rotation (NO in S606), the image processing device 104 generates and outputs a virtual viewpoint image without placing the marker object (S611).

[0048] When it is determined in S603 that the marker input operation has been received, the image processing device 104 acquires two-dimensional coordinates of the marker operation in the virtual viewpoint image (S608). That is, the image processing device 104 acquires two-dimensional coordinates of the marker operation input to the virtual viewpoint operation area 402 as described with reference to FIG. 4(A). Then, the image processing device 104 converts the two-dimensional coordinates of the marker input acquired in S608 into a marker object, which is three-dimensional data in a target plane in the virtual space (S609). The method of converting the marker input into a marker object has been described with reference to FIGS. 5(A) to 5(E), and therefore will not be described again here. When the image processing device 104 acquires the marker object in S609, the image processing device 104 holds the marker object together with a time code (S610). Then, the image processing device 104 generates a virtual viewpoint image using a three-dimensional model of the subject corresponding to the time code updated in S602 and the marker object arranged in the virtual space according to the virtual viewpoint (S611).

[0049] According to the above processing, the markers input to the two-dimensional virtual viewpoint image can be converted into three-dimensional data using the target plane in the virtual space, and the three-dimensional model of the subject and the marker object can be arranged in the same virtual space. Therefore, it is possible to generate a virtual viewpoint image in which the positional relationship between the subject and the marker is maintained regardless of the position and orientation of the virtual viewpoint.

[0050] (Example of screen display) Using FIG. 7(A) to FIG. 7(D), examples of screen display in the case where the above-mentioned processing is performed and the case where it is not performed will be described. FIG. 7(A) shows a state where the image processing device 104 displays a virtual viewpoint image before a marker is input. In this state, for example, by a touch operation 701, the user can arbitrarily move and rotate the virtual viewpoint. Here, as an example, the time code is paused by a user operation, and a marker input can be accepted in this state. Note that the marker input may correspond to a virtual viewpoint image for a certain period of time, and for example, the marker input may be maintained for a predetermined number of frames from the frame in which the operation is accepted. Also, for example, the input marker may be maintained until an explicit instruction to erase the marker is given from the user. FIG. 7(B) shows a state in which a marker operation by a user is being accepted. Here, an example is shown in which two marker inputs (a circle 711 and a curve 712) are accepted by a pencil. When such marker inputs are accepted, the image processing device 104 converts the marker inputs into 3D model marker objects on the surface of interest, and draws the marker objects together with the virtual viewpoint image. When the virtual viewpoint is not changed, the display is the same as when the user inputs a marker, and the user may recognize that the marker has simply been input. While the conversion process to the marker object is being performed, the accepted 2D marker inputs may be displayed as is, and redrawing based on the 3D marker object may be performed in response to the completion of the conversion process to the marker object.

[0051] FIG. 7(C) shows an example in which the input marker is converted into a three-dimensional marker object and drawn, and then the virtual viewpoint is moved and rotated by the user's touch operation 721. According to the method of this embodiment, a three-dimensional marker object is placed in the vicinity of the three-dimensional position of the subject to be focused on, to which the marker is attached, in the virtual space, so that the marker is displayed near the subject to be focused on even after the virtual viewpoint is moved and rotated. Also, since the marker is converted into a three-dimensional marker object and drawn in the virtual viewpoint image, the position and orientation of the marker are changed and observed according to the movement and rotation of the virtual viewpoint. Thus, according to the method of this embodiment, the consistency between the input content of the marker and the content of the virtual viewpoint image can be maintained even after the virtual viewpoint is moved and rotated. On the other hand, if the method of this embodiment is not applied, after the virtual viewpoint is moved and rotated, the marker is displayed as it is at the input position, and only the content of the virtual viewpoint image changes, so that the input content of the marker and the content of the virtual viewpoint image become inconsistent.

[0052] Thus, according to this embodiment, the marker input to the virtual viewpoint image is generated as a marker object on a surface that is easily noticed by the viewer in the virtual space where the three-dimensional model of the subject is placed. Then, this marker object is rendered as the virtual viewpoint image together with the three-dimensional model of the subject. This maintains the positional relationship between the subject and the marker, and can eliminate or reduce the sense of incongruity caused by the marker being displaced from the subject when the virtual viewpoint is moved to an arbitrary position and orientation.

[0053] (Sharing marker objects) The marker object generated by the above-mentioned method can be shared among other devices. FIG. 8 shows an example of the configuration of a system that performs such sharing. In FIG. 8, the virtual viewpoint image generating system 100 is configured to be able to execute at least a part of the functions described with reference to FIG. 1(A), and generates a multi-viewpoint image and a three-dimensional model from photographing a subject. The management server 801 is a storage device that manages and stores the three-dimensional model of the subject and the shared marker object described later for each time code, and can also be a distribution device that distributes the marker object for each time code. Each of the image processing devices 811 to 813 has a function similar to that of the image processing device 104, for example, and acquires a three-dimensional model from the management server 801 to generate and display a virtual viewpoint image. Each of the image processing devices 811 to 813 can also receive a marker input from a user to generate a marker object, and upload the generated marker object to the management server 801 for storage. By downloading the marker object uploaded by the second image processing device, the first image processing device can generate a virtual viewpoint image that combines the marker input by the user of the second image processing device with a three-dimensional model of the subject.

[0054] (Data structure of marker object) The image processing device 104 holds the marker object in the configuration shown in Fig. 9(A) to Fig. 9(C) by the marker management unit 205, for example. Then, the image processing device 104 may upload the data of the marker object to the management server 801 with the same data configuration. Also, the image processing device 104 downloads the data of the marker object with the same data configuration from the management server 801. Note that the data configurations shown in Fig. 9(A) to Fig. 9(C) are only examples, and data in any format capable of identifying the position and shape of the marker object in the virtual space may be used. Also, different data formats may be used between the management server 801 and the image processing device 104, and the data of the marker object may be reproduced as necessary according to a predetermined rule.

[0055] 9(A) shows an example of the configuration of marker object data 900. The type of object is stored in a header 901 at the beginning of a data set. For example, information indicating that this data set is a "marker" data set is stored in the header 901. The type of data set may be specified as "foreground" or "background". The marker object data 900 also includes a frame number 902 and one or more pieces of data 903 corresponding to each of the one or more frames (time codes).

[0056] The data 903 for each frame (time code) has a configuration as shown in FIG. 9B, for example. The data 903 includes, for example, a time code 911 and a data size 912. The data size 912 allows the boundary between the data of this frame (time code) and the data of the next frame to be specified, and the data size for each of the multiple frames may be prepared outside the data 903. In this case, for example, information on the data size may be stored between the header 901 and the frame number 902, between the frame number 902 and the data 903, after the data 903, etc. The data 903 further includes a marker object number 913. The marker object number 913 indicates the number of markers included in the time code. For example, in the example of FIG. 7B, two markers, a circle 711 and a curve 712, are input. Therefore, in this case, the marker object number 913 indicates that two marker objects exist. The data 903 includes data 914 of the number of marker objects indicated by the number of marker objects 913 .

[0057] The data 914 for each marker object has a configuration as shown in FIG. 9C, for example. The data 914 for each marker object includes a data size 921 and a data type 922. The data size 921 indicates the size of each of the one or more pieces of data 914, and is used to enable identification of the boundary of the data for each marker object. The data type 922 indicates the type of shape of the three-dimensional model. The data type 922 specifies information such as "point cloud" or "mesh". Note that the shape of the three-dimensional model is not limited to these. As an example, when information indicating "point cloud" is stored in the data type 922, the data 914 includes the number of point clouds 922, and one or more combinations of coordinates 924 of all point clouds and texture 934. Note that in each of the one or more pieces of data 914, in addition to the three-dimensional data of the marker object, the center coordinates of all point cloud coordinates, minimum and maximum values ​​in each three-dimensional axis, and the like may be stored (not shown). In addition, in each of the one or more pieces of data 914, further other data may be stored. The data of the marker object does not have to be prepared on a frame-by-frame basis, and may be configured as an animation in a format such as a scene graph.

[0058] 9(A) to 9(C) can be used to manage multiple marker objects at different time codes in a single sporting event, for example, and can also manage multiple markers for one time code.

[0059] Next, the management of marker objects by the management server 801 will be described. FIG. 9(D) shows an example of the configuration of a table for managing marker objects. Time codes are arranged on the horizontal axis of the table, and information indicating marker objects corresponding to each time code is stored on the vertical axis. For example, data of "1st object" indicated in each of data 903 of a plurality of time codes (frames) is stored in a cell of a row corresponding to "1st object". For example, data shown in FIG. 9(C) is stored in one cell. By using the table configuration of the database as shown in FIG. 9(D), when a marker object is specified, the management server 801 can obtain data of all time codes related to the marker object. Furthermore, when a range of time codes is specified, the management server 801 can obtain data of all marker objects included in the time codes. Note that these are merely examples, and the above-mentioned data configuration and management method do not have to be used as long as the marker object can be managed and shared.

[0060] (Sharing and displaying marker objects) An example of sharing and displaying a marker object will be described with reference to Figs. 10(A) to 10(C). Here, a case where a marker object is shared among the image processing device 811 to the image processing device 813 shown in Fig. 8 will be described. Here, Fig. 10(A) shows an operation screen displayed on the image processing device 811, Fig. 10(B) shows an operation screen displayed on the image processing device 812, and Fig. 10(C) shows an operation screen displayed on the image processing device 813. As an example, as shown in Fig. 10(A), it is assumed that a marker in the shape of a circle 1001 is accepted for a virtual viewpoint image of a certain time code in the image processing device 811. For example, based on accepting a marker sharing instruction (not shown) from a user, the image processing device 811 transmits data of a marker object converted from the input marker as described above to the management server 801. Also, as shown in Fig. 10(B), it is assumed that a marker in the shape of a curve 1002 is accepted for a virtual viewpoint image of the same time code as that in Fig. 10(A) in the image processing device 812. Upon receiving a marker sharing instruction (not shown) from a user, the image processing device 812 transmits a marker object converted from the input marker as described above to the management server 801. The marker object is associated with a time code and transmitted to the management server 801 for storage. For this reason, the image processing devices 811 to 813 each have a storage control function for associating the time code with the marker object and storing them in the management server 801.

[0061] In this state, it is assumed that the image processing device 813 receives an instruction to update the marker (not shown). In this case, the image processing device 813 acquires a marker object corresponding to the marker from the management server 801 with respect to the time code at which the image processing device 811 and the image processing device 812 input the marker as shown in Fig. 10(A) and Fig. 10(B). Then, the image processing device 813 generates a virtual viewpoint image using the three-dimensional model of the subject of the time code and the acquired marker object as shown in Fig. 10(C). Note that the marker object acquired here is three-dimensional data arranged in the same virtual space as the three-dimensional model of the subject as described above. Therefore, in the image processing device 813 of the sharing destination, the three-dimensional model of the subject and the marker object are arranged in the same positional relationship as the positional relationship at the time of input in the image processing device 811 and the image processing device 812 of the sharing source. Therefore, even if the position of the virtual viewpoint and the direction of the line of sight are arbitrarily operated in the image processing device 813, the virtual viewpoint image is generated while maintaining those positional relationships.

[0062] For example, when the image processing device 811 receives a marker update instruction, the image processing device 811 may acquire information on the marker object of the marker (curve 1002) input in the image processing device 812. Then, the image processing device 811 may draw a virtual viewpoint image including the acquired marker object in addition to the marker object of the marker (circle 1001) input in the image processing device 811 itself and a three-dimensional model of the subject.

[0063] For example, when the image processing device 813 receives an instruction to update a marker, the marker objects managed in the management server 801 may be displayed as a list using thumbnails or the like, and a marker object to be downloaded may be selected. In this case, when the number of marker objects managed in the management server 801 is large, the priority of the marker object near the time code where the update instruction was received may be increased and the marker objects may be displayed as a list using thumbnails or the like. That is, when the image processing device 813 receives an instruction to update a marker, the marker input near the time code corresponding to that time may be preferentially displayed so that the user can easily select the marker. In addition, the number of times a download instruction has been received may be managed for each marker object, and the priority of the marker object with the highest number of download instructions may be increased and the marker object may be displayed as a list using thumbnails or the like. The higher the display priority, the closer to the top of the list the thumbnail or the like may be displayed, or the thumbnail may be displayed in a larger size.

[0064] In this way, the marker input to the virtual viewpoint image can be shared among multiple devices while maintaining the positional relationship between the subject and the marker. At this time, since the marker has the form of a three-dimensional object in the virtual space as described above, even if the position of the virtual viewpoint or the direction of the line of sight is different between the sharing source device and the sharing destination device, the virtual viewpoint image can be generated while maintaining the positional relationship between the subject and the marker.

[0065] (Controlling marker transparency according to time code) The marker object displayed as described above will lose its positional relationship with the subject when the subject moves over time. In such a case, if the marker object is left displayed, it may give the user a sense of incongruity. In response to this, by increasing the transparency of the marker object as time passes from the time code at which the marker object is input, the user can be informed of the state of change from the original point of interest and the sense of incongruity can be reduced. As described above, the marker object is stored in association with the time code. Here, in the example of FIG. 10(C), it is assumed that the user executes an operation 751 on the time code. In this case, as shown in FIG. 11(A), the image processing device 813 may change the transparency (α value) of the marker as the marker object transitions to a time code farther away from the time code at which it was generated. As a result, the marker is drawn so that the marker is displayed lighter as the timing at which the marker is actually input is shifted, making it possible to reduce the sense of incongruity felt by the user. Note that when the user returns the time code, control may be performed to return the transparency to the original transparency as shown in FIG. 10(C). The marker object may also be drawn in a virtual viewpoint image corresponding to a time code earlier than the time code at which the marker object was generated. In this case, too, control may be performed so that the transparency is increased (displayed lighter) as the difference between the time code at which the marker object was generated and the time code of the virtual viewpoint image to be drawn increases. Control may also be performed so that the transparency before the time code at which the marker object was generated is higher (displayed lighter) than the transparency after the time code at which the marker object was generated.

[0066] In the above example, the transparency changes as the time code transitions, but the present invention is not limited to this. For example, the marker may be displayed more faintly by changing at least one of the brightness, saturation, and chromaticity of the marker object.

[0067] (Marker object control according to the coordinates of a 3D model) In addition to or instead of the above-described embodiment, the image processing device may receive an operation to attach a marker to a three-dimensional model in the foreground. For example, as shown in FIG. 11(B), a marker addition instruction 1101 is received for a three-dimensional model in the foreground (person). In this case, the image processing device generates a circular marker object 1102 having a predetermined radius centered on the XY coordinates in the target plane of the three-dimensional model, as shown in FIG. 11(C), for example. This makes it possible to generate a marker object similar to the marker object described using FIG. 5(A) to FIG. 5(E). Note that the shape of the marker object generated when an instruction to add a marker object (attach a marker) to a three-dimensional model in the foreground is received is not limited to a circle. The shape of the marker object in this case may be, for example, a rectangle or the like, or may be another shape that can attract attention to the three-dimensional model.

[0068] The position of the marker attached to the foreground three-dimensional model may be changed according to the movement of the person, ball, etc. in the foreground. In the marker object described in FIG. 5(A) to FIG. 5(E), the position of the marker object may be changed according to the position of the foreground three-dimensional model by associating the marker object with the foreground three-dimensional model at the position where the marker is attached or around the position. That is, regardless of the method of attaching the marker, the coordinates of the foreground three-dimensional model for each time code may be acquired, and the coordinates of the marker object may be changed for each time code. In this way, when the person, etc. in the foreground moves according to the passage of the time code, the marker object can be moved following the three-dimensional model. Also, for example, the position of the marker object may be changed according to the change in the time code by a user operation. In this case, the position of the marker object may be specified for each time code according to the user operation, and the position may be stored and managed.

[0069] In the above embodiment, the process when a marker is attached to a virtual viewpoint image based on a multi-viewpoint image captured by a plurality of imaging devices is described, but the present invention is not limited to this. That is, even when a marker is attached to a virtual viewpoint image generated based on a three-dimensional virtual space that is entirely artificially created on a computer, the marker may be converted into a three-dimensional object in the virtual space. In addition, in the above embodiment, an example in which a marker object associated with a time code corresponding to a virtual viewpoint image to which a marker is attached is generated and stored is described, but the time code may not be associated with the marker object. For example, when the virtual viewpoint image is a still image or when the virtual viewpoint image is used only for temporarily attaching a marker in a conference or the like, the marker object may be displayed or erased by, for example, a user's operation regardless of the time code. In addition, at least a part of the image processing device used in, for example, a conference system may not have the ability to specify a virtual viewpoint. That is, after a marker is attached to a virtual viewpoint image, it is sufficient that only a specific user, such as a person who is in charge of proceeding with the conference, can specify a virtual viewpoint, and image processing devices owned by other users may not accept operations of the virtual viewpoint. In this case as well, since the marker object is drawn according to the virtual viewpoint specified by a specific user, it is possible to prevent the relationship between the marker and the subject of the virtual viewpoint image from becoming inconsistent.

[0070] [Other embodiments] In the above-mentioned embodiment, an example of displaying a marker object as additional information to be displayed on the virtual viewpoint image has been described. However, the additional information to be displayed on the virtual viewpoint image is not limited to this. For example, at least any of additional information of a marker, an icon, an avatar, an illustration, etc. designated by a user may be configured to be displayed on the virtual viewpoint image. In addition, a plurality of such additional information may be prepared in advance, and an arbitrary one may be selected by the user and arranged on the virtual viewpoint image. In addition, a configuration may be made such that a user can drag an icon, etc. on a touch panel display and arrange it at an arbitrary position. The arranged additional information such as an icon is converted into three-dimensional data by a method similar to that of the above-mentioned embodiment. Note that this is not limited to this method, and the additional information of two-dimensional data and the additional information of three-dimensional data may be associated in advance, and the additional information of two-dimensional data may be replaced with the corresponding additional information of three-dimensional data at the timing when the additional information of two-dimensional data is arranged. In this way, this embodiment is applicable to the case where various additional information is displayed on a virtual viewpoint image.

[0071] In the above embodiment, the additional information (marker object) is converted into three-dimensional data, but the three-dimensional data does not necessarily have to represent a three-dimensional shape. In other words, the three-dimensional data is data having at least a three-dimensional position in a virtual space, and the shape of the additional information may be a plane, a line, or a point.

[0072] Furthermore, it is not necessary to have all of the functions described in the above embodiments, and any combination of functions may be used.

[0073] The present disclosure can also be realized by a process in which a program for implementing one or more functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more functions.

[0074] The present disclosure is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the present disclosure. [Explanation of symbols]

[0075] 104: image processing device, 201: virtual viewpoint control unit, 202: model generation unit, 203: image generation unit, 204: marker control unit, 205: marker management unit

Claims

Claims 1. First viewpoint information indicating the position of a first virtual viewpoint and the line-of-sight direction from the first virtual viewpoint, second viewpoint information indicating the position of a second virtual viewpoint different from the first virtual viewpoint and the line-of-sight direction from the second virtual viewpoint, and acquisition means for acquiring an input of additional information for a first virtual viewpoint image generated based on the first viewpoint information and a virtual space. Generation means for generating a three-dimensional additional object arranged in the virtual space based on the additional information. Display control means for controlling the display means to display a second virtual viewpoint image generated based on the virtual space in which the additional object is arranged and the second viewpoint information. The information processing apparatus having: The generation means generates the additional object on a plane parallel to a plane corresponding to the ground of a background object arranged in the virtual space. The second virtual viewpoint image includes the additional object. An information processing apparatus characterized by the above. Claims 2. The information processing apparatus according to claim 1, wherein the additional information is information indicating at least one of a curve, a straight line, and a point. Claims 3. The information processing apparatus according to claim 1, wherein the additional information is information indicating a subject to be noted. Claims 4. The information processing apparatus according to claim 3, wherein the additional information is a curve surrounding a subject to be noted. Claims 5. The information processing apparatus according to claim 1, wherein the acquisition means acquires the input of the additional information by acquiring an input to a screen on which the first virtual viewpoint image is displayed. Claims 6. The information processing apparatus according to claim 5, wherein the additional object is generated at a position of an intersection of a straight line passing through a position where the additional information is input on a plane in the virtual space corresponding to the first virtual viewpoint and the screen, and the parallel plane. Claims 7. The information processing apparatus according to claim 1, further comprising recording means for associating and recording a time code corresponding to the first virtual viewpoint image and the additional object. Claims 8. The information processing apparatus according to claim 1, wherein the additional object included in the second virtual viewpoint image is displayed thinner as the difference between the time code corresponding to the first virtual viewpoint image and the time code corresponding to the second virtual viewpoint image is larger.

9. An information processing device as described in claim 1, characterized in that the first virtual viewpoint image and the second virtual viewpoint image are generated based on a plurality of captured images.

10. An information processing method executed by an information processing device, comprising: an acquisition step of acquiring first viewpoint information indicating a position of a first virtual viewpoint and a line of sight direction from the first virtual viewpoint, second viewpoint information indicating a position of a second virtual viewpoint different from the first virtual viewpoint and a line of sight direction from the second virtual viewpoint, and additional information for a first virtual viewpoint image generated based on the first viewpoint information and a virtual space; a generating step of generating an additional object having a three-dimensional shape to be placed in the virtual space based on the additional information; a display control step of controlling display on a display means of the virtual space in which the additional object is arranged and a second virtual viewpoint image generated based on the second viewpoint information; having In the generating step, the additional object is generated on a plane parallel to a plane corresponding to a ground of a background object arranged in the virtual space; the second virtual viewpoint image includes the additional object, 23. An information processing method comprising:

11. A program for causing a computer to function as each of the means possessed by an information processing device described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Image processing device, image processing method, and image processing program

    JP2014032443A

  • Rendering, viewing and annotating panoramic images, and applications thereof

    JP2014170580A

  • Image display device, image processing system, image processing method, and image processing program

    JP2017151491A

  • Superimposed image display device and computer program

    JP2019095215A

  • Drawing in a 3D virtual reality environment

    US20180101986A1