Image processing device, image processing method, and program

The image processing apparatus addresses miscommunication in virtual coaching by generating synchronized virtual viewpoint images that adapt to the player's viewpoint changes, ensuring effective instruction delivery.

JP2026084343APending Publication Date: 2026-05-21CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-11-11
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing technologies fail to accurately convey coaching instructions in virtual viewpoint images when the instructor and the instructed person are in different locations within a wide virtual space or when the object being instructed changes over time, leading to miscommunication.

Method used

An image processing apparatus that generates virtual viewpoint images based on multiple camera inputs, allowing a coach to specify a virtual viewpoint and display instructional content to a player wearing an HMD, while dynamically adjusting the viewpoint to match the player's gaze and head movements.

Benefits of technology

Enables smooth communication of coaching instructions by synchronizing virtual viewpoint images between multiple users, ensuring the player accurately receives instructions despite changes in location or focus within the virtual space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026084343000001_ABST
    Figure 2026084343000001_ABST
Patent Text Reader

Abstract

To enable smooth communication among multiple users while sharing virtual viewpoint images. [Solution] An image processing device that generates and outputs a virtual viewpoint image generates a virtual viewpoint image based on multiple images captured by multiple imaging devices, reflecting instructions from the first user to the second user, based on the operation input of the first user. Then, based on the operation input of the first user, the generated virtual viewpoint image is output to a user terminal used by the second user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image processing apparatus, a control method, and a program for virtual viewpoint images.

Background Art

[0002] In recent years, there has been a technology of installing a plurality of cameras at different positions and synchronously imaging from multiple viewpoints, and using the plurality of images obtained by the imaging to generate not only an image at the camera installation position but also a virtual viewpoint image at any one or more viewpoints. In services using virtual viewpoint images, for example, in a basketball game, from a plurality of images obtained by synchronously imaging with a plurality of cameras, a video producer can produce compelling virtual viewpoint content such as a video from the perspective of a player. Also, it is possible for a user who is viewing virtual viewpoint content to freely move the virtual viewpoint, and the user can watch the game while viewing virtual viewpoint images corresponding to various viewpoints.

[0003] It has also been proposed to use such virtual viewpoint images, for example, in sports coaching. In the case of sports coaching applications, it is assumed that a player (the instructed person) receiving coaching watches a virtual viewpoint image with an HMD, for example, while receiving instructions from a coach (the instructor). In such a situation, it is necessary to enable the player to accurately understand the instruction content from the coach. In this regard, Patent Document 1 discloses a video presentation system that reflects the instruction content from another user who watches a virtual space video with a non-HMD to the virtual space video of the HMD for a user who watches the virtual space video with the HMD.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In the technology described in Patent Document 1, the display position of the instruction content is determined based on the virtual space image viewed by the person giving the instruction (who is not wearing an HMD), and the instruction content is displayed on the virtual space image viewed by the person receiving the instruction (who is wearing an HMD). Now, consider, for example, when the virtual space is wide and the person receiving the instruction and the person giving the instruction are in different locations within the virtual space, or when the position of the object being instructed changes over time, causing the person receiving the instruction and the person giving the instruction to focus on different things within the virtual space. In such cases, the technology described in Patent Document 1 may not correctly convey the instruction content from the person giving the instruction to the person receiving the instruction. [Means for solving the problem]

[0006] The image processing apparatus according to this disclosure is characterized by comprising: a receiving means for receiving operation input from a first user; an image generation means for generating a virtual viewpoint image based on a plurality of images captured by a plurality of imaging devices, which reflects instructions from the first user to a second user based on the operation input from the first user; and an output control means for outputting the generated virtual viewpoint image to a user terminal used by the second user based on the operation input from the first user. [Effects of the Invention]

[0007] According to this disclosure, it is possible to communicate smoothly with multiple users while sharing virtual viewpoint images. [Brief explanation of the drawing]

[0008] [Figure 1] (a) and (b) are diagrams showing an example of the configuration of an image processing system. [Figure 2] (a) and (b) are diagrams showing an example of the hardware configuration of a user terminal and an image processing device. [Figure 3] A diagram showing an example of the software configuration (functional configuration) of an image processing device. [Figure 4] (a) and (b) are diagrams showing examples of the GUI of the first user terminal. [Figure 5]A diagram showing an example of a second virtual viewpoint image displayed on an HMD. [Figure 6] A flowchart illustrating the generation and output process of virtual viewpoint images performed by an image processing device. [Figure 7] A flowchart detailing the image output process when screen sharing is not being accepted. [Figure 8] A flowchart showing the details of the image output process when screen sharing is in progress. [Figure 9] A diagram showing an example of the GUI of the first user terminal. [Modes for carrying out the invention]

[0009] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims. While the embodiments describe multiple features, not all of these features are essential to the invention, and the features may be combined in any way. Furthermore, in the attached drawings, the same or similar configurations are given the same reference numeral, and redundant descriptions may be omitted.

[0010] [Embodiment 1] This embodiment describes an image processing system that generates a virtual viewpoint image representing the view from a specified virtual viewpoint, based on multiple images captured by multiple imaging devices and a specified virtual viewpoint. In this embodiment, the "virtual viewpoint image" is not limited to an image corresponding to a virtual viewpoint freely (arbitrarily) specified by the user; for example, an image corresponding to a virtual viewpoint selected by the user from among multiple virtual viewpoint candidates is also included as a virtual viewpoint image. Furthermore, this embodiment mainly describes the case where the virtual viewpoint is specified by user input, but the virtual viewpoint may be specified automatically based on the results of image analysis, etc. In addition, this embodiment mainly describes the case where the virtual viewpoint image is a video, but the virtual viewpoint image may be a still image.

[0011] Furthermore, in this embodiment, the term "virtual camera" will be used for explanation. A virtual camera is a virtual imaging device that is different from the multiple imaging devices actually installed around the imaging area, and is a concept used to conveniently explain the virtual viewpoint related to the generation of virtual viewpoint images. In other words, a virtual viewpoint image can be considered as an image captured from a virtual viewpoint set in a virtual space associated with the imaging area. The position and orientation of the viewpoint in the virtual imaging can be represented as the position and orientation (orientation) of the virtual camera. To put it another way, a virtual viewpoint image can be said to be an image that simulates the image obtained by a camera, assuming that the camera exists at the position of a virtual viewpoint set in a virtual space corresponding to the real space where actual imaging takes place.

[0012] <System Configuration> This embodiment will be explained using a use case in which a basketball coach (instructor) gives instructions to a player (instructor) using a virtual viewpoint image. In this use case, it is assumed that the coach uses a desktop PC as the first user terminal to give instructions, and the player wears an HMD as the second user terminal to receive instructions from the coach. If the coach tries to give instructions to the player using a virtual viewpoint image displayed on the desktop PC's display, the player will have to remove the HMD each time, and will not be able to concentrate on viewing the virtual viewpoint image displayed on the HMD. Therefore, it is conceivable to generate a virtual viewpoint image by placing a CG 3D model representing the instructions from the coach to the player (hereinafter referred to as the "instruction model") in a virtual space. However, if the player's line of sight is not directed towards the direction in which the instruction model is placed, and the instruction model is outside the field of view of the virtual viewpoint image that the player is looking at, the player will ultimately not be able to recognize the coach's instructions. Furthermore, even if it is possible to place the instruction model while displaying the virtual viewpoint image that the player is looking at on the desktop PC's display, it is difficult to place the instruction model in the appropriate position while viewing an image that shakes in accordance with the player's head movements. Furthermore, it is necessary to continuously display the instruction model on the virtual viewpoint image in accordance with the player's gaze direction, which may change moment by moment. Therefore, in this embodiment, an image processing system particularly suitable for such coaching will be described.

[0013] First, the configuration of the image processing system according to this embodiment will be described with reference to Figures 1(a) and (b). The image processing system of this embodiment has n sensor systems 10a to 10n, and each sensor system has at least one imaging device, which is a camera. Hereafter, unless otherwise specified, the n sensor systems will not be distinguished and will be referred to as "multiple sensor systems 10".

[0014] Figure 1(a) shows an example of the installation of multiple sensor systems 10 surrounding an area 12 to be imaged in a real three-dimensional space, and a non-existent virtual camera 11. Each of the multiple sensor systems 10 images the area 12 from different directions. In this embodiment, the area 12 to be imaged is assumed to be a basketball court, and n (for example, 100) sensor systems 10 are installed around the court. The area 12 to be imaged may include seating areas in addition to the basketball court. Furthermore, the area 12 to be imaged is not limited to indoors, but may be an outdoor stadium or stage. Also, the multiple sensor systems 10 do not have to be installed around the entire perimeter of the area 12, and may be installed only around a part of the area 12 depending on the installation location constraints. In addition, the multiple cameras of the multiple sensor systems 10 may include imaging devices with different functions, such as telephoto cameras and wide-angle cameras. The multiple cameras of the multiple sensor systems 10 perform imaging in synchronous motion. The multiple images obtained by imaging from these multiple cameras are called "multi-view images". In this embodiment, each of the multi-view images may be an captured image, or it may be an image obtained by performing image processing on the captured image, such as foreground extraction (foreground image).

[0015] The virtual camera 11 is set up in a virtual space associated with region 12 and can be set to a viewpoint position different from any camera of the multiple sensor systems 10. The virtual viewpoint image generated by the image processing device 200 is an image representing the view from the virtual camera 11. Here, it is possible to set different virtual cameras for the virtual viewpoint image provided to the first user terminal 100A and the virtual viewpoint image provided to the second user terminal 100B. In addition to cameras, the multiple sensor systems 10 may also have microphones (not shown). The microphones of each of the multiple sensor systems 10 synchronize to pick up sound. Based on this picked-up sound, an acoustic signal can be generated that will be played back together with the display of the virtual viewpoint image on the user terminal described later. For the sake of simplicity, the description of sound will be omitted from the following explanation, but it will be assumed that images and sound are basically processed together.

[0016] FIG. 1(b) is a diagram showing the configuration of the entire image processing system according to the present embodiment. The image processing system includes, in addition to the plurality of sensor systems 10 described above, a first user terminal 100A, a second user terminal 100B, and an image processing device 200.

[0017] The first user terminal 100A is an information processing device such as a desktop PC or a tablet terminal used for a first user (in this embodiment, a coach who guides a player) to specify a virtual viewpoint or view a virtual viewpoint image. The first user terminal 100A receives an operation signal of a virtual camera via a mouse or the like by the first user and transmits it to the image processing device 200. Further, the first user terminal 100A displays the virtual viewpoint image received from the image processing device 200 on a display device (not shown) such as an external or built-in liquid crystal display.

[0018] The image processing device 200 is an information processing device as an image processing server that generates a virtual viewpoint image and provides it to the first user terminal 100A / second user terminal 100B. The image processing device 200 acquires multi-viewpoint images from the plurality of sensor systems 10 and stores them in a database (not shown) together with the time code at the time of imaging. The time code is information for uniquely identifying the time when the imaging device captured an image, and is held in a format such as 'year:month:day:hour:minute:second:frame number'. Then, the image processing device 200 generates a virtual viewpoint image corresponding to the specified virtual viewpoint using the multi-viewpoint images stored in the database and provides it to the first user terminal 100A and the second user terminal 100B. Generation of the virtual viewpoint image is performed, for example, by model-based rendering (MBR). MBR is a method of generating a virtual viewpoint image using three-dimensional shape data (3D model) of an object generated based on a plurality of images obtained by imaging the object (subject) from a plurality of directions. The 3D model can be obtained by a three-dimensional shape restoration method such as a volume intersection method.

[0019] The second user terminal 100B is an information processing device, such as an HMD or tablet terminal, for a second user (in this embodiment, an athlete) to view the virtual viewpoint image generated by the image processing device 200. It is also possible to specify the virtual viewpoint based on the operation input of the second user terminal 100B by the second user. For example, if the second user terminal 100B is an HMD, the second user wearing the HMD can move their head to generate a so-called first-person viewpoint operation signal, which is sent to the image processing device 200 as information to specify the virtual viewpoint.

[0020] <Hardware configuration of user terminal and image processing device> Next, an example of the hardware configuration of the user terminal and image processing device in this embodiment will be described with reference to Figures 2(a) and (b).

[0021] Figure 2(a) shows an example of the hardware configuration of a user terminal, which is an information processing device. Here, we will describe the first user terminal 100A, but the second user terminal 100B has a similar configuration.

[0022] The CPU 101 uses the RAM 102 as work memory to execute programs stored in the ROM 103 and / or the hard disk drive (HDD) 105, and controls the various configurations described later via the system bus 112. This enables the execution of various processes described later.

[0023] The HDD interface (I / F) 104 is an interface, such as Serial ATA (SATA), that connects the user terminal 100A and the HDD 105. The CPU 101 can read data from and write data to the HDD 105 via the HDD interface (I / F) 104. Furthermore, the CPU 101 loads the data stored in the HDD 105 into the RAM 102. The CPU 101 can also save various data obtained from the RAM 102 through program execution to the HDD 105. Note that the HDD is just one example of a secondary storage device; an optical disc drive, SSD, flash memory, etc., may also be used.

[0024] The input interface (I / F) 106 connects the user terminal 100A to an input device 107, such as a touch panel, keyboard, mouse, digital camera, or scanner, for inputting one or more coordinates. The input interface (I / F) 106 is a serial bus interface, such as USB or IEEE1394. The CPU 101 can read data from the input device 107 via the input I / F 106.

[0025] The output interface (I / F) 108 connects the user terminal 100A to an output device 109, such as a display. The output interface (I / F) 108 is a video output interface such as DVI or HDMI (registered trademark). The CPU 101 sends data related to the virtual viewpoint image to the output device 109 via the output I / F 108, thereby executing the display of the virtual viewpoint image. The network interface (I / F) 110 connects the user terminal 100A to an external server 111 and is a network card such as a LAN card. The CPU 101 can read data from the external server 111 via the network interface 110.

[0026] Figure 2(b) shows an example of the hardware configuration of the image processing device 200. The CPU 201 uses the RAM 202 as work memory to execute the program stored in the ROM 203 and controls the various components described later.

[0027] The communication unit 204 connects to external devices such as the first user terminal 100A and the second user terminal 100B and performs data communication. The communication unit 204 communicates according to communication standards such as Ethernet and IEEE 802.11 (so-called wireless LAN). The CPU 201 sends and receives data from external devices via the communication unit 204.

[0028] The input / output unit 205 performs data input and output via input and output interfaces (not shown). Devices such as a mouse, keyboard, display, and digital camera are connected to the input / output unit 205.

[0029] The GPU206 is a computing unit specialized for image processing. The GPU206 performs rendering processes to generate a virtual viewpoint image from multi-view images input from multiple sensor systems 10.

[0030] The HDD207 is a secondary storage device for storing image data and other data. Note that other types of storage devices such as optical disc drives, SSDs, and flash memory can also be used.

[0031] <Software configuration of the image processing device> Next, the functional configuration of the image processing apparatus 200 according to this embodiment will be described with reference to Figure 3. The image processing apparatus 200 consists of a first playback section determination unit 301, a first viewpoint determination unit 302, a first video generation unit 303, a shared information generation unit 304, a shared information storage unit 305, a shared processing unit 306, a second playback section determination unit 307, a second viewpoint determination unit 308, and a second video generation unit 309.

[0032] The first playback section determination unit 301 determines the playback section of the virtual viewpoint image generated by the first video generation unit 303 according to the operation signal of the first user input from the first user terminal 100A. Here, the “playback section” means the time range of the virtual viewpoint image within the total time range of the input multi-view images, and is defined, for example, using a start time code indicating the start time and an end time code indicating the end time.

[0033] The first viewpoint determination unit 302 determines external parameters representing the virtual viewpoint, i.e., the position and orientation of the virtual camera, for the first video generation unit 303 to generate a virtual viewpoint image, in accordance with the operation signals of the first user input from the first user terminal 100A. Here, the position of the virtual camera is represented, for example, by a three-dimensional coordinate system (x, y, z) consisting of three axes: the x, y, and z axes. The orientation of the virtual camera is specified, for example, by the values ​​of three axes: the pan, tilt, and roll axes (pan, tilt, roll). The pan axis represents the left-right movement of the camera, the tilt axis represents the up-down movement of the camera, and the roll axis represents the rotation of the camera around the optical axis. It is assumed that internal parameters such as the focal length and field of view (imaging area) of the virtual camera are predetermined.

[0034] The first video generation unit 303 generates a virtual viewpoint image based on the input multi-viewpoint image, the playback interval determined by the first playback interval determination unit 301, and the position and orientation of the virtual camera determined by the first viewpoint determination unit 302.

[0035] The shared information generation unit 304 generates image generation information (hereinafter referred to as "shared information") for sharing a virtual viewpoint image between the first user and the second user and for conveying instructions from the first user to the second user. This shared information includes information indicating the content of the instructions from the first user to the second user, information indicating the display position of the instructions, and information related to the virtual viewpoint, such as the position, orientation, and playback interval of the virtual camera used to generate the virtual viewpoint image. Here, "information indicating the content of the instructions" includes, for example, CG of a string of characters representing advice that a coach wants to convey to a player, CG representing shapes such as arrows that indicate areas to be noticed in the image, and, for example, changes in the color of a specific area in the image. The "display position (of the instructions)" is the position in the virtual space where the instructions are displayed, and is represented, for example, by 3D coordinates (x, y, z). The shared information may include other elements (for example, the playback speed of the virtual viewpoint image), or it may not include all of the elements described above. The generated shared information is stored in the RAM 202 of the image processing device 200.

[0036] The sharing processing unit 306 retrieves all the shared information generated by the shared information generation unit 304 and stored in the RAM 202, and selects one piece of shared information desired by the first user. The selected shared information is sent to the second playback section determination unit 307 and the second viewpoint determination unit 308.

[0037] The second playback interval determination unit 307 determines the playback interval of the virtual viewpoint image generated by the second video generation unit 303. In this case, while a screen sharing operation by the first user is being accepted (while the sharing processing unit 306 has selected sharing information), the playback interval is determined based on that sharing information. When a screen sharing operation is not being accepted, the playback interval is determined according to the operation input of the first user input from the first user terminal 100A. Alternatively, when a screen sharing operation is not being accepted, the playback interval may be determined according to the operation input of the second user input from the second user terminal 100B. This allows the player to view the desired virtual viewpoint image at any playback interval when the coach is not giving a screen sharing instruction.

[0038] The second viewpoint determination unit 308 determines the virtual viewpoint, i.e., the external parameters representing the position and orientation of the virtual camera, for the second video generation unit 303 to generate a virtual viewpoint image. In this case, while the first user is accepting screen sharing operations, the virtual viewpoint's position (x, y, z) and orientation (pan, tilt, roll), except for the virtual viewpoint's height (x), are determined to the position and orientation values ​​included in the shared information. The virtual viewpoint's height (x) is determined according to the operation signals of the second user input from the second user terminal 100B. As a result, a virtual viewpoint image is generated that, while largely following the virtual viewpoint determined by the first user (coach), is matched to the eye level of the second user (player). On the other hand, when the first user is not accepting screen sharing operations, the virtual viewpoint is determined according to the operation input of the second user input from the second user terminal 100B (in the case of an HMD, operation signals representing the movement of the second user's head detected by the internally mounted acceleration sensor and gyroscope sensor). Similar to the first viewpoint determination unit 302, internal parameters such as the focal length and field of view (imaging area) of the virtual camera are assumed to be predetermined.

[0039] The second video generation unit 309 generates a virtual viewpoint image based on the input multi-viewpoint image, the playback section determined by the second playback section determination unit 307, and the position and orientation of the virtual camera determined by the second viewpoint determination unit 308.

[0040] The output control unit 305 performs output control of the virtual viewpoint images generated by the first video generation unit 303 and the second video generation unit 309. Specifically, via the communication unit 204, the virtual viewpoint image (hereinafter referred to as the "first virtual viewpoint image") generated by the first video generation unit 303 is sent to the first user terminal 100A, and the virtual viewpoint image (hereinafter referred to as the "second virtual viewpoint image") generated by the second video generation unit 309 is sent to the first user terminal 100A and the second user terminal 100B.

[0041] <Explanation of GUI> Next, the graphical user interface (GUI) of the first user terminal 100A will be described. Here, a scenario where a basketball coach as the first user gives instructions to a player as the second user will be used as an example for the explanation. FIG. 4(a) shows the GUI when the aforementioned shared information has not been generated and the screen sharing operation has not been received from the coach, and FIG. 4(b) shows the GUI when the aforementioned shared information has been generated and the screen sharing operation has been received from the coach. The following will explain each GUI.

[0042] ≪GUI When Screen Sharing Operation is Not Accepted≫ The GUI 400 shown in FIG. 4(a) is composed of UI elements such as three image areas ۴۰۱~۴۰۳, a seek bar 404, a play / pause button 405, a speed button 406, a text input field 407, a save button 408, and a delete button 409.

[0043] Image area 401 is the image area where the second virtual viewpoint image displayed on the second user terminal 100B is shown. This allows the coach to check the virtual viewpoint image that the player is viewing. Currently, image area 401 displays a virtual viewpoint image of a specific moment during a basketball game, with the vertically elongated rectangular prism being a simplified representation of a human and the sphere being a representation of a basketball.

[0044] Image area 402 is an image area for the first user to prepare a second virtual viewpoint image, and displays a first virtual viewpoint image that can only be viewed by the first user. For example, the coach can operate the virtual camera using the input device 107 of the first user terminal 100A to specify a desired position and orientation. For example, if the input device 107 is a mouse, the coach can change the position of the virtual camera by dragging with a left click on image area 402, and change the orientation of the virtual camera by dragging with a right click.

[0045] The seek bar 404 is a UI element that indicates the playback interval of a virtual viewpoint image. For example, a coach can set an arbitrary playback interval by manipulating the left circle corresponding to the start point and the right circle corresponding to the end point on the seek bar 404, specifying the start / end timecodes of the second virtual viewpoint image that they want the player to see.

[0046] The play / pause button 405 controls the playback or pause of the first virtual viewpoint image displayed in the image area 402. The speed button 406 switches the playback speed of the first virtual viewpoint image displayed in the image area 402, allowing the user to specify any playback speed from options presented by a pull-down menu, for example. For instance, a playback speed of "1.0" plays at normal speed, a value less than "1.0" plays in slow motion, and a value greater than "1.0" plays at high speed.

[0047] The text input field 407 is a UI element for entering a string of text as instructions from the first user to the second user. For example, the coach uses the keyboard, which is the input device 107, to enter brief sentences such as reminders or situational explanations for the players. Based on the string entered here, a text box, which is one form of the instruction model, is generated.

[0048] Image area 403 is an image area that displays a virtual viewpoint image (hereinafter referred to as the "overhead view image") corresponding to a fixed virtual viewpoint that provides an overhead view of the imaging area (in this embodiment, a basketball court) captured by the multiple sensor systems 10. Similar to image area 1, the overhead view image displayed in image area 3 can only be viewed by the first user, the coach. Currently, image area 403 displays a virtual viewpoint image of the entire basketball court as seen from directly above, and the black rectangles in the image are figures that represent players in play. The overhead view image displayed in image area 403 allows the coach to easily grasp the positional relationships between players during a game.

[0049] The Save button 408 is used to save the position and orientation of the virtual camera for the second virtual viewpoint image and the instructions for the second user as shared information after the first user has specified these settings. Note that multiple shared information items can be saved by repeatedly performing the shared information setting operation described above and pressing the Save button 408 while screen sharing is not being accepted (when no shared information is selected). Once the shared information is saved, as shown in Figure 4(b) below, the same number of marks as the number of saved shared information items will be displayed on the overhead image in image area 403.

[0050] ≪GUI accepting screen sharing requests≫ The GUI400' shown in Figure 4(b) consists of the following UI elements: three image areas 401-403, a seek bar 404, a play / pause button 405, a speed button 406, a text input field 407, an unshare button 409, and a delete button 410. The image areas 401-403, the seek bar 404, the play / pause button 405, the speed button 406, and the text input field 407 are the same as those in the GUI400 in Figure 4(a). The differences from the GUI400 in Figure 4(a) will be explained below.

[0051] Currently, within the overhead view displayed in image area 403 of GUI400', there are three star-shaped marks 411 indicating shared information. The placement of the three marks 411 indicates the display position in the virtual space of the instruction model contained in each of the saved shared information. The coach places a mark 411 at any position in the overhead view, for example, by left-clicking. This determines the position on the court where the instruction model will be displayed (two-dimensional coordinate values ​​in the xy plane horizontal to the court). Here, the coordinate value in the direction perpendicular to the court (z-axis direction) is set to a predetermined value, for example, a height of 2m. Note that the operation method described here is just one example, and left-clicking may be assigned to another function (for example, moving the virtual camera in image area 402). In that case, the display position of the instruction model should be determined by another operation method, for example, by left-clicking while holding down the Ctrl key. Through these operations, the position in the virtual space where the instruction model will be displayed (three-dimensional coordinate values) is determined. After the mark 411 is placed in the overhead image and the display position of the instruction model is determined, when the coach presses the aforementioned save button 408, the mark 411 will continue to be displayed in the overhead image of image area 403. Then, when the coach performs a screen sharing operation (for example, a mark selection operation such as hovering the cursor over any mark 411 and right-clicking), a second virtual viewpoint image is generated based on the shared information related to the screen sharing operation and is displayed on both the first user terminal 100A and the second user terminal 100B. Figure 5 shows the second virtual viewpoint image displayed on the HMD as the second user terminal 100B. In this way, the virtual viewpoint image including the text box 412 of the string "Pay attention to a feint of the number 3" entered by the coach in the text input field 407 is displayed on both the image area 401 of GUI400' and the HMD. This means that the screen display on the HMD is forcibly switched to a virtual viewpoint image with instructions from the viewpoint the coach wants to show, as a result of the coach's screen sharing operation. This allows coaches to ensure that players see the virtual viewpoint images they want to show them, which include the instructions they want to convey.

[0052] The Unshare button 409 is used by the first user to cancel screen sharing. This Unshare button 409 is displayed while screen sharing is being accepted from the first user. When the Unshare button 409 is pressed, the sharing information related to the currently accepted screen sharing operation changes from selected to unselected, and the generation and output of the second virtual viewpoint image based on the sharing information stops.

[0053] The Delete button 410 is used to delete any shared information from the saved shared information. For example, if this Delete button 410 is pressed while screen sharing is in progress (while selecting shared information), all the contents of the selected shared information will be deleted. This allows coaches to delete shared information that is no longer needed.

[0054] <Processing by the image processing device 200> Next, the virtual viewpoint image generation and output processing performed by the image processing device 200 will be explained using the flowcharts in Figures 6 to 8. The series of processes shown in the flowcharts in Figures 6 to 8 are realized by the CPU 201 or GPU 206 loading the software stored in ROM 203 into RAM 202 and executing it.

[0055] ≪Main Flow≫ Figure 6 is a flowchart showing the general flow of the image generation process in the image processing device 200 according to this embodiment, and is executed frame by frame. The following explanation will follow the flowchart in Figure 6. In the following explanation, the symbol "S" represents a step.

[0056] In S601, multi-view images, which serve as source data necessary for generating virtual viewpoint images, are acquired from a database (not shown).

[0057] In S602, the next process to be executed is determined based on whether or not a screen sharing operation by the first user (in this embodiment, the operation of selecting a mark displayed on the overhead image) is currently being accepted. If a screen sharing operation is not being accepted, S603 is executed next; if it is being accepted, S604 is executed next.

[0058] In S603, image output processing is performed when the first user's screen sharing operation is not being accepted. On the other hand, in S604, image output processing is performed when the first user's screen sharing operation is being accepted. Details of the image output processing for S603 and S604 will be described later.

[0059] In S605, it is determined whether or not to continue the output processing of the virtual viewpoint image. If it is to continue, the process returns to S602 and continues for the next frame. On the other hand, if it is not to continue (for example, if the application for playing the virtual viewpoint image is terminated), the processing in this flowchart is terminated. The above is a general overview of the image output processing in the image processing device 200.

[0060] ≪Image output processing when screen sharing is not being accepted≫ Figure 7 is a flowchart detailing the image output process in the S603 described above when the first user's screen sharing operation is not being accepted. The following explanation will follow the flowchart in Figure 7. In the following explanation, the symbol "S" represents a step.

[0061] In S701, the next process to be executed is determined by whether or not a control value related to the operation input by the first user (hereinafter referred to as the "first input value") has been input from the first user terminal 100A. If the first input value has not been received, S712 is executed next; if it has been received, S702 is executed next.

[0062] In S702, the first playback section determination unit 301 determines the playback section of the first virtual viewpoint image based on the first input value received in S701. The first input value assumed here is, for example, a control value such as a time code corresponding to the mouse or other operation input by the first user to the seek bar 404 in the GUI 400 shown in Figure 4 above. For example, the first user specifies the start time code and end time code by dragging both ends of the seek bar 404. If the first input value received in S701 is not an input value related to the playback section, this step is skipped.

[0063] In S703, the first video generation unit 303 causes the GPU 206 to perform rendering processing on the target frame from the multi-view images acquired in S601, based on the playback section determined in S702. This generates a virtual viewpoint image (overhead view image) that represents the view from a pre-set overhead viewpoint. Note that the generation of the overhead view image may be performed by the second video generation unit 309, or a third video generation unit (not shown) may be provided separately for generating the overhead view image.

[0064] In S704, the first viewpoint determination unit 302 determines the position and orientation of the virtual camera based on the first input value received in S701. The first input value assumed here is, for example, a control value corresponding to an input operation by the first user using a mouse or the like to specify the position and orientation of the virtual camera relative to the image area 402 in the GUI 400 shown in Figure 4(a) above. If the first input value received in S701 is not an input value related to the position and orientation of the virtual camera, this step is skipped.

[0065] In S705, the shared information generation unit 304 generates an instruction model based on the first input value received in S701. The first input value assumed here is, for example, a string entered in the text input field 407 in the GUI 400 shown in Figure 4(a) above, and a text block containing that string is generated by computer graphics (CG). The generated text block is stored in RAM 202. If the first input value received in S701 is not an input value related to the generation of an instruction model, this step is skipped.

[0066] In S706, the shared information generation unit 304 determines the display position (position on the xy plane) of the instruction model generated in S705 within the virtual space based on the first input value received in S701. The first input value assumed here is, for example, an input operation signal that specifies the position of the virtual camera using a mouse or the like by the first user relative to the image area 403 in the GUI 400 shown in Figure 4(a) above. If the first input value received in S701 is not an input value related to the display position of the instruction model, this step is skipped.

[0067] In S707, the first video generation unit 303 causes the GPU 206 to perform rendering processing on the target frame from the multi-view images acquired in S601, based on the playback interval determined in S702. This generates a first virtual viewpoint image representing the view from the virtual camera determined in S704.

[0068] In S708, the next process to be executed is determined by whether the first input value received in S701 is a shared information save operation or not. If the received first input value is a shared information save operation (for example, a signal value indicating that the save button 408 in the GUI 400 shown in Figure 4(a) above has been pressed), then S709 is executed next. On the other hand, if the received first input value is a shared information save operation, then S711 is executed next.

[0069] In S709, the shared information generation unit 304 associates the playback section determined in S702, the position and orientation of the virtual camera determined in S704, and the instruction model and its display position generated / determined in S705 / S706 with each other and saves them as shared information. When saving, an ID or similar is assigned to distinguish it from other shared information, and it is saved, for example, on HDD207.

[0070] In S710, a mark representing the shared information saved in S709 (a star-shaped mark in this embodiment) is added to the overhead image generated in S703, based on the position in the virtual space (position on the xy plane) determined in S706.

[0071] In S711, the next process to be executed is determined based on whether the first input value received in S701 is a control value for a screen sharing operation by the first user (in this embodiment, an operation to select a mark displayed on the overhead image). If the received first input value is a screen sharing operation, this process is exited and the system returns to the flowchart in Figure 6. On the other hand, if the received first input value is not a screen sharing operation, S712 is executed next.

[0072] In S712, the next process to be executed is determined by whether or not a value related to the operation input by the second user (hereinafter referred to as the "second input value") has been input from the second user terminal 100B. If the second input value has not been received, S715 is executed next; if it has been received, S713 is executed next.

[0073] In S713, the second viewpoint determination unit 308 determines the position and orientation of the virtual camera based on the second input value received in S712. The second input value assumed here is, for example, a sensor signal value corresponding to the head movement of the second user wearing the HMD as the second user terminal 100B. If the second input value received in S712 is not an input value related to the position and orientation of the virtual camera, this step and the following S714 are skipped.

[0074] In S714, the second video generation unit 309 performs rendering processing on the target frame from the multi-view images acquired in S601, based on the playback interval determined in S702. This generates a virtual viewpoint image (hereinafter referred to as the "second virtual viewpoint image") that represents the view from the virtual camera determined in S712.

[0075] In S715, the overhead image generated in S703 and the first virtual viewpoint image generated in S707 are transmitted to the first user terminal 100A via the communication unit 204. The second virtual viewpoint image generated in S714 is also transmitted to both the first user terminal 100A and the second user terminal 100B via the communication unit 204. The first user terminal 100A then displays the received overhead image, the first virtual viewpoint image, and the second virtual viewpoint image in designated image areas on the GUI. The second user terminal 100B displays the received second virtual viewpoint image. After this step is completed, the flow exits and returns to the flowchart in Figure 6.

[0076] The above is a description of the flowchart when screen sharing operations are not being accepted. Through this process, when screen sharing operations by the first user are not being accepted, a virtual viewpoint image is generated and output according to the playback section specified by the first user, and the virtual viewpoint image is looped and played on both the first user terminal 100A and the second user terminal 100B. As a result, once the virtual viewpoint image is generated, the user can repeatedly view the virtual viewpoint image of the same scene without any further operation.

[0077] ≪Image output processing while screen sharing requests are being accepted≫ Figure 8 is a flowchart detailing the image output process when the first user is requesting screen sharing in the S604 mentioned above. The following explanation will follow the flowchart in Figure 8. In the following explanation, the symbol "S" represents a step.

[0078] In S801, the sharing processing unit 306 reads the sharing information related to the currently requested screen sharing operation (in this embodiment, the sharing information associated with the selected mark) from the HDD 207 and stores it in the RAM 202. Once the sharing information related to the currently requested screen sharing operation has been read, this step is skipped thereafter.

[0079] In S802, the second playback section determination unit 307 determines the playback section of the second virtual viewpoint image based on the shared information read in S801. Specifically, the start timecode and end timecode included in the read shared information are set as the playback section. Based on the playback section set in this way, playback starts from the position of the start timecode on the seek bar 404 and continues until the position of the end timecode. If the shared information related to the currently accepted screen sharing operation has already been read, this step is skipped thereafter.

[0080] In S803, the second image generation unit 303 causes the GPU 206 to perform rendering processing based on the multi-view images acquired in S601 to generate an overhead image representing the view from a pre-set overhead viewpoint. Note that the generation of the overhead image may be performed by the first image generation unit 303, or a third image generation unit (not shown) may be provided separately for generating the overhead image.

[0081] In S804, the second viewpoint determination unit 302 determines the position and orientation of the virtual camera based on the shared information read in S801 and the second input value input from the second user terminal. As described above, the position (x,y) and orientation (pan,tilt,roll) of the virtual camera are determined by the position and orientation values ​​included in the shared information, and the height (x) of the virtual camera is determined according to the operation signal value of the second user.

[0082] In S805, the second video generation unit 309 causes the GPU 206 to perform rendering processing based on the multi-view images acquired in S601, generating a second virtual viewpoint image that represents the view from the virtual camera determined in S804. Specifically, the instruction model (a text box in this embodiment) included in the shared information read in S801 is placed in the virtual space based on the display position included in the shared information and rendered together with the player's 3D model. At this time, the text box is positioned to face the virtual camera so that the text information within the text box is easily recognizable by the viewer. In this way, a virtual viewpoint image including an instruction model representing the coach's instructions is generated. Note that it is sufficient for the instruction content to be reflected in the virtual viewpoint image, so for example, rendering processing may be performed without placing the instruction model in the virtual space, and a two-dimensional CG corresponding to the instruction model may be synthesized onto the obtained virtual viewpoint image.

[0083] In S806, the next process to be executed is determined by whether or not a control value (first input value) related to the operation input by the first user has been input from the first user terminal 100A. If the first input value has not been received, S809 is executed next; if it has been received, S807 is executed next.

[0084] In S807, the first viewpoint determination unit 302 determines the position and orientation of the virtual camera based on the first input value received in S806. Subsequently, in S808, the first image generation unit 303 causes the GPU 206 to perform rendering processing based on the multi-view images acquired in S601 to generate a first virtual viewpoint image representing the view from the virtual camera determined in S807. At this time, a virtual viewpoint image including the instruction model may be generated, similar to S805 described above. In this case, the coach can check how the instruction model will be displayed on the first virtual viewpoint image before showing it to the second user.

[0085] In S809, the overhead view image generated in S803 and the first virtual viewpoint image generated in S808 are transmitted to the first user terminal 100A via the communication unit 204. In addition, the second virtual viewpoint image including the instruction model generated in S805 is transmitted to both the first user terminal 100A and the second user terminal 100B via the communication unit 204. On the first user terminal 100A, the received overhead view image, the first virtual viewpoint image, and the second virtual viewpoint image including the instruction model are displayed in predetermined image areas on the GUI. On the second user terminal 100B, the received second virtual viewpoint image including the instruction model is displayed.

[0086] In S810, the next process to be executed is determined by whether or not the first user has received an input from the first user terminal 100A to cancel screen sharing (in this embodiment, the operation of pressing the delete button 410). If no operation to cancel screen sharing has been received from the first user terminal 100A, this process is exited and the system returns to the flowchart in Figure 6. On the other hand, if an operation to cancel screen sharing has been received, S811 is executed. In S811, the sharing processing unit 306 clears the sharing information held in RAM 202. After clearing, the system returns to the flowchart in Figure 6.

[0087] The above is a description of the flowchart when screen sharing is being accepted. Through this process, while the first user is accepting screen sharing requests, a virtual viewpoint image is generated and output according to the playback section included in the shared information, and the virtual viewpoint image is looped and played on both the first user terminal 100A and the second user terminal 100B. As a result, once the virtual viewpoint image is generated, the user can repeatedly view the same scene without any further action required.

[0088] Through this series of processes, for example, a coach can force players to view virtual viewpoint images from a perspective of their choosing, including instructions for the players. Note that for both the first and second virtual viewpoint images, processes such as playback, pause, and changing the playback speed are executed as needed via interrupt processing during the flowcharts shown in Figures 6 to 8 above.

[0089] <Variation> While accepting a screen sharing operation, the second viewpoint determination unit 308 may determine the position and orientation of the virtual camera based on the display position of the instruction model included in the shared information. For example, by determining the position and orientation of the virtual camera so that the instruction model placed at a specified position in the virtual space is at a predetermined position in the virtual viewpoint image (e.g., the center of the screen or the right corner of the screen), a virtual viewpoint image in which the instruction content is displayed in a position easily recognizable by the second user can be easily output.

[0090] In the above-described embodiment, one image processing device 200 generated and output both the first virtual viewpoint image and the second virtual viewpoint image. However, for example, two image processing devices 200 may be provided, each generating and outputting the first and second virtual viewpoint images. In this case, the two image processing devices 200 communicate with each other to share information, and synchronized output control of the virtual viewpoint images handled by each device is performed. Alternatively, for example, the first user terminal 100A may also have the functions of the image processing device 200.

[0091] When the first user selects mark 411 on the overhead view image, it is not necessary to immediately force a switch to the virtual viewpoint image based on the shared information related to that selection. For example, after the first user selects mark 411, a message indicating the change in viewpoint may be superimposed for a few seconds within the second virtual viewpoint image that the second user is viewing before the forced switch is performed. This reduces confusion caused by a sudden change in viewpoint while viewing the virtual viewpoint image. Alternatively, the forced switching of the virtual viewpoint image by the first user's screen sharing operation may be limited, for example, only while the selection operation of mark 411 is continuing, and after the first user stops the selection operation, the second user may be allowed to control the viewpoint. In this case, the second user wearing the HMD can reduce VR sickness caused by unintended viewpoint movements.

[0092] In the above-described embodiment, the display position of the instruction model and the acceptance of screen sharing operations were performed based on input operations using a mouse or the like on the overhead image, but the system is not limited to this. For example, the display position of the instruction model may be determined by the first user directly inputting three-dimensional coordinate values ​​on the GUI. Also, the acceptance of screen sharing operations may be performed by displaying a list of saved sharing information as a pull-down menu on the GUI, and the first user selecting from the options in the list.

[0093] In the above embodiment, the second virtual viewpoint image generated while the screen sharing operation is being accepted will include the instructions given by the first user. However, the range in which the instructions are displayed on the screen may be set separately. In this case, for example, after setting the start / end timecodes that control the playback section on the seek bar 404, the start / end timecodes for displaying the instructions within that range are determined. This allows, for example, a coach to have a player view the virtual viewpoint image of the playback section they want to show for instruction, while displaying the instructions only in a certain section, thereby making the player more aware of the instructions. Furthermore, multiple display positions for the instructions may be specified in association with timecodes, so that the instructions move from the start to the end of the second virtual viewpoint image. This allows, for example, if an object that the player wants to draw attention to (for example, a specific player on the opposing team) moves within the playback section, the instructions can be made to follow that object. Alternatively, object recognition / following technology may be used to specify, for example, a specific player that the coach wants to follow, so that the instructions move in accordance with that specific player.

[0094] Furthermore, it may be possible to accept changes to the shared information while screen sharing operations are being accepted, and to generate and output a second virtual viewpoint image that reflects the changed content. Figure 9 is an example of GUI400'' in this modified example. A stylus 900 has been added to GUI400'' in Figure 4(b) above. A new string "Attention here" has been entered into the text input field 407, and the corresponding instruction model (text block 911) is displayed on the first virtual viewpoint image in image area 401. Now, suppose the first user uses the stylus 900 to select mark 901 on the overhead image and drags it to the position of mark 902. Then, in conjunction with the movement of the mark, text block 911 moves to text block 912. Such changes and movements of instruction models are performed at S805 in the flow of Figure 8 above. Specifically, a text block 912 of the newly entered string is generated, and the xy plane indicated by the mark after the movement operation is displayed. This is done by placing the text block 912 at the two-dimensional coordinate position and rendering it. Note that the movement of the mark is not limited to a stylus and can be done with a mouse or other device. In addition, the position and orientation of the virtual camera may change in conjunction with the movement of the instruction model. For example, the virtual camera's position may be fixed while its orientation changes as the instruction model moves to maintain its position in the center of the second virtual viewpoint image. This allows the first user to draw the second user's attention to specific areas by moving the instruction model in accordance with the movement of objects in the virtual viewpoint image.

[0095] <Other Embodiments> The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0096] Furthermore, this disclosure includes the following configurations and methods.

[0097] [Configuration 1] A reception means for receiving operation input from the first user, Image generation means that generates a virtual viewpoint image based on multiple images captured by multiple imaging devices, which reflects instructions from the first user to the second user based on the operation input of the first user, An output control means that outputs the generated virtual viewpoint image to a user terminal used by the second user based on the operation input of the first user, An image processing apparatus characterized by having

[0098] [Configuration 2] The image processing apparatus according to Configuration 1, wherein the image generation means places a 3D model representing the instructions in a virtual space and performs rendering processing according to the virtual viewpoint specified by the first user to generate the virtual viewpoint image.

[0099] [Configuration 3] The output control means further outputs the generated virtual viewpoint image to the user terminal used by the first user. The virtual viewpoint image is displayed on both the user terminal used by the first user and the user terminal used by the second user. The image processing apparatus according to configuration 2, characterized in that...

[0100] [Structure 4] The image generation means generates the virtual viewpoint image based on a screen sharing operation from the first user, which is received as an operation input from the first user. The image processing apparatus according to configuration 3, characterized in that

[0101] [Composition 5] The system further includes storage means for storing information including the content of instructions from the first user to the second user received as operation input from the first user, the position of the 3D model in the virtual space, and the position and orientation of the virtual camera corresponding to the virtual viewpoint. The aforementioned screen sharing operation is an operation to select the saved information. The image processing apparatus according to configuration 4, characterized by the following:

[0102] [Composition 6] The reception means receives the operation input of the first user via a GUI (Graphical User Interface), The GUI displays a mark corresponding to the information saved by the saving means. The image processing apparatus according to configuration 5, characterized by the features described herein.

[0103] [Composition 7] The GUI includes an overhead image corresponding to a viewpoint that provides an overview of the area in which the multiple imaging devices perform imaging. The aforementioned mark is displayed on the overhead image. The image processing apparatus according to configuration 6, characterized by the features described therein.

[0104] [Structure 8] The image processing apparatus according to configuration 7, characterized in that the position of the mark displayed on the overhead image indicates the position of the 3D model in the virtual space.

[0105] [Composition 9] The image processing apparatus according to any one of configurations 6 to 8, wherein the GUI includes an input field for the first user to input a string of characters as the content of instructions from the first user to the second user.

[0106] [Configuration 10] The image generation means generates another virtual viewpoint image based on the operation input of the first user, using a plurality of images captured by the plurality of imaging devices. The output control means does not output the generated alternative virtual viewpoint image to the user terminal used by the second user, but outputs it to the user terminal used by the first user. The GUI displays the other virtual viewpoint image. The image processing apparatus according to configuration 6, characterized by the features described therein.

[0107] [Composition 11] The image processing apparatus according to configuration 10, characterized in that the position and orientation of the virtual camera corresponding to the virtual viewpoint are determined based on the operation input of the first user using the other virtual viewpoint image on the GUI.

[0108] [Composition 12] The image processing apparatus according to configuration 6, characterized in that the position and orientation of the virtual camera corresponding to the virtual viewpoint are determined based on the position of the 3D model in the virtual space included in the selected information.

[0109] [Composition 13] The aforementioned virtual viewpoint image and the aforementioned other virtual viewpoint image are videos. The information stored by the storage means further includes playback sections of the virtual viewpoint image and the other virtual viewpoint image, which are received as operation input from the first user. The GUI includes UI elements for the first user to input the playback interval, The image processing apparatus according to configuration 10, characterized in that...

[0110] [Method 1] A reception step for receiving input from the first user, An image generation step that generates a virtual viewpoint image based on multiple images captured by multiple imaging devices, which reflects instructions from the first user to the second user based on the operation input of the first user, An output control step which outputs the generated virtual viewpoint image to a user terminal used by the second user based on the operation input of the first user, An image processing method characterized by including

[0111] [Composition 14] A program for causing a computer to function as an image processing device as described in any one of configurations 1 to 13.

Claims

1. A reception means for receiving operation input from the first user, Image generation means that generates a virtual viewpoint image based on multiple images captured by multiple imaging devices, which reflects instructions from the first user to the second user based on the operation input of the first user, Output control means for outputting the generated virtual viewpoint image to a user terminal used by the second user based on the operation input of the first user, An image processing apparatus characterized by having

2. The image processing apparatus according to claim 1, characterized in that the image generation means places a 3D model representing the instructions in a virtual space and generates the virtual viewpoint image by performing rendering processing according to the virtual viewpoint specified by the first user.

3. The output control means further outputs the generated virtual viewpoint image to the user terminal used by the first user. The virtual viewpoint image is displayed on both the user terminal used by the first user and the user terminal used by the second user. The image processing apparatus according to claim 2.

4. The image generation means generates the virtual viewpoint image based on a screen sharing operation from the first user, which is received as an operation input from the first user. The image processing apparatus according to claim 3.

5. The system further includes storage means for storing information including the content of instructions from the first user to the second user received as operation input from the first user, the position of the 3D model in the virtual space, and the position and orientation of the virtual camera corresponding to the virtual viewpoint. The aforementioned screen sharing operation is an operation to select the saved information. The image processing apparatus according to claim 4, characterized in that it does so.

6. The receiving means receives the operation input of the first user via a GUI (Graphical User Interface), The GUI displays a mark corresponding to the information stored by the storage means. The image processing apparatus according to feature 5.

7. The GUI includes an overhead image corresponding to a viewpoint that provides an overview of the area in which the multiple imaging devices perform imaging. The aforementioned mark is displayed on the overhead image. The image processing apparatus according to claim 6.

8. The image processing apparatus according to claim 7, characterized in that the position of the mark displayed on the overhead image indicates the position of the 3D model in the virtual space.

9. The image processing apparatus according to claim 6, wherein the GUI includes an input field for the first user to input a string of characters as the content of an instruction from the first user to the second user.

10. The image generation means generates another virtual viewpoint image based on the operation input of the first user, using a plurality of images captured by the plurality of imaging devices. The output control means does not output the generated alternative virtual viewpoint image to the user terminal used by the second user, but outputs it to the user terminal used by the first user. The GUI displays the other virtual viewpoint image. The image processing apparatus according to claim 6.

11. The image processing apparatus according to claim 10, characterized in that the position and orientation of the virtual camera corresponding to the virtual viewpoint are determined based on the operation input of the first user using the other virtual viewpoint image on the GUI.

12. The image processing apparatus according to claim 6, characterized in that the position and orientation of the virtual camera corresponding to the virtual viewpoint are determined based on the position of the 3D model in the virtual space included in the selected information.

13. The aforementioned virtual viewpoint image and the aforementioned other virtual viewpoint image are videos. The information stored by the storage means further includes playback sections of the virtual viewpoint image and the other virtual viewpoint image, which are received as operation input from the first user. The GUI includes UI elements for the first user to input the playback section, The image processing apparatus according to feature 10.

14. A reception step for receiving input from the first user, An image generation step that generates a virtual viewpoint image based on multiple images captured by multiple imaging devices, which reflects instructions from the first user to the second user based on the operation input of the first user, An output control step which outputs the generated virtual viewpoint image to a user terminal used by the second user based on the operation input of the first user, An image processing method characterized by including [a certain element].

15. A program for causing a computer to perform the image processing method described in claim 14.