Image processing system, image processing method, and computer program
The image processing system generates and displays virtual viewpoint images from multiple camera perspectives, addressing the lack of compelling digital content by integrating flexible viewpoint selection and display control, thereby enhancing user engagement.
Patent Information
- Application Number
- JP2023038750
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-04-27
- Filing Date
- 2023-03-13
- Publication Date
- 2025-08-28
- Estimated Expiration
- 2043-03-13
AI Technical Summary
Existing technologies fail to provide compelling digital content that includes virtual viewpoint images and other images effectively.
An image processing system that generates and displays virtual viewpoint images based on images captured by multiple cameras from different directions, allowing for flexible viewpoint selection and integration with other images, using a display control mechanism to manage the display of these images.
Enables the creation of attractive digital content that includes virtual viewpoint images and other images, enhancing user interaction and engagement.
Smart Images

Figure 0007730851000001 
Figure 0007730851000002 
Figure 0007730851000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image processing system, an image processing method, a computer program, and the like. [Background technology]
[0002] A technology that generates a virtual viewpoint image from a specified virtual viewpoint using a plurality of images captured by a plurality of imaging devices has attracted attention. Patent Document 1 describes a method of capturing images of a subject using a plurality of imaging devices installed at different positions, and generating a virtual viewpoint image using a three-dimensional shape of the subject estimated from the captured images. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-45920 Summary of the Invention [Problem to be solved by the invention]
[0004] However, it has not been possible to provide attractive digital content that includes virtual viewpoint images and other images.
[0005] The present disclosure aims to provide an image processing system for providing a display technique for compelling digital content, including virtual viewpoint images and other images. [Means for solving the problem]
[0006] An image processing system according to one embodiment of the present disclosure includes: a specifying means for specifying a virtual viewpoint image associated with a first surface of a three-dimensional digital content, the virtual viewpoint image being generated based on a plurality of images captured by a plurality of imaging devices and a virtual viewpoint, and an image of a viewpoint different from the virtual viewpoint corresponding to the virtual viewpoint image associated with a second surface of the digital content; a display control means for controlling the display of an image corresponding to the virtual viewpoint image and an image corresponding to an image of a viewpoint different from the virtual viewpoint in a display area; an acquisition means for acquiring input information for selecting the display area; and The display control means displays an image corresponding to the selected display area in a selected image display area that is a display area different from the selected display area, based on the input information acquired by the acquisition means. of It is characterized by: [Effects of the Invention]
[0007] According to the present disclosure, it is possible to display attractive digital content that includes virtual viewpoint images and other images. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a diagram showing an example of the device configuration of an image processing system 100 according to a first embodiment. [Figure 2] 1 is a diagram showing a hardware configuration of an image processing system 100 according to a first embodiment. [Figure 3] 1 is a flowchart illustrating the operation flow of the image processing system 100 according to the first embodiment. [Figure 4] 4A to 4C are diagrams showing examples of stereoscopic images as content generated by the content generation unit 4 in the fourth embodiment. [Figure 5] 10 is a flowchart illustrating the operation flow of the image processing system 100 according to the fifth embodiment. [Figure 6] 10 is a flowchart illustrating the operation flow of the image processing system 100 according to the third embodiment. [Figure 7] This is a continuation of the flowchart in Figure 6. [Figure 8]8 is a continuation of the flowchart of FIG. 6 and FIG. 7. [Figure 9] FIG. 10 is a diagram illustrating an example of a graphical user interface displayed on a user device according to the fourth embodiment. [Figure 10] 10 is a flowchart illustrating the operation flow of the fourth embodiment. [Figure 11] FIG. 13 is a diagram illustrating an example of a graphical user interface displayed on a user device according to the fifth embodiment. [Figure 12] 13 is a flowchart illustrating the operation flow of the fifth embodiment. [Figure 13] FIG. 20 is a diagram showing an example of a graphical user interface displayed on a user device according to the sixth embodiment. [Figure 14] 13 is a flowchart illustrating the operation flow of the sixth embodiment. [Figure 15] 13 is a diagram illustrating each surface of a three-dimensional digital content according to a seventh embodiment. FIG. [Figure 16] FIG. 20 is a diagram illustrating the shooting direction of a player according to the seventh embodiment. [Figure 17] FIG. 20 is a diagram showing an example of a three-dimensional digital content generated by the content generation unit 4 in the seventh embodiment. [Figure 18] FIG. 13 is a diagram showing an example of the device configuration of an image processing system 101 according to a seventh embodiment. [Figure 19] 13 is a flowchart for explaining the operation flow of the image processing system 101 according to the seventh embodiment. [Figure 20] 13 is a flowchart for explaining the operation flow of the image processing system 101 according to the eighth embodiment. [Figure 21] FIG. 13 is a diagram showing the system configuration of an image processing system 103 according to a ninth embodiment. [Figure 22] FIG. 13 is a diagram showing the flow of data transmission in the ninth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. However, the present disclosure is not limited to the following embodiments. In each drawing, the same members or elements are designated by the same reference numerals, and duplicate descriptions will be omitted or simplified.
[0010] (Embodiment 1) The image processing system of the first embodiment generates a virtual viewpoint image seen from a virtual viewpoint based on captured images captured by a plurality of imaging devices (cameras) from different directions, the state of the imaging devices, and a specified virtual viewpoint. The virtual viewpoint image is then displayed on the surface of a virtual three-dimensional image. Note that the imaging devices may have not only cameras but also functional units for performing image processing. Furthermore, the imaging devices may have sensors for acquiring distance information in addition to cameras.
[0011] The multiple cameras capture images of the imaging area from multiple directions. The imaging area may be, for example, an area surrounded by a stadium field at an arbitrary height. The imaging area may correspond to the three-dimensional space from which the three-dimensional shape of the subject is estimated. The three-dimensional space may be the entire imaging area or a part of it. The imaging area may also be a concert venue, an imaging studio, etc.
[0012] The multiple cameras are installed in different positions and directions (attitudes) surrounding the imaging area, and capture images synchronously. Note that the multiple cameras do not need to be installed around the entire circumference of the imaging area, and may be installed in only a part of the imaging area depending on installation location restrictions, etc. The number of cameras is not limited, and for example, if the imaging area is a rugby stadium, several tens to several hundreds of cameras may be installed around the stadium.
[0013] The multiple cameras may also include cameras with different angles of view, such as a telephoto camera and a wide-angle camera. For example, using a telephoto camera to capture high-resolution images of players can improve the resolution of the generated virtual viewpoint image. In ball games, the ball moves over a wide range, so using a wide-angle camera for capture can reduce the number of cameras. Combining the imaging areas of a wide-angle camera and a telephoto camera for capture can improve the degree of freedom in installation position. The cameras are synchronized to a common time, and capture time information is added to each captured image frame.
[0014] A virtual viewpoint image, also called a free viewpoint image, allows an operator to monitor an image corresponding to a viewpoint freely (arbitrarily) designated by the operator. For example, a virtual viewpoint image includes an image corresponding to a viewpoint selected by the operator from a limited number of viewpoint candidates. The virtual viewpoint may be designated by an operator or automatically by AI based on the results of image analysis, etc. The virtual viewpoint image may be a video or a still image.
[0015] The virtual viewpoint information used to generate a virtual viewpoint image is information including the position and direction (posture) of the virtual viewpoint, as well as the angle of view (focal length), etc. Specifically, the virtual viewpoint information includes parameters representing the three-dimensional position of the virtual viewpoint, parameters representing the direction (line of sight) from the virtual viewpoint in the pan, tilt, and roll directions, focal length information, etc. However, the contents of the virtual viewpoint information are not limited to those described above.
[0016] The virtual viewpoint information may also have parameters for each of a plurality of frames. In other words, the virtual viewpoint information may have parameters corresponding to each of a plurality of frames constituting the video of the virtual viewpoint image, and may be information indicating the position and orientation of the virtual viewpoint at each of a plurality of consecutive points in time.
[0017] A virtual viewpoint image is generated, for example, by the following method. First, multiple camera images are acquired by capturing images from different directions using cameras. Next, a foreground image is acquired from the multiple camera images, extracting a foreground area corresponding to a subject such as a person or a ball, and a background image is acquired from the background area other than the foreground area. The foreground image and background image contain texture information (such as color information).
[0018] Then, a foreground model representing the three-dimensional shape of the subject and texture data for coloring the foreground model are generated based on the foreground image. Also, texture data for coloring a background model representing the three-dimensional shape of the background, such as a stadium, is generated based on the background image. Then, the texture data is mapped onto the foreground model and background model, and rendering is performed according to the virtual viewpoint indicated by the virtual viewpoint information, thereby generating a virtual viewpoint image.
[0019] However, the method for generating the virtual viewpoint image is not limited to this, and various methods can be used, such as a method for generating the virtual viewpoint image by projective transformation of the captured image without using a foreground or background model.
[0020] A foreground image is an image in which a subject area (foreground area) is extracted from an image captured by a camera. A subject extracted as a foreground area refers to a dynamic subject (moving object) that moves (its absolute position and shape may change) when images are captured from the same direction in a time series. For example, in a sport, the subject includes people such as players and referees on the field where the sport is being played, and in a ball game, it includes not only people but also the ball. In concerts and entertainment events, foreground subjects include singers, musicians, performers, and presenters.
[0021] A background image is an image of at least an area (background area) different from the foreground subject. Specifically, a background image is an image in which the foreground subject has been removed from the captured image. The background also refers to an object that remains stationary or nearly stationary when images are captured from the same direction in chronological order.
[0022] Such imaging objects include, for example, a stage for a concert, a stadium where an event such as a sport is held, a structure such as a goal used in a ball game, a field, etc. However, the background is at least an area different from the subject, which is the foreground. Note that the imaging objects may include other objects in addition to the subject and background.
[0023] Fig. 1 is a diagram showing an image processing system 100 according to this embodiment. Some of the functional blocks shown in Fig. 1 are realized by causing a computer included in the image processing system 100 to execute a computer program stored in a memory serving as a storage medium. However, some or all of these functions may be realized by hardware. Examples of hardware that can be used include a dedicated circuit (ASIC) and a processor (reconfigurable processor, DSP).
[0024] Furthermore, the individual functional blocks of the image processing system 100 do not have to be built into the same housing, and may be configured as separate devices connected to each other via signal paths. The image processing system 100 is connected to multiple cameras. The image processing system 100 also has a shape estimation unit 2, an image generation unit 3, a content generation unit 4, a storage unit 5, a display unit 115, an operation unit 116, etc. The shape estimation unit 2 is connected to multiple cameras 1 and the image generation unit 3, and the display unit 115 is connected to the image generation unit 3. The individual functional blocks may be implemented in separate devices, or all or some of the functional blocks may be implemented in the same device.
[0025] Multiple cameras 1 are placed at different locations around a stage for a concert or other event, a stadium for a sporting event, a structure such as a goal used in a ball game, or a field, and each camera captures images from a different viewpoint. Each camera has an identification number (camera number) to identify the camera. Camera 1 may also include other functions, such as a function to extract a foreground image from captured images, and hardware (circuits, devices, etc.) that realizes those functions. The camera number may be set based on the installation location of camera 1, or may be set based on other criteria.
[0026] The image processing system 100 may be located within the venue where the camera 1 is installed, or may be located outside the venue, for example, at a broadcasting station. The image processing system 100 is connected to the camera 1 via a network.
[0027] The shape estimation unit 2 acquires images from multiple cameras 1. Then, the shape estimation unit 2 estimates the three-dimensional shape of the subject based on the images acquired from the multiple cameras 1. Specifically, the shape estimation unit 2 generates three-dimensional shape data expressed using a known representation method. The three-dimensional shape data may be point cloud data made up of points, mesh data made up of polygons, or voxel data made up of voxels.
[0028] The image generation unit 3 acquires information indicating the position and posture of the three-dimensional shape data of the subject from the shape estimation unit 2, and can generate a virtual viewpoint image including the two-dimensional shape of the subject that is expressed when the three-dimensional shape of the subject is viewed from a virtual viewpoint. In order to generate the virtual viewpoint image, the image generation unit 3 can also receive designation of virtual viewpoint information (such as the position of the virtual viewpoint and the line of sight from the virtual viewpoint) from an operator, and generate the virtual viewpoint image based on the virtual viewpoint information. Here, the image generation unit 3 functions as a virtual viewpoint image generation means that generates a virtual viewpoint image based on multiple images obtained from multiple cameras.
[0029] The virtual viewpoint image is sent to the content generation unit 4, and the content generation unit 4 generates, for example, a three-dimensional digital content as described later. The digital content including the virtual viewpoint image generated by the content generation unit 4 is output to the display unit 115.
[0030] The content generation unit 4 can also directly receive images from multiple cameras and supply the image from each camera to the display unit 115. Furthermore, based on an instruction from the operation unit 116, it can also switch on which surface of the virtual stereoscopic image the image from each camera and the virtual viewpoint image are to be displayed.
[0031] The display unit 115 is configured with, for example, a liquid crystal display or an LED, and acquires and displays digital content including virtual viewpoint images from the content generation unit 4. It also displays a GUI (Graphical User Interface) and the like for an operator to operate each camera 1.
[0032] The operation unit 116 is composed of a joystick, a jog dial, a touch panel, a keyboard, a mouse, etc., and is used by the operator to operate the camera 1, etc. The operation unit 116 is also used by the operator to select an image to be displayed on the surface of the digital content (stereoscopic image) generated by the content generation unit 4. Furthermore, the operation unit 116 can specify the position and posture of the virtual viewpoint for generating a virtual viewpoint image in the image generation unit 3.
[0033] The position and orientation of the virtual viewpoint may be directly specified on the screen by an operator's operation instruction, or when a predetermined subject is specified on the screen by an operator's operation instruction, the predetermined subject may be image-recognized and tracked, and virtual viewpoint information from the subject or virtual viewpoint information from positions around the subject in an arc shape centered on the subject may be automatically specified.
[0034] Furthermore, it is also possible to automatically specify virtual viewpoint information from the subject or from positions around the subject in an arc, by image recognition of a subject that meets conditions previously specified by an operator's operation instructions. In this case, the specified conditions may include, for example, the name of a specific athlete, the shooter, the person who made the fine play, the position of the ball, etc.
[0035] The storage unit 5 includes a memory for storing the digital content generated by the content generation unit 4, the virtual viewpoint images, the camera images, etc. The storage unit 5 may also have a removable recording medium. The removable recording medium may store, for example, a plurality of camera images captured at other venues or other sporting scenes, a virtual viewpoint image generated using those images, or digital content generated by combining those images.
[0036] The storage unit 5 may also be configured to store a plurality of camera images downloaded from an external server or the like via a network, virtual viewpoint images generated using those images, digital content generated by combining those images, etc. Furthermore, those camera images, virtual viewpoint images, digital content, etc. may be created by a third party.
[0037] FIG. 2 is a diagram showing the hardware configuration of the image processing system 100 according to the first embodiment, and the hardware configuration of the image processing system 100 will be described with reference to FIG.
[0038] The image processing system 100 includes a CPU 111, a ROM 112, a RAM 113, an auxiliary storage device 114, a display unit 115, an operation unit 116, a communication I / F 117, and a bus 118. The CPU 111 controls the entire image processing system 100 using computer programs stored in the ROM 112, the RAM 113, the auxiliary storage device 114, etc., thereby realizing each functional block of the image processing system shown in FIG.
[0039] The RAM 113 temporarily stores computer programs and data supplied from the auxiliary storage device 114, and data supplied from the outside via the communication I / F 117. The auxiliary storage device 114 is configured, for example, with a hard disk drive, and stores various data such as image data, audio data, and digital content including virtual viewpoint images from the content generation unit 4.
[0040] As described above, the display unit 115 displays digital content including virtual viewpoint images, GUIs, etc. As described above, the operation unit 116 receives operation inputs from the operator and inputs various instructions to the CPU 111. The CPU 111 operates as a display control unit that controls the display unit 115 and an operation control unit that controls the operation unit 116.
[0041] The communication I / F 117 is used for communication with devices external to the image processing system 100 (for example, the camera 1 or an external server). For example, if the image processing system 100 is connected to an external device via a wired connection, a communication cable is connected to the communication I / F 117. If the image processing system 100 has a function for wireless communication with an external device, the communication I / F 117 is equipped with an antenna. The bus 118 connects each unit of the image processing system 100 to transmit information.
[0042] In this embodiment, an example is shown in which the display unit 115 and the operation unit 116 are included inside the image processing system 100, but at least one of the display unit 115 and the operation unit 116 may exist as a separate device outside the image processing system 100. In addition, the image processing system 100 may be in the form of, for example, a PC terminal.
[0043] Fig. 3 is a flowchart for explaining the operation flow of the image processing system 100 of the embodiment 1. Also, Figs. 4(A) to 4(C) are diagrams showing examples of three-dimensional digital content generated by the content generation unit 4 in the embodiment 1.
[0044] The CPU 111 serving as a computer in the image processing system 100 executes a computer program stored in a memory such as the ROM 112 or the auxiliary storage device 114, thereby performing the operations of the steps in the flowchart of FIG.
[0045] In this embodiment, the image processing system 100 is installed in a broadcasting station or the like, and may produce and broadcast cubic digital content 200 as shown in Fig. 4(A), or may provide it via the Internet. In this case, an NFT can be assigned to the digital content 200.
[0046] That is, in order to increase asset value, scarcity can be achieved by, for example, limiting the quantity of content distributed and managing it with a serial number. NFT stands for Non-fungible Token, and is a token issued and circulated on a blockchain. Examples of NFT formats include token standards called ERC-721 and ERC-1155. Tokens are usually stored in association with a wallet managed by the operator.
[0047] In step S31, CPU 111 associates the main camera image (first image) with first surface 201 of digital content 200, which has a three-dimensional shape, as shown in FIG. 4(A), for example. The main camera image associated with first surface 201 may be displayed for confirmation by the operator. Furthermore, as shown in FIG. 4(A), if the line of sight from a virtual viewpoint of the digital content (specifically, the direction perpendicular to the paper surface of FIG. 4(A)) is not parallel to the normal direction of first surface 201, the following display may be performed. That is, the main camera image displayed on first surface 201 may be generated by projective transformation according to the angle of the normal direction of first surface 201 relative to the display surface of the digital content. Here, the main camera image (main image, first image) is an image selected for TV broadcasting or the like from among multiple images acquired by multiple cameras installed at a sports venue. The main image is an image that includes a predetermined subject within its angle of view. The main camera image does not have to be captured by a camera installed at the sports venue. For example, the image may be captured by a handheld camera brought by a cameraman, or may be captured by a camera brought by a spectator in the venue or an electronic device such as a smartphone equipped with a camera. The main camera image may be one of the multiple cameras used to generate the virtual viewpoint image, or may be a camera not included in the multiple cameras.
[0048] An operator at a broadcasting station or the like sequentially selects which camera's image to broadcast or distribute over the Internet as the main image using the operation unit 116. For example, when broadcasting or distributing the moment a goal is scored, an image from a camera near the goal is often broadcast as the main image.
[0049] In this embodiment, as shown in Figures 4(A) to 4(C), the surface seen on the left side is the first surface, the surface seen on the right side is the second surface, and the surface seen on the upper side is the third surface. However, this is not limited to this. Which surfaces are the first to third surfaces can be arbitrarily set in advance.
[0050] In step S32, the content generation unit 4 associates data such as the name of the player who shot the goal, the name of the team he belongs to, and the final game result of the game in which he scored, as associated data with the third side 203 of the digital content 200. The associated data associated with the third side 203 may be displayed for confirmation by the operator. When an NFT is granted, data indicating its rarity, such as the number of issuances, may be displayed as associated data on the third side 203. The number of issuances may be determined by an operator who generates the digital content using an image generation system, or may be determined automatically by the image generation system.
[0051] In step S33, image generation unit 3 acquires, from among images from multiple cameras 1, an image that includes, for example, a goal or a shooter, and whose viewpoint direction differs by a predetermined angle (e.g., 90 degrees) from the viewpoint direction of the camera capturing the main camera image. At this time, since the positions and attitudes of the multiple cameras are known in advance, CPU 111 can determine from which camera an image with a viewpoint direction that differs by a predetermined angle from the main camera image as described above can be acquired. Note that, although the expression "viewpoint of an image" is sometimes used below, this refers to the viewpoint of the camera capturing the image or a virtual viewpoint specified for generating the image.
[0052] Alternatively, in step S33, a virtual viewpoint image from a virtual viewpoint that includes the image-recognized subject and is a predetermined virtual viewpoint (for example, a viewpoint that differs by 90 degrees as described above) may be acquired in the image generation unit 3. In this case, the predetermined virtual viewpoint (a viewpoint that differs by 90 degrees as described above, i.e., a posture that differs) may be specified for the image generation unit 3, and the virtual viewpoint image may be generated and acquired.
[0053] Alternatively, virtual viewpoint images from a plurality of viewpoints may be generated in advance by the image generation unit 3, and the corresponding image may be selected from among them. In this embodiment, an image with a viewpoint that differs by a predetermined angle from the main camera image is an image with a viewpoint that differs by, for example, 90 degrees, but this angle can be set in advance.
[0054] Furthermore, the virtual viewpoint image may be an image corresponding to a virtual viewpoint specified based on the orientation of a subject included in the main camera image (for example, the orientation of the face or body of a person). Note that if there are multiple subjects included in the main camera image, the virtual viewpoint may be set for one of the subjects, or virtual viewpoints may be set for multiple subjects.
[0055] Although the above description has been given of an example in which a viewpoint at a predetermined angle is selected with respect to the main image, it is also possible to select and acquire a virtual viewpoint image from a predetermined viewpoint, such as the subject viewpoint, a viewpoint from behind the subject, or one of the virtual viewpoints located on an arc centered on the subject.
[0056] The subject viewpoint is a virtual viewpoint in which the position of the subject is the position of the virtual viewpoint and the direction of the subject is the line of sight from the virtual viewpoint. For example, when a person is the subject, the subject viewpoint is a viewpoint in which the position of the person's face is the position of the virtual viewpoint and the direction of the person's face is the line of sight from the virtual viewpoint. Alternatively, the line of sight of the person may be the line of sight from the virtual viewpoint.
[0057] A viewpoint from behind a subject is a virtual viewpoint in which a position behind the subject at a predetermined distance is set as the virtual viewpoint, and the direction from that position toward the position of the subject is set as the line of sight from the virtual viewpoint. Alternatively, the line of sight from the virtual viewpoint may be determined according to the orientation of the subject. For example, when a person is set as the subject, a viewpoint from behind the subject is a virtual viewpoint in which a position behind and above the person's back at a predetermined distance is set as the line of sight from the virtual viewpoint, and the direction of the person's face is set as the line of sight from the virtual viewpoint.
[0058] A virtual viewpoint within a position on an arc centered on the subject is a virtual viewpoint that is set at a position on a sphere defined by a predetermined radius centered on the position of the subject, and the direction from that position toward the position of the subject is the line of sight from the virtual viewpoint.
[0059] For example, when a person is the subject, the virtual viewpoint is a position on a sphere defined by a predetermined radius centered on the position of the person, and the direction from that position toward the position of the subject is the line of sight from the virtual viewpoint.
[0060] Step S33 thus functions as a virtual viewpoint image generating step for acquiring a virtual viewpoint image of a viewpoint having a predetermined relationship with the first image and using it as the second image. Note that the virtual viewpoint image of a viewpoint having a predetermined relationship with the first image is a virtual viewpoint image captured at the same time (image capturing timing) as the viewpoint of the first image. In this embodiment, the viewpoint having a predetermined relationship with the first image is a viewpoint that has a predetermined angular relationship or a predetermined positional relationship with the viewpoint of the first image, as described above.
[0061] Then, in step S34, CPU 111 associates the second image with second surface 202 of digital content 200. Note that the second image may be displayed for the operator's confirmation. Note that at this time, the main image associated with first surface 201 and the second image associated with second surface 202 are synchronized so that they are images captured at the same time (time code), as described above. In this way, steps S31 to S34 associate the first image with the first surface of the three-dimensional digital content, which will be described later, and associate a virtual viewpoint image of a virtual viewpoint having a predetermined relationship with the first image with second surface 202. Also, steps S31 to S34 function as a content generation step (content generation means).
[0062] Next, in step S35, the CPU 111 determines whether an operation to change the viewpoint of the second image displayed on the second screen 202 has been performed via the operation unit 116. That is, the operator may change the viewpoint of the second image to be displayed on the second screen while watching the ever-changing sports scene, for example, by selecting a camera image with a desired viewpoint from among the multiple cameras 1.
[0063] Alternatively, a virtual viewpoint image from a desired viewpoint may be acquired by specifying the desired viewpoint from among the virtual viewpoints to the image generator 3. In step S35, if such a viewpoint change operation is performed, the result becomes Yes, and the process proceeds to step S36.
[0064] In step S36, CPU 111 selects a viewpoint image after the viewpoint has been changed from among the multiple cameras 1, or acquires a virtual viewpoint image after the viewpoint has been changed from image generator 3. The virtual viewpoint image may be acquired by acquiring a pre-generated virtual viewpoint image, or by generating a new virtual viewpoint image based on the changed viewpoint. The acquired image is then designated as the second image, and the process proceeds to step S34, where it is associated with the second surface. In this state, the first image, the second image, and the associated data are associated with the first to third surfaces of digital content 200, respectively, on display unit 115. The operator may confirm, through the display, the state in which the first image, the second image, and the associated data are associated with the first to third surfaces of digital content 200, respectively. In this case, the surface numbers may also be displayed to indicate which surfaces are the first to third surfaces.
[0065] If there is no viewpoint change in step S35, the process proceeds to step S37, where the CPU 111 determines whether or not to assign an NFT to the digital content 200. To do this, for example, a display image (GUI) asking whether or not to assign an NFT to the digital content 200 is displayed on the display unit 115. If the operator selects to assign an NFT, this is determined, and the process proceeds to step S38, where the NFT is assigned to the digital content 200, encryption is performed, and then the process proceeds to step S39.
[0066] If the answer in step S37 is No, the process proceeds directly to step S39.
[0067] The digital content 200 in step S37 may be a three-dimensional image having a shape such as that shown in Figures 4(B) and 4(C). In addition, in the case of a polyhedron, it is not limited to a hexahedron as shown in Figure 4(A) and may be, for example, an octahedron.
[0068] In step S39, CPU 111 determines whether or not to end the flow for generating digital content 200 in Fig. 3. If the operator has not operated operation unit 116 to end it, the process returns to step S31 and repeats the above processing, and if it has ended it ends the flow in Fig. 3. Note that even if the operator has not operated operation unit 116 to end it, the process may be automatically ended when a predetermined period of time (e.g., 30 minutes) has elapsed since the last operation of operation unit 116.
[0069] 4(B) and (C) are diagrams showing modified examples of digital content 200, with Fig. 4(B) being a sphere, as compared to the digital content 200 of Fig. 4(A). For example, when viewed from the front of sphere 200, a first image is displayed on first surface 201, which is the left spherical surface, and a second image is displayed on second surface 202, which is the right spherical surface. Furthermore, the above-mentioned associated data is displayed on third surface 203, which is the upper spherical surface.
[0070] Fig. 4(C) is a diagram showing an example in which each plane of the digital content 200 in Fig. 4(A) is curved to a desired curvature. In this way, the digital content in this embodiment may be displayed using a sphere as shown in Fig. 4(B) or (C), or a cube with a spherical surface, to display an image.
[0071] (Embodiment 2) Next, a second embodiment will be described with reference to FIG.
[0072] Fig. 5 is a flowchart for explaining the operation flow of the image processing system 100 of embodiment 2. Note that the operation of each step in the flowchart of Fig. 5 is performed by the CPU 111 serving as the computer of the image processing system 100 executing a computer program stored in a memory such as the ROM 112 or the auxiliary storage device 114.
[0073] In FIG. 5, steps with the same reference numerals as those in FIG. 3 are the same processes, and the explanation thereof will be omitted.
[0074] 5, the CPU 111 acquires a camera image from a viewpoint specified by the operator or a virtual viewpoint image from a virtual viewpoint specified by the operator from the image generation unit 3. The acquired image is then set as the second image. The rest of the process is the same as the flow in FIG. 3.
[0075] In the first embodiment, a second image having a predetermined relationship (different angle) with respect to the main image (first image) is acquired. However, in the second embodiment, the operator selects a desired camera or acquires a virtual viewpoint image of a desired subject from a desired viewpoint to be used as the second image.
[0076] The camera image or virtual viewpoint image selected by the operator in step S51 includes, for example, an image from a bird's-eye view diagonally above the sports venue or an image from a view diagonally below the sports venue, etc. In this way, in the second embodiment, the operator can select the virtual viewpoint image to be displayed on the second screen.
[0077] Furthermore, the virtual viewpoint image selected by the operator in step S51 may be a virtual viewpoint image of a viewpoint from a position distant from the subject, such as a zoomed-out image.
[0078] Additionally, previously generated camera images and virtual viewpoint images generated based on them may be stored in the storage unit 5, and these may be read out and displayed on the first to third screens as the first image, second image, and associated data, respectively.
[0079] Between steps S38 and S39, a step may be inserted in which the CPU 111 automatically switches to a default stereoscopic image display when a predetermined period of time (e.g., 30 minutes) has elapsed since the last operation of the operation unit 116. The default stereoscopic image display may display, for example, a main image on the first screen, associated data on the third screen, and a camera image or a virtual viewpoint image from the viewpoint most frequently used in past statistics on the second screen.
[0080] (Embodiment 3) The third embodiment will be described with reference to Figures 6 to 8. Figure 6 is a flowchart for explaining the operation flow of image processing system 100 of the third embodiment, Figure 7 is a flowchart continuing from Figure 6, and Figure 8 is a flowchart continuing from Figures 6 and 7. Note that the operation of each step in the flowcharts of Figures 6 to 8 is performed by CPU 111, which serves as a computer in image processing system 100, executing a computer program stored in a memory such as ROM 112 or auxiliary storage device 114.
[0081] In the third embodiment, when the operator selects the number of virtual viewpoints from 1 to 3, the display of the first to third planes of the digital content 200 is automatically switched accordingly.
[0082] In step S61, the operator selects the number of virtual viewpoints from among 1 to 3, and this is accepted by the CPU 111. In step S62, the CPU 111 acquires the selected number of virtual viewpoint images from the image generation unit 3.
[0083] At this time, a representative virtual viewpoint is automatically selected. That is, the scene is analyzed, and the virtual viewpoint that is most frequently used for that scene, for example, based on past statistics, is set as the first virtual viewpoint, the next most frequently used virtual viewpoint is set as the second virtual viewpoint, and the next most frequently used virtual viewpoint is set as the third virtual viewpoint. Note that the second virtual viewpoint may be set in advance to differ, for example, by +90° from the first virtual viewpoint, and the third virtual viewpoint may be set in advance to differ, for example, by -90° from the first virtual viewpoint. Here, +90° and -90° are examples, and the present invention is not limited to these angles.
[0084] In step S63, the CPU 111 determines whether the number of selected virtual viewpoints is 1, and if so, proceeds to step S64. In step S64, the CPU 111 acquires a main image from a main camera among the multiple cameras 1 and associates it with the first level 201 of the digital content 200.
[0085] Then, in step S65, the CPU 111 associates the associated data with the third surface 203 of the digital content 200. The associated data may be the same as the associated data displayed in step S32 of FIG. 3 in the first embodiment, such as the name of the player who shot the ball into the goal.
[0086] In step S66, the CPU 111 associates the first virtual viewpoint image from the first virtual viewpoint described above with the second surface of the digital content 200, and then proceeds to step S81 in FIG.
[0087] If the determination in step S63 is No, the CPU 111 determines in step S67 whether or not the number of selected virtual viewpoints is two, and if it is two, the process proceeds to step S68.
[0088] In step S68, the CPU 111 associates the associated data with the third side 203 of the digital content 200. The associated data may be the same as the associated data associated in step S65, such as the name of the player who shot the ball into the goal.
[0089] Then, in step S69, the CPU 111 associates the first virtual viewpoint image from the first virtual viewpoint with the first plane 201 of the digital content 200. Also, the CPU 111 associates the second virtual viewpoint image from the second virtual viewpoint with the second plane 202 of the digital content 200. Thereafter, the process proceeds to step S81 in FIG. 8.
[0090] If it is determined No in step S67, the process proceeds to step S71 in Fig. 7. In step S71, CPU 111 determines whether or not the operator has selected to associate associated data with third surface 203. If Yes, the process proceeds to step S72, and if No, the process proceeds to step S73.
[0091] In step S72, CPU 111 associates the virtual viewpoint image from the first virtual viewpoint with first plane 201 of digital content 200, associates the virtual viewpoint image from the second virtual viewpoint with second plane 202, and associates the virtual viewpoint image from the third virtual viewpoint with third plane 203. Then, the process proceeds to step S81 in FIG. 8.
[0092] If the answer is No in step S71, in step S73, the CPU 111 associates the associated data with the third side 203 of the digital content 200. The associated data may be the same as the associated data associated in step S65, such as the name of the player who shot the ball into the goal.
[0093] Then, in step S74, CPU 111 associates the first virtual viewpoint image from the first virtual viewpoint with first plane 201 of digital content 200. Furthermore, in step S75, CPU 111 associates the second virtual viewpoint image from the second virtual viewpoint and the third virtual viewpoint image from the third virtual viewpoint with second plane 202 of digital content 200 so that they can be displayed side by side. That is, second plane 202 is divided into two areas for displaying the second virtual viewpoint image and the third virtual viewpoint image, and a virtual viewpoint image is associated with each area. Thereafter, the process proceeds to step S81 in FIG. 8.
[0094] 8, the CPU 111 determines whether or not to assign an NFT to the digital content 200. To do this, for example, a display image (GUI) asking whether or not to assign an NFT to the digital content 200 is displayed on the display unit 115. If the operator selects to assign an NFT, the process proceeds to step S82, where the NFT is assigned to the digital content 200, encryption is performed, and the process proceeds to step S83.
[0095] If the answer in step S81 is No, the process proceeds directly to step S83. Note that the digital content 200 in step S81 may be in the form shown in Figures 4(B) and 4(C), as described above.
[0096] In step S83, the CPU 111 determines whether or not to end the flow of FIGS. 6 to 8, and if the operator has not operated the operation unit 116 to end it, the process proceeds to step S84.
[0097] In step S84, the CPU 111 determines whether the number of virtual viewpoints has changed. If the number has changed, the process returns to step S61. If the number has not changed, the process returns to step S62. If the determination in step S83 is Yes, the flow of FIGS. 6 to 8 ends.
[0098] In the third embodiment, an example has been described in which, when the operator selects the number of virtual viewpoints from 1 to 3, the images to be associated with the first to third planes of the digital content 200 are automatically switched accordingly. However, the operator may select the number of camera images to be associated with the planes constituting the digital content 200 from among images from multiple cameras. Then, a predetermined camera may be automatically selected from multiple cameras 1 accordingly, and the camera images from that camera may be automatically associated with the first to third planes of the digital content 200. The number of viewpoints does not have to be a maximum of three. For example, the number of viewpoints may be determined within a range that maximizes the number of planes constituting the digital content or the number of planes to which images can be associated. Furthermore, if multiple images can be associated with one plane, the maximum number of viewpoints can be further increased.
[0099] Furthermore, between steps S82 and S83, a step may be inserted in which the CPU 111 automatically switches to content consisting of a default stereoscopic image display when a predetermined period (e.g., 30 minutes) has elapsed since the last operation of the operation unit 116. The default stereoscopic image display displays, for example, a main image on the first plane, and a camera image or a virtual viewpoint image from the viewpoint most frequently used according to past statistics on the second plane. The third plane may be, for example, associated data.
[0100] As described above, in the third embodiment, in steps S69, S72, and S95, it is possible to associate a virtual viewpoint image different from the virtual viewpoint image to be displayed on the second surface with the first surface.
[0101] (Embodiment 4) Next, a fourth embodiment will be described with reference to Figures 9 and 10. In this embodiment, the system configuration is the same as that described in the first embodiment, and therefore a description thereof will be omitted. The hardware configuration of the system is also the same as that shown in Figure 2, and a description thereof will also be omitted.
[0102] This embodiment is a display image (GUI) that displays three-dimensional digital content generated by any of the methods of embodiments 1 to 3 on a user device. The user device may be, for example, a PC, a smartphone, or a tablet terminal with a touch panel (not shown). In this embodiment, a tablet terminal with a touch panel will be described as an example. This GUI is generated by the image processing system 100 and transmitted to the user device. Note that this GUI may also be generated by the user device that has acquired the necessary information.
[0103] The image processing system includes a CPU, ROM, RAM, auxiliary storage device, display unit, operation unit, communication I / F, bus, etc. (not shown). The CPU controls the entire image processing system using computer programs, etc. stored in the ROM, RAM, auxiliary storage device, etc.
[0104] The image processing system identifies captured images obtained by capturing images from different directions using multiple imaging devices (cameras), virtual viewpoint video based on a specified virtual viewpoint, sound information associated with the virtual viewpoint video, and information about the subject included in the captured images and virtual viewpoint video from three-dimensional digital content.
[0105] In this embodiment, an example of a three-dimensional digital content generated in the third embodiment will be described, in which the number of virtual viewpoints is three and three virtual viewpoint images are associated with the second surface. The three virtual viewpoint images are images, and will be referred to as virtual viewpoint images hereinafter. In this embodiment, the three-dimensional digital content is a hexahedron, but it may also be a sphere or an octahedron.
[0106] The sound information associated with the virtual viewpoint video is sound information acquired at the venue when the image is captured. Alternatively, sound information corrected based on the virtual viewpoint may be used. The sound information corrected based on the virtual viewpoint is, for example, sound information acquired at the venue when the image is captured, adjusted so that it can be heard when facing the line of sight from the virtual viewpoint at the position of the virtual viewpoint. Separate sound information may also be prepared.
[0107] FIG. 9 is a diagram showing a graphic user interface (GUI) of this embodiment displayed on a user device. The GUI image includes a first area 911 and a second area 912. The first area 911 includes a first display area 901, a second display area 902, a third display area 903, a fourth display area 904, a fifth display area 905, and a sixth display area 906, each of which displays an image showing information associated with three-dimensional digital content. Each display area displays an image or information assigned to it. The second area 912 includes a seventh display area 907, which displays an image showing information associated with a display area selected by the user from the first display area 901 to the sixth display area 906. The images showing information associated with the first display area 901 to the sixth display area 906 may be still images or videos.
[0108] 9 shows an example in which the second display area 902 associated with three virtual viewpoint images is selected. When the display area associated with three virtual viewpoint images is selected, GUIs 908, 909, and 910 corresponding to the virtual viewpoint images of the respective viewpoints are displayed in the second area. Note that the GUIs 908, 909, and 910 may be displayed superimposed on the seventh display area 907.
[0109] In this embodiment, each surface of the three-dimensional digital content is associated with each display area. That is, an image showing information associated with each surface of the digital content is displayed in each display area. Note that an image showing information associated with the digital content may be displayed in each display area regardless of the shape of the digital content.
[0110] The number of display surfaces and display areas for digital content may be different. For example, for hexahedral digital content, only the first display area 901 to the fourth display area 904 may be displayed on the user device. Part of the information associated with the digital content is displayed in the first display area 901 to the third display area 903, and information associated with a display area selected by the user is displayed in the fourth display area 904.
[0111] Information identified from the three-dimensional digital content is associated with each display area, and an image indicating the identified information is displayed. In this embodiment, the subject is a basketball player, and a main image showing the player is displayed in the first display area 901. The image displayed in the first display area 901 may be a virtual viewpoint image or a captured image. The second display area 902 displays superimposed images showing three virtual viewpoint images related to the player displayed in the main image and an icon 913 indicating the virtual viewpoint image. The second display area 902 displays a virtual viewpoint image corresponding to a viewpoint different from that of the image displayed in the first display area 901. The third display area 903 displays an image indicating information about the team to which the player displayed in the main image belongs. The fourth display area 904 displays an image indicating the performance information of the player displayed in the main image for the season in which the main image was captured. The fifth display area 905 displays an image indicating the final score of the game at the time of capture. The sixth display area 906 displays an image indicating copyright information for the digital content.
[0112] If the information associated with the first display area 901 to the sixth display area 906 includes information indicating video, an illustration or icon may be superimposed on the image in the display area corresponding to the video. In this case, different icons are used for video generated from images captured by one imaging device and virtual viewpoint video generated from virtual viewpoint images generated by multiple imaging devices. In this embodiment, an icon 913 is superimposed on the image showing the virtual viewpoint video in the second display area 902. Furthermore, when the second display area 902 is selected by the user, the icon 913 is displayed on the virtual viewpoint video displayed in the seventh display area 907. Note that the illustrations and icons may be arranged around the display areas.
[0113] Among the first display area 901 to the sixth display area 906, a plurality of images or a plurality of videos may be associated with one display area. In this embodiment, a case will be described in which three virtual viewpoint videos with different viewpoints are associated with the second display area 902. In this case, virtual viewpoint videos with different viewpoints are associated with GUI 908, GUI 909, and GUI 910. A virtual viewpoint video of the subject viewpoint (viewpoint 1) is associated with GUI 908. A virtual viewpoint video corresponding to a viewpoint from behind the subject (viewpoint 2) is associated with GUI 909. A virtual viewpoint video of a virtual viewpoint (viewpoint 3) within a position on a sphere centered on the subject is associated with GUI 910. Note that if a plurality of images or a plurality of videos cannot be associated with one display area, GUI 908, GUI 909, and GUI 910 do not need to be displayed.
[0114] An initial image is set in the seventh display area 907 as information to be displayed before the user selects one of the first to sixth display areas 901 to 906. The initial image may be an image associated with the first to sixth display areas 901 to 906, or may be an image different from the images associated with the first to sixth display areas 901 to 906. In this embodiment, the main image in the first display area 901 is set as the initial image.
[0115] Fig. 10 is a flowchart for explaining the operation flow of the image processing system in this embodiment. Specifically, this is processing executed by CPU 111 in Fig. 2. In step S1001, content information associated with the first to sixth surfaces of the three-dimensional digital content is identified.
[0116] In step S1002, the content information identified in step S1001 is associated with the first display area 901 to the sixth display area 906. After that, images showing the associated information are displayed in the first display area 901 to the sixth display area 906. Furthermore, a preset initial image is displayed in the seventh display area 907. In this embodiment, the main image in the first display area 901 is displayed in the seventh display area 907 as the initial image.
[0117] In step S1003, it is determined whether a predetermined time (for example, 30 minutes) has passed since the latest input was received. If Yes, the process proceeds to step S1017. If No, the process proceeds to step S1004.
[0118] In step S1004, it is determined whether or not an input from the user to select any of the first to sixth display areas 901 to 906 has been received. If Yes, the step to proceed varies depending on the received input. If an input to select the first display area 901 has been received, proceed to step S1005. If an input to select the second display area 902 has been received, proceed to step S1006. If an input to select the third display area 903 has been received, proceed to step S1007. If an input to select the fourth display area 904 has been received, proceed to step S1008. If an input to select the fifth display area 905 has been received, proceed to step S1009. If an input to select the sixth display area 906 has been received, proceed to step S1010. If No, return to step S1003.
[0119] In step 1005, a main image showing the player corresponding to first display area 901 is displayed in seventh display area 907. If a main image showing the player associated with first display area 901 is already displayed in seventh display area 907, the process returns to step S1003 with the main image still displayed in seventh display area 907. If information related to the display area selected by the user is already displayed in seventh display area 907, steps 1007 to 1010 are similar and therefore will not be described further below. Furthermore, if the main image is a video and the main image is already displayed in seventh display area 907, the video may be played again from a predetermined playback time, or the currently displayed video may be played as is.
[0120] In step S1006, a virtual viewpoint video related to the player corresponding to the second display area 902 is displayed in the seventh display area 907. If multiple virtual viewpoint videos are associated with the second display area 902, a preset virtual viewpoint video is displayed in the seventh display area 907. In this embodiment, a virtual viewpoint video of the subject viewpoint (viewpoint 1) is displayed in the seventh display area 907. After the preset virtual viewpoint video is displayed in the seventh display area 907, the process proceeds to step S1011.
[0121] In step S1011, it is determined whether a predetermined time (for example, 30 minutes) has passed since the latest input was received. If Yes, the process proceeds to step S1017. If No, the process proceeds to step S1012.
[0122] In step S1012, it is determined whether an input has been received to select a virtual viewpoint video of a virtual viewpoint that the user wishes to view from among the multiple virtual viewpoint videos. Specifically, a GUI indicating each virtual viewpoint is displayed as in the seventh display area 907 of FIG. 9, and the virtual viewpoint video of each viewpoint is selected when the user selects a GUI. If Yes, the process proceeds to the next step depending on the selected GUI. If GUI 908 is selected, the process proceeds to step S1013; if GUI 909 is selected, the process proceeds to step S1014; and if GUI 910 is selected, the process proceeds to step S1015. If No, the process proceeds to step S1016.
[0123] Note that a GUI indicating each virtual viewpoint may not be provided, and a virtual viewpoint video of each viewpoint may be selected by a flick operation or a touch operation on the seventh display area 907. Furthermore, when multiple virtual viewpoint videos are associated with the second display area 902, the multiple virtual viewpoint videos may be connected to form a single virtual viewpoint video and played back continuously. In this case, the process proceeds to step S1016 without going through step S1012 for selecting a virtual viewpoint.
[0124] In step S1013, the virtual viewpoint video from the subject's viewpoint (viewpoint 1) is displayed in the seventh display area 907. Then, the process returns to step S1011. If the virtual viewpoint video from the subject's viewpoint is already displayed in the seventh display area 907, the video may be played again from the default playback time, or the currently displayed video may be played as is. If the video is already displayed in the seventh display area 907, the same applies to steps S1014 and S1015, and therefore further description is omitted below.
[0125] In step S1014, a virtual viewpoint video corresponding to a viewpoint (viewpoint 2) from behind the subject is displayed in the seventh display area 907. After that, the process returns to step S1011.
[0126] In step S1015, a virtual viewpoint video of a virtual viewpoint (viewpoint 3) within a position on a sphere centered on the subject is displayed in the seventh display area 907. After that, the process returns to step S1011.
[0127] In step S1016, it is determined whether an input from the user to select one of the first to sixth display areas 901 to 906 has been received. If Yes, the same process as in step S1004 is performed. If No, the process returns to step S1011.
[0128] In step S1007, information about the team to which the player corresponding to the third display area 903 belongs is displayed in the seventh display area 907.
[0129] In step S1008, the performance information for this season of the player corresponding to the fourth display area 904 is displayed in the seventh display area 907.
[0130] In step S1009, the final score of the match corresponding to the fifth display area 905 is displayed in the seventh display area 907.
[0131] In step S1010, the copyright information corresponding to the sixth display area 906 is displayed in the seventh display area 907.
[0132] In step S1017, an initial image is displayed in the seventh display area 907. In this embodiment, the main image in the first display area 901 is displayed as the initial image in the seventh display area 907. Thereafter, the processing flow ends.
[0133] (Embodiment 5) Next, a fifth embodiment will be described with reference to Figures 11 and 12. In this embodiment, the system configuration is the same as that described in the first embodiment, and therefore a description thereof will be omitted. The hardware configuration of the system is also the same as that shown in Figure 2, and a description thereof will also be omitted.
[0134] FIG. 11 is a diagram showing a graphic user interface (GUI) of this embodiment. As in the fourth embodiment, the following description will be given taking as an example a three-dimensional digital content generated in the third embodiment, in which the number of virtual viewpoints is three and three virtual viewpoint images are associated with the second surface. Also, unlike the fourth embodiment, the number of display areas displayed in the first area 1107 is different from the number of surfaces of the digital content. Specifically, the number of display areas displayed in the first area 1107 is five, while the number of surfaces of the digital content is six. This GUI is generated by the image processing system 100 and transmitted to the user device. Note that this GUI may also be generated by the user device that has acquired the necessary information.
[0135] In this embodiment, the digital content includes three virtual viewpoint images from different viewpoints, and the virtual viewpoint images from each viewpoint are respectively associated with second display area 1102 to fourth display area 1104. As in embodiment 4, the three virtual viewpoints are a virtual viewpoint image from the subject viewpoint (viewpoint 1), a viewpoint from behind the subject (viewpoint 2), and a virtual viewpoint within a position on a sphere centered on the subject (viewpoint 3). In order to associate the three virtual viewpoint images with second display area 1102 to fourth display area 1104, icons 1109 indicating the virtual viewpoint images are also superimposed and displayed in second display area 1102 to fourth display area 1104.
[0136] An image of a player is associated with the first display area 1101, and copyright information is associated with the fifth display area 1105. The first display area 1101 may contain a virtual viewpoint video or a captured image from a viewpoint different from the virtual viewpoint images in the second display area 1102 to the fourth display area 1104. Note that the information displayed in the display areas is not limited to these, and may be any information associated with digital content.
[0137] Fig. 12 is a flowchart for explaining the operation flow of image processing system 100 of this embodiment. Note that the operation of each step in the flowchart in Fig. 12 is performed by CPU 111, which serves as a computer in image processing system 100, executing a computer program stored in a memory such as ROM 112 or auxiliary storage device 114. Note that steps in Fig. 12 with the same reference numerals as in Fig. 10 are the same processes, and their explanations will be omitted.
[0138] 12, if an input to select the second display area 1102 is received, the process proceeds to step S1201. If an input to select the third display area 1103 is received, the process proceeds to step S1202. If an input to select the fourth display area 1104 is received, the process proceeds to step S1203.
[0139] In step S1201, a virtual viewpoint video of the subject viewpoint (viewpoint 1) is displayed in the sixth display area 1106. After that, the process returns to step S1003.
[0140] In step S1202, a virtual viewpoint video corresponding to a viewpoint (viewpoint 2) from behind the subject is displayed in the sixth display area 1106. After that, the process returns to step S1003.
[0141] In step S1203, a virtual viewpoint video of a virtual viewpoint (viewpoint 3) within a position on a sphere centered on the subject is displayed in the sixth display area 1106. After that, the process returns to step S1003.
[0142] (Embodiment 6) Next, a sixth embodiment will be described with reference to Figures 13 and 14. In this embodiment, the system configuration is the same as that described in the first embodiment, and therefore a description thereof will be omitted. The hardware configuration of the system is also the same as that shown in Figure 2, and a description thereof will also be omitted.
[0143] FIG. 13 is a diagram showing a graphic user interface (GUI) according to the sixth embodiment. In this embodiment, an example of three-dimensional digital content generated in the third embodiment, in which the number of virtual viewpoints is six and six virtual viewpoint images are associated with the second surface, will be described. The six virtual viewpoint images are images, and will be referred to as virtual viewpoint images hereinafter. This GUI is generated by the image processing system 100 and transmitted to the user device. Note that this GUI may also be generated by the user device that has acquired the necessary information.
[0144] This embodiment differs from the fourth embodiment in that the number of display areas displayed in the first area 1307 is different from the number of screens of the digital content. Specifically, the number of display areas displayed in the first area 1307 is three, while the number of screens of the digital content is six. In addition, six virtual viewpoint videos are associated with the second display area 1302.
[0145] A main image showing a player is associated with the first display area 1301, and copyright information is associated with the third display area 1303. The first display area 1301 may contain a virtual viewpoint video or a captured image from a viewpoint different from the virtual viewpoint image in the second display area 1302. Note that the information displayed in the display areas is not limited to these, and may be any information associated with digital content.
[0146] This embodiment differs from the fourth and fifth embodiments in that the second area 1308 includes a fourth display area 1304, a fifth display area 1305, and a sixth display area 1306. When the user selects one of the first to third display areas 1301 to 1303, an image corresponding to the selected display area is displayed in the fifth display area.
[0147] The fifth display region 1305 is always displayed in the second region 1308. On the other hand, the fourth display region 1304 and the sixth display region 1306 are displayed in the second region 1308 when a virtual viewpoint video is displayed in the fifth display region.
[0148] The display area located at the center of the second area 1308 is different in shape from the other display areas. Specifically, the fourth display area 1304 and the sixth display area 1306 are different in shape or size from the fifth display area 1305. In this embodiment, the fifth display area is rectangular, while the fourth display area and the sixth display area are trapezoidal. This can improve the visibility of the fifth display area 1305 located at the center of the second area 1308.
[0149] Each of the six virtual viewpoints has a different virtual viewpoint. Three virtual viewpoints have three subjects, and the positions of the three subjects are used as the virtual viewpoints. The other three virtual viewpoints have a virtual viewpoint located a certain distance behind and above the positions of the three subjects. For example, in a basketball game, with player A on offense, player B on defense, and a basketball as the subjects, the following examples can be given. One is a first virtual viewpoint image in which the position of player A's face is used as the virtual viewpoint and the direction of player A's face is used as the line of sight from the virtual viewpoint. Another is a second virtual viewpoint image in which the virtual viewpoint is used as the virtual viewpoint located a certain distance behind (e.g., 3 m behind) and above (e.g., 1 m above) the position of player A's face and the line of sight from the virtual viewpoint is set to a direction that includes player A within the angle of view. Finally, a third virtual viewpoint image is used as the virtual viewpoint located at the position of player B's face and the direction of player B's face is used as the line of sight from the virtual viewpoint. Another one is a fourth virtual viewpoint image in which the virtual viewpoint is a position a certain distance behind and a certain distance above the position of player B's face, and the line of sight from the virtual viewpoint is set to a direction that includes player B within the angle of view. Another one is a fifth virtual viewpoint image in which the virtual viewpoint is the position of the center of gravity of the basketball, and the line of sight from the virtual viewpoint is the direction of movement of the basketball. Still another is a sixth virtual viewpoint image in which the virtual viewpoint is a position a certain distance behind and a certain distance above the position of the center of gravity of the basketball, and the direction of movement of the basketball. Note that the position a certain distance behind and a certain distance above the position of the subject may be determined based on the shooting scene or the proportion of the subject's occupancy in the angle of view of the virtual viewpoint image. The line of sight from the virtual viewpoint is set based on any one of the subject's posture, the subject's direction of movement, or the subject's position within the angle of view.
[0150] In this embodiment, the six virtual viewpoint videos have the same playback time, but the virtual viewpoint videos may have different playback times.
[0151] When second display area 1302 corresponding to six virtual viewpoint images is selected, fourth display area 1304 and sixth display area 1306 are displayed in second area 1308. Virtual viewpoint images showing three subjects are respectively associated with the three display areas displayed in second area 1308. Specifically, fourth display area 1304 is associated with the first and second virtual viewpoint images with player A as the subject, fifth display area 1305 is associated with the third and fourth virtual viewpoint images with player B as the subject, and sixth display area 1306 is associated with the fifth and sixth virtual viewpoint images with a basketball as the subject.
[0152] The images displayed in each display area are not all played simultaneously. Only the image displayed in the display area located at the center of the second area 1308 is played. In this embodiment, only the virtual viewpoint image displayed in the fifth display area 1305 is played.
[0153] Fig. 14 is a flowchart for explaining the operation flow of image processing system 100 of embodiment 6. Note that the operation of each step in the flowchart in Fig. 5 is performed by CPU 111 as a computer of image processing system 100 executing a computer program stored in a memory such as ROM 112 or auxiliary storage device 114. Note that in Fig. 14, steps with the same reference numerals as in Fig. 10 are the same processes, and explanations thereof will be omitted.
[0154] In step S1401, preset virtual viewpoint images are displayed in the fourth display region 1304, the fifth display region 1305, and the sixth display region 1306. In this embodiment, three virtual viewpoint images are displayed, with the positions of three subjects as the virtual viewpoints. Specifically, the fourth display region 1304 displays a first virtual viewpoint image with the position of player A's face as the virtual viewpoint, the fifth display region 1305 displays a third virtual viewpoint image with the position of player B's face as the virtual viewpoint, and the sixth display region 1306 displays a fifth virtual viewpoint image with the position of the center of gravity of the basketball as the virtual viewpoint. At this time, only the fifth display region 1305 located at the center of the second region 1308 plays an image, and the fourth display region 1304 and the sixth display region 1306 do not play an image. After the preset virtual viewpoint images are displayed in the fourth display region 1304, the fifth display region 1305, and the sixth display region 1306, the process proceeds to step S1011.
[0155] In step S1402, it is determined whether an operation to change the subject of the virtual viewpoint video displayed in fifth display area 1305 has been input. Specifically, it is determined whether an operation to switch to a virtual viewpoint video with a different subject has been input by a horizontal sliding operation on fifth display area 1305. If the answer is Yes, the process proceeds to the next step based on the sliding direction. If input information of a leftward sliding operation has been received, the process proceeds to step S1410. If input information of a rightward sliding operation has been received, the process proceeds to step S1411. If the answer is No, the process proceeds to step S1405.
[0156] In step S1403, the virtual viewpoint video associated with each display area is reassociated with the display area to the left. For example, when a leftward slide operation is received on the third virtual viewpoint video associated with the fifth display area 1305, the third and fourth virtual viewpoint videos associated with the fifth display area 1305 are associated with the fourth display area 1304, which is to the left of the fifth display area 1305. The fifth and sixth virtual viewpoint videos associated with the sixth display area 1306 are associated with the fifth display area 1305, which is to the left of the sixth display area 1306. In the second area 1308, since there is no display area to the left of the fourth display area 1304, the first and second virtual viewpoint videos associated with the fourth display area 1304 are associated with the sixth display area 1306, which has no display area to the right. After the re-association, among the virtual viewpoint videos associated with the fifth display area 1305, a virtual viewpoint video is played back in which the position of the virtual viewpoint relative to the subject is the same as the virtual viewpoint video played back in the fifth display area 1305 before the association. For example, if the virtual viewpoint video played back in the fifth display area 1305 before the association was the third virtual viewpoint video, after the association, the fifth virtual viewpoint video and the sixth virtual viewpoint video are associated with the fifth display area 1305. Since the third virtual viewpoint video is a virtual viewpoint video in which the position of the subject is the virtual viewpoint position, the fifth virtual viewpoint video, which is also a virtual viewpoint video in which the position of the subject is the virtual viewpoint position, is displayed in the fifth display area. This allows the user to intuitively switch between virtual viewpoint videos with different subjects. After the above processing is performed, the process returns to step S1011.
[0157] In step S1404, the virtual viewpoint video associated with each display area is reassociated with the display area to the right. For example, when a rightward slide operation is received on the third virtual viewpoint video associated with the fifth display area 1305, the third and fourth virtual viewpoint videos associated with the fifth display area 1305 are associated with the sixth display area 1306 located to the right of the fifth display area 1305. The first and second virtual viewpoint videos associated with the fourth display area 1304 are associated with the fifth display area 1305 located to the right of the fourth display area 1304. In the second area 1308, since there is no display area to the right of the sixth display area 1306, the fifth and sixth virtual viewpoint videos associated with the sixth display area 1306 are associated with the fourth display area 1304, which has no display area to the left. After the re-association, the virtual viewpoint video associated with the fifth display area 1305 is played back, the virtual viewpoint video having the same position of the virtual viewpoint relative to the subject as the virtual viewpoint video played back in the fifth display area 1305 before the association. After the above processing is performed, the process returns to step S1408. After the above processing is performed, the process returns to step S1011.
[0158] In step S1405, it is determined whether an operation to change the virtual viewpoint position of the virtual viewpoint video displayed in fifth display area 1305 has been input. Specifically, an input to switch to a virtual viewpoint video of the same subject but with a different virtual viewpoint position is accepted by a double tap operation on fifth display area 1305. If the answer is Yes, proceed to step S1406; if the answer is No, proceed to step S1016.
[0159] In step S1406, a process for changing the position of the subject in the virtual viewpoint video is performed. Specifically, the process switches to a virtual viewpoint video that shows the same subject but with a different virtual viewpoint. For example, assume that a third virtual viewpoint video, in which the position of player B's face is the virtual viewpoint, and a fourth virtual viewpoint video, in which the virtual viewpoint is a certain distance behind and higher than the position of player B's face, are associated in fifth display area 1305. If a double-tap operation is received while the third virtual viewpoint video is displayed in fifth display area 1305, a process for switching to the fourth virtual viewpoint video and displaying it in fifth display area 1305 is performed. This allows the position of the virtual viewpoint for the same subject to be intuitively switched. After performing the above process, the process returns to step S1011.
[0160] Furthermore, when switching the virtual viewpoint video using a slide operation or a double tap operation, the time code of the virtual viewpoint video being played in the fifth display area 1305 may be recorded, and the virtual viewpoint video after switching may be played from the time of the recorded time code.
[0161] In this embodiment, a double tap operation is used to switch the position of the virtual viewpoint on the same subject, but other operations may be used, such as a pinch-in / pinch-out operation or a sliding operation in the up / down direction on the fifth display area 1305.
[0162] In the present embodiment, multiple virtual viewpoint images showing the same subject are associated with one display area. However, multiple virtual viewpoint images showing the same subject may be associated with multiple display areas. Specifically, a seventh display area, an eighth display area, and a ninth display area may be newly provided above the fourth display area 1304, the fifth display area 1305, and the sixth display area 1306, respectively (not shown). In this case, the first virtual viewpoint image is displayed in the fourth display area 1304, the third virtual viewpoint image in the fifth display area 1305, the fifth virtual viewpoint image in the sixth display area 1306, the second virtual viewpoint image in the seventh display area, the fourth virtual viewpoint image in the eighth display area, and the sixth virtual viewpoint image in the ninth display area. In the sixth embodiment, the operation of switching between virtual viewpoint images showing the same subject was performed by a tap gesture. In this case, the display area is switched by a sliding operation in the up or down direction. This allows the user to intuitively control the virtual viewpoint.
[0163] (Embodiment 7) In the first embodiment, a second image having a predetermined relationship with a main image (first image) associated with a first surface 201 of digital content 200 is associated with a second surface 202 of digital content 200. In the present embodiment, a plurality of virtual viewpoint images with the same time code are associated with each surface of a three-dimensional digital content. Specifically, an example will be described in which a virtual viewpoint image is associated with each surface of a three-dimensional digital content based on the line of sight of the virtual viewpoint with respect to the subject.
[0164] 15 is a diagram showing an image processing system 101 of this embodiment. Note that the same blocks as those in FIG. 1 are given the same numbers and their explanations will be omitted.
[0165] The image generation unit 117 analyzes the correspondence between the position of the virtual viewpoint, the line of sight from the virtual viewpoint, and the coordinates of the object based on the virtual viewpoint specified by the operation unit 116 and the coordinate information of the object displayed in the virtual viewpoint video. The image generation unit 117 identifies an object of interest from the virtual viewpoint video viewed from the virtual viewpoint specified by the operation unit 116. In this embodiment, the object that appears at the center of the virtual viewpoint video or is closest to the center is identified, but this is not limiting. For example, the object that occupies the largest proportion of the virtual viewpoint video may be identified, or the object may be selected without generating a virtual viewpoint video. Next, the image generation unit 117 determines the shooting direction for shooting the identified object of interest and generates multiple virtual viewpoints corresponding to each shooting direction. The shooting directions are up, down, left, right, front, back (front and rear). The multiple virtual viewpoints generated correspond to the same time code as the virtual viewpoint specified by the operator. After generating virtual viewpoint videos corresponding to the generated virtual viewpoints, the image generation unit 117 determines from which shooting direction (up, down, left, right, front, back) the object of interest is captured in each virtual viewpoint video, and assigns shooting direction information to the virtual viewpoint video. The shooting direction information indicates the direction from which the image was shot relative to the orientation of the target object. The shooting direction is determined based on the positional relationship between the object captured at the center of the virtual viewpoint video and a predetermined point at the start of the virtual viewpoint video. Details will be explained in FIG. 17.
[0166] The content generation unit 118 determines which surface of the three-dimensional content the virtual viewpoint video received from the image generation unit 117 is to be associated with based on the attached shooting direction information, and generates three-dimensional digital content.
[0167] 16 is a diagram illustrating the faces of three-dimensional digital content in this embodiment. In three-dimensional digital content 1600, face 1601 is defined as the front face, face 1602 as the right face, face 1603 as the top face, face 1604 as the left face, face 1605 as the back face, and face 1606 as the bottom face.
[0168] FIG. 17 is a diagram illustrating the shooting direction relative to the players. A court 1700 is provided with a goal 1701, a goal 1702, and a player 1703. The player 1703 is assumed to attack toward the goal 1701. Here, the direction connecting the player 1703 and the goal 1701, and the direction parallel to the ground from the goal 1701 toward the player 1703, is defined as the "front." Here, a method for determining the front is explained. First, a line segment connecting the player 1703 and the goal 1701 is derived. The goal is a predetermined point, and the player is the center of gravity of the 3D model. Next, a surface that is perpendicular to the derived line segment and in contact with the player's 3D model is derived, and this surface is defined as the front. After determining the front, a bounding box surrounding the player is determined. The top, bottom, right, left, and back faces are determined based on the front of the bounding box. In FIG. 17 , the direction in which a player is viewed from the front is represented by arrow 1704, and the shooting direction information of a virtual viewpoint video created using this direction as the viewing direction from the virtual viewpoint is “front.” A virtual viewpoint video with “front” shooting direction information assigned is associated with plane 1601 in FIG. 16 . Similarly, the shooting direction information of a virtual viewpoint video created using the direction of arrow 1706 capturing a player from the right as the viewing direction from the virtual viewpoint is “right side.” A virtual viewpoint video with “right side” shooting direction information assigned is associated with plane 1602 in FIG. 16 . The shooting direction information of a virtual viewpoint video created using the direction of arrow 1707 capturing a player from the left as the viewing direction from the virtual viewpoint is “left side.” A virtual viewpoint video with “left side” shooting direction information assigned is associated with plane 1604 in FIG. 16 . The shooting direction information of a virtual viewpoint video created using the direction of arrow 1708 capturing a player from above as the viewing direction from the virtual viewpoint is “top side.” The virtual viewpoint video to which the shooting direction information of "top" is assigned is associated with plane 1603 in FIG. 16. The shooting direction information of the virtual viewpoint video created with the direction of arrow 1709 capturing the player from below as the line of sight from the virtual viewpoint is "bottom." The virtual viewpoint video to which the shooting direction information of "bottom" is assigned is associated with plane 1606 in FIG. 16. The shooting direction information of the virtual viewpoint video created with the direction of arrow 1705 capturing the player from behind as the line of sight from the virtual viewpoint is "rear."The virtual viewpoint video to which the shooting direction information of "rear" is assigned is associated with plane 1605 in Fig. 16. Note that in this embodiment, the direction is determined based on the relationship between the player and the goal at a specific moment, but the direction may change based on the relationship between the player and the goal as the player moves on the field.
[0169] In this embodiment, the shooting direction is determined based on the positions of the player and the goal, but this is not limiting. For example, the direction in which a player is viewed from the player's direction of travel is set as the "front," and directions obtained by rotating the "front" direction by ±90 degrees on the XY plane are set as the "left side" and the "right side." Directions obtained by rotating the "front" direction by ±90 degrees on the YZ plane are set as the "top side" and the "bottom side." Directions obtained by rotating the "front" direction by +180 degrees or −180 degrees on the XY plane are set as the "rear side." As another example of setting the front, the direction as viewed from the direction of the player's face may be set as the "front," or the direction as viewed from the position of the goal closest to the straight line in the player's direction of travel may be set as the "front."
[0170] Note that the shooting direction may be changed depending on the player's position relative to the goal each time the player moves. For example, the front may be redefined if the player moves a certain distance (e.g., 3 meters) or more from a position where the front was previously determined, or the shooting direction may be changed after a certain amount of time has passed. Alternatively, the front may be redefined if, when comparing the initially determined front with the front calculated after the player has moved, the angle when viewed from above changes by 45 degrees or more. Also, the front may be redefined each time the ball is passed to another player, triggering the redefinition of the front based on the relationship between the player receiving the pass and the goal position.
[0171] 18 is a diagram showing an example of three-dimensional digital content generated by the content generation unit 4. A plurality of virtual viewpoint images viewed from a plurality of virtual viewpoints corresponding to the same time code as the virtual viewpoint set by the operator correspond to each surface of the three-dimensional digital content. By displaying in this manner, the user can intuitively grasp from which position the object was captured in the virtual viewpoint image.
[0172] Fig. 19 is a flowchart for explaining the operational flow of the image processing system 101 of the seventh embodiment. Note that the same steps as S37 to S39 in Fig. 3 are given the same numbers and the explanation will be omitted.
[0173] In S1901, the image generation unit 117 acquires virtual viewpoint information indicating the position of the virtual viewpoint designated by the user via the operation unit 116 and the line of sight direction from the virtual viewpoint.
[0174] In S1902, the image generation unit 117 identifies an object of interest in a virtual viewpoint image viewed from a virtual viewpoint corresponding to the acquired virtual viewpoint information. In this embodiment, the image generation unit 117 identifies an object that appears at the center of the virtual viewpoint image or that is closest to the center.
[0175] In S1903, the image generation unit 117 determines the shooting direction for the target object. In this embodiment, the front is defined as the surface that is perpendicular to the line connecting the target object's position and a predetermined position and that is in contact with the 3D model of the target object. Using this front as a reference, the shooting directions for up, down, left, right, and behind are determined.
[0176] In S1904, the image generation unit 117 generates a plurality of virtual viewpoints corresponding to the plurality of shooting directions determined in S1903. In this embodiment, shooting directions corresponding to front, back, up, down, left, and right with respect to the object of interest are determined, and therefore virtual viewpoints corresponding to each of these directions are generated. Note that the generated virtual viewpoint only needs to have the line of sight from the virtual viewpoint set in the same direction as the shooting direction, and the object of interest does not need to be positioned on the optical axis of the virtual viewpoint. Furthermore, the position of the generated virtual viewpoint is set at a position a predetermined distance away from the position of the object of interest. In this embodiment, the virtual viewpoint is set at a position 3 m away from the object of interest.
[0177] In S1905, the image generation unit 117 generates a virtual viewpoint image corresponding to the generated virtual viewpoint, and then adds imaging direction information indicating the imaging direction corresponding to the virtual viewpoint to the generated virtual viewpoint image.
[0178] In S1906, the image generation unit 117 determines whether or not virtual viewpoint videos have been generated for all of the virtual viewpoints generated in S1904. If all of the virtual viewpoint videos have been generated, the image generation unit 117 transmits the generated virtual viewpoint videos to the content generation unit 118 and proceeds to S1907. If all of the virtual viewpoint videos have not been generated, the process proceeds to S1905, and loops until all of the virtual viewpoint videos have been generated.
[0179] In S1907, the content generation unit 118 associates the virtual viewpoint video with each surface of the three-dimensional digital content based on the received shooting direction information of the virtual viewpoint video.
[0180] In S1908, the content generation unit 118 determines whether the received virtual viewpoint video is associated with all surfaces of the three-dimensional digital content. If so, the process proceeds to S37; if not, the process proceeds to S1907. Note that in this embodiment, it is assumed that the virtual viewpoint video is associated with all surfaces, but this is not limiting, and the virtual viewpoint video may be associated with a specific surface. In this case, in S1908, it is determined whether the virtual viewpoint video is associated with the specific surface.
[0181] The above process allows a virtual viewpoint video corresponding to the shooting direction to be associated with each surface of the three-dimensional digital content, so that a user viewing the virtual viewpoint video using the digital content can intuitively grasp the virtual viewpoint video corresponding to each surface when they want to switch the virtual viewpoint video.
[0182] (Embodiment 8) In the seventh embodiment, a plurality of virtual viewpoints corresponding to the same time code are generated based on a virtual viewpoint specified by an operator, and virtual viewpoint images corresponding to the shooting direction of each virtual viewpoint are associated with each surface of the digital content. However, there may be cases where the virtual viewpoint image seen from the virtual viewpoint specified by the operator is desired to be associated with a surface corresponding to the shooting direction. In the eighth embodiment, the shooting direction from which the object of interest is shot in the virtual viewpoint image seen from the virtual viewpoint specified by the operator is identified, and the virtual viewpoint image is associated with each surface of the digital content according to the shooting direction.
[0183] Fig. 20 is a flowchart for explaining the operational flow of the image processing system 101 of embodiment 8. Note that the same steps as S37 to S39 in Fig. 3, and S1901, S1902, S1907, and S1908 in Fig. 19 are designated by the same numbers, and their explanations will be omitted.
[0184] In step S2001, the image generation unit 117 generates a virtual viewpoint video based on the virtual viewpoint information acquired in S1901.
[0185] In step S2002, the image generation unit 117 determines the shooting direction of the target object for each frame of the virtual viewpoint video. For example, if the virtual viewpoint video has a total of 1000 frames, then 800 frames are assigned with shooting direction information of "front." 100 frames are assigned with shooting direction information of "rear." 50 frames are assigned with shooting direction information of "left." 30 frames are assigned with shooting direction information of "right." 20 frames are assigned with shooting direction information of "top." 10 frames are assigned with shooting direction information of "bottom." In this way, shooting direction information is assigned to each frame with a different time code.
[0186] In this embodiment, when a user watches digital content, if a frame captured in a different direction is displayed, the screen rotates to show the corresponding digital content. This allows the three-dimensional shape to be utilized, giving the virtual viewpoint image a sense of dynamism.
[0187] Furthermore, in this embodiment, the planes associated with the three-dimensional digital content are determined in advance for each shooting direction, but this is not limiting. For example, the planes associated with the three-dimensional digital content may be determined in descending order of the frame ratio of each virtual viewpoint video capturing an object. Specifically, the second to sixth directions are set in descending order of the frame ratio. Assume that, of 1,000 frames of the virtual viewpoint video, 800 frames are from the front 1604, 100 frames are from the rear 1605, 50 frames are from the left 1606, 30 frames are from the right 1607, 20 frames are from the top 1608, and 10 frames are from the bottom 1609. In this case, the first direction is the front, and the subsequent second to sixth directions are determined in the order of rear, left, right, top, and bottom. The virtual viewpoint video with the shooting direction information added is then output to the content generation unit 118.
[0188] (Embodiment 9) In the fourth embodiment, an example was described in which the storage unit 5 for storing digital content is incorporated into the image processing device 100. In the ninth embodiment, an example will be described in which digital content is stored in an external device 2102. Note that the image processing device 102 is a device in which the storage unit 5 is removed from the image processing device 100 (not shown).
[0189] 21 is a diagram showing the system configuration of an image processing system 103 according to the ninth embodiment. The image processing system 103 includes an image processing device 102, a user device 2101, and an external device 2102.
[0190] The image processing device 102 generates digital content using the method described in any one of the first to third embodiments. The generated digital content, media data such as virtual viewpoint images used in the generation, icons indicating each virtual viewpoint image, metadata of each virtual viewpoint image, etc. are transmitted to an external device 2102. The image processing device 102 also generates a display image described in any one of the fourth to sixth embodiments. The generated display image is transmitted to a user device 2101.
[0191] The user device 2101 may be, for example, a PC, a smartphone, or a tablet terminal with a touch panel (not shown). In this embodiment, a tablet terminal with a touch panel will be described as an example.
[0192] The external device 2102 stores digital content generated by the method described in any one of the first to third embodiments. Like the storage unit 5 in FIG. 1, the external device 2102 stores not only digital content but also virtual viewpoint images, camera images, icons corresponding to the virtual viewpoint images displayed in each digital content, and the like. When a specific digital content is requested by the image processing device 102, the external device 2102 transmits the corresponding digital content to the image processing device 102. In addition to the digital content, the external device 2102 may transmit the virtual viewpoint images, metadata of each virtual viewpoint image, and the like to the image processing device 102.
[0193] 22 is a diagram showing the flow of data transmission in embodiment 9. Note that this embodiment will explain the flow of generating a display image in response to an instruction from a user, and further changing the virtual viewpoint image to be displayed in response to an operation on the display image.
[0194] In S2201, the user device 2101 receives input from the user and transmits an instruction to view digital content to the image processing device 102. This instruction includes information for identifying the digital content to be viewed. Specifically, this information includes the NFT of the digital content, an address stored in the external device 2102, and the like.
[0195] In S2202, the image processing device 102 requests the external device 2102 to provide the digital content to be viewed, based on the acquired viewing instruction.
[0196] In S2203, the external device 2102 identifies digital content corresponding to the acquired request. Note that in addition to the digital content, metadata of the digital content and related virtual viewpoint images are also identified according to the request content.
[0197] In S2204, the external device 2102 transmits the identified digital content to the image processing device 102. If there is identified information in addition to the digital content, it is also transmitted.
[0198] In S2205, the image processing device 102 generates a display image corresponding to the acquired digital content. For example, if three virtual viewpoint images are associated with the acquired digital content, the image processing device 102 creates the display image shown in FIG. 9 of the fourth embodiment. Alternatively, the image processing device 102 may create the display image shown in FIG. 11 of the fifth embodiment, or the display image shown in FIG. 13 of the sixth embodiment. In this embodiment, the display image to be generated is assumed to be set in advance as the display image shown in FIG. 9 of the fourth embodiment. A display image corresponding to the digital content may be set in advance. In this case, the creator sets the display image when creating the digital content and stores it in the metadata of the digital content. As another example, the user may specify the display image to be generated. In this case, the type of display image to be displayed is also specified when issuing an instruction to view the digital content in S2201. Each aspect of the acquired digital content is then associated with the generated display image. In this embodiment, the subject is a basketball player, and an image corresponding to a main image showing the player is displayed in the first display area 901. In the second display area 902, images showing three virtual viewpoint images related to the player displayed in the main image and an icon 913 showing the virtual viewpoint image are superimposed and displayed. An image showing information about the team to which the player displayed in the main image belongs is displayed in the third display area 903. An image showing the performance information of the player displayed in the main image for the season in which the main image was captured is displayed in the fourth display area 904. An image showing the final score of the game at the time of capture is displayed in the fifth display area 905. An image showing copyright information for the digital content is displayed in the sixth display area 906. Note that the first display area 901 to the sixth display area 906 are images that can be selected by user operation. In S2205, the seventh display area 907 displays the main image showing the player displayed in the first display area 901 as the initial image.
[0199] In S2206, the image processing apparatus 102 transmits the display image generated in S2205 to the user device 2101.
[0200] In S2207, the user device 2101 displays the received display image.
[0201] In S2208, when the user device 2101 receives an operation to select a display area of a display image by a user operation, the user device 2101 transmits information specifying the selected display area to the image processing device 102. For example, when the display image shown in FIG. 9 of the fourth embodiment is displayed and the user selects an image corresponding to the display area 902, the user device 2101 transmits information indicating that the display area 902 has been selected to the image processing device 102. Note that in the present embodiment, the information specifying the display area is transmitted to the image processing device 102. This is to enable an image or video different from the images displayed in the first display area 901 to the sixth display area 906 to be displayed in the seventh display area 907. Note that this is not a limitation, and the image selected by the user operation may be displayed in the seventh display area 907. In this case, the user device 2101 transmits information indicating the selected image, rather than information indicating the selected display area, to the image processing device 102.
[0202] In S2209, the image processing device 102 determines which image of the digital content has been selected from the icon corresponding to the selected display unit. Since the display area 902 was selected in S2208, the virtual viewpoint video corresponding to the display area 902 is selected. Therefore, the virtual viewpoint video included in the digital content is displayed in the seventh display area 907. If the digital content includes multiple virtual viewpoint videos, the virtual viewpoint video to be displayed first is set in advance. Through the above processing, information specifying the display area corresponding to the image selected by the user operation among the first display area 901 to the sixth display area 906 is received, and the display screen is updated by displaying the image or video corresponding to the selected display area in the seventh display area 907. Furthermore, when multiple virtual viewpoint videos are displayed in the second area 1308, as in the display screen shown in FIG. 13 of the sixth embodiment, it is determined whether a touch input or a flick input has been received for the displayed virtual viewpoint video. If a touch input has been received, the virtual viewpoint video being played is paused. When a flick input is received, the virtual viewpoint video currently being played back is switched to a different virtual viewpoint video displayed in the fifth display area 1305.
[0203] In S2210, the image processing apparatus 102 transmits the updated display image to the user device 2101.
[0204] Through the above process, a display image showing the digital content desired by the user can be created and displayed on the user device.
[0205] Although the present disclosure has been described in detail above based on a number of embodiments, the present disclosure is not limited to the above embodiments, and various modifications are possible based on the gist of the present disclosure, and are not excluded from the scope of the present disclosure. For example, the above-described embodiments 1 to 7 may be combined as appropriate.
[0206] Note that a computer program that realizes part or all of the control in this embodiment and the functions of the above-described embodiment may be supplied to an image processing system or the like via a network or various storage media. A computer (or a CPU, MPU, or the like) in the image processing system or the like may then read and execute the program. In this case, the program and the storage medium storing the program constitute the present disclosure.
[0207] The disclosure of this embodiment includes the following configurations, methods, and programs.
[0208] (Configuration 1) A means for identifying a virtual viewpoint image associated with a first surface of a three-dimensional digital content, the virtual viewpoint image being generated based on a plurality of images captured by a plurality of imaging devices and a virtual viewpoint, and an image of a viewpoint different from the virtual viewpoint corresponding to the virtual viewpoint image associated with a second surface of the digital content; a display control means for controlling the display of an image corresponding to the virtual viewpoint image and an image corresponding to an image of a viewpoint different from the virtual viewpoint in a display area; An apparatus comprising:
[0209] (Configuration 2) Further, the display device has an acquisition means for acquiring input information for selecting the display area, The display control means displays an image corresponding to the selected display area in a selected image display area that is a display area different from the selected display area, based on the input information acquired by the acquisition means. 2. The apparatus of claim 1,
[0210] (Configuration 3) The device described in configuration 1 or 2, characterized in that the digital content includes a plurality of virtual viewpoint images generated based on the plurality of images and a plurality of virtual viewpoints including the virtual viewpoint.
[0211] (Configuration 4) The device according to any one of configurations 1 to 3, wherein the display control means displays images corresponding to the plurality of virtual viewpoint images in different display areas.
[0212] (Configuration 5) Associating images corresponding to the plurality of virtual viewpoint images with specific display areas; The device described in any one of configurations 1 to 4, characterized in that when the display control means acquires an input from the acquisition means to select a display area in which to display images corresponding to the plurality of virtual viewpoint images, the display control means displays a specific virtual viewpoint image from the plurality of virtual viewpoint images in the selected image display area.
[0213] (Configuration 6) The device described in any one of configurations 1 to 5, characterized in that at least one of the multiple virtual viewpoints corresponding to the multiple virtual viewpoint images is determined based on a subject contained in at least one of the multiple images obtained by capturing images using the multiple imaging devices.
[0214] (Configuration 7) The device according to any one of configurations 1 to 6, wherein the position of the virtual viewpoint is determined based on the position of a three-dimensional shape representing the subject.
[0215] (Configuration 8) The device described in any one of configurations 1 to 7, characterized in that the position of at least one of the multiple virtual viewpoints corresponding to the multiple virtual viewpoint images is determined based on the position of the subject, and the line of sight direction from the virtual viewpoint is determined based on the orientation of the subject.
[0216] (Configuration 9) A device described in any one of configurations 1 to 7, characterized in that the position of at least one of the multiple virtual viewpoints corresponding to the multiple virtual viewpoint images is determined based on a position behind the subject at a predetermined distance, and the line of sight direction from the virtual viewpoint is determined based on the orientation of the subject.
[0217] (Configuration 10) A device described in any one of configurations 1 to 7, characterized in that the position of at least one of multiple virtual viewpoints corresponding to the multiple virtual viewpoint images is determined based on a position on a sphere centered on the subject, and the line of sight direction from the virtual viewpoint is determined based on the direction from the position of the virtual viewpoint toward the subject.
[0218] (Configuration 11) The subject is a person, 9. The device according to any one of configurations 1 to 8, wherein the orientation of the subject is the orientation of the subject's face.
[0219] (Configuration 12) The device described in any one of configurations 1 to 5, characterized in that when specific operation information is input to the selected image display area, the display control means switches between the virtual viewpoint image displayed in the selected image display area and a virtual viewpoint image among the multiple virtual viewpoint images that is different from the virtual viewpoint image displayed in the selected image display area.
[0220] (Configuration 13) The device described in Configuration 12, wherein the specific operation information is operation information relating to at least one of a keyboard typing operation, a mouse clicking operation, a mouse scrolling operation, a touch operation on a display device on which a virtual viewpoint image is displayed, a slide operation, a flick gesture, and a pinch-in / pinch-out operation.
[0221] (Configuration 14) The device described in Configuration 12 is characterized in that the display control means superimposes icons corresponding to each of the plurality of virtual viewpoint images on the selected image display area, and when input to select an icon is received, switches between the virtual viewpoint image displayed in the selected image display area and the virtual viewpoint image corresponding to the input icon.
[0222] (Configuration 15) The device according to any one of configurations 1 to 14, wherein the display control means displays an icon representing the virtual viewpoint image superimposed on an image representing the virtual viewpoint image.
[0223] (Method) A step of identifying a virtual viewpoint image associated with a first surface of a three-dimensional digital content, the virtual viewpoint image being generated based on a plurality of images captured by a plurality of imaging devices and a virtual viewpoint, and an image of a viewpoint different from the virtual viewpoint corresponding to the virtual viewpoint image associated with a second surface of the digital content; a display control step of controlling the display of an image corresponding to the virtual viewpoint image and an image corresponding to an image of a viewpoint different from the virtual viewpoint in a display area; An image processing method comprising:
[0224] (Program) A program for controlling each means described in the device described in any one of configurations 1 to 15 by a computer. [Explanation of symbols]
[0225] 1 camera 2 Shape estimation part 3. Image generation section 4 Content Generation Unit 5 Storage section 115 Display section 116 Operation section 100 Image processing device
Claims
1. a specifying means for specifying a virtual viewpoint image associated with a first surface of a three-dimensional digital content, the virtual viewpoint image being generated based on a plurality of images captured by a plurality of imaging devices and a virtual viewpoint, and an image of a viewpoint different from the virtual viewpoint corresponding to the virtual viewpoint image associated with a second surface of the digital content; a display control means for controlling the display of an image corresponding to the virtual viewpoint image and an image corresponding to an image of a viewpoint different from the virtual viewpoint in a display area; an acquisition means for acquiring input information for selecting the display area; and The display control means displays an image corresponding to the selected display area in a selected image display area that is a display area different from the selected display area, based on the input information acquired by the acquisition means. An image processing system characterized by:
2. 2. The image processing system according to claim 1, wherein the digital content includes a plurality of virtual viewpoint images generated based on the plurality of images and a plurality of virtual viewpoints including the virtual viewpoint.
3. 3. The image processing system according to claim 2, wherein the display control means displays images corresponding to the plurality of virtual viewpoint images in different display areas.
4. Associating images corresponding to the plurality of virtual viewpoint images with specific display areas; The image processing system according to claim 2, characterized in that when the display control means acquires an input from the acquisition means to select a display area in which to display images corresponding to the plurality of virtual viewpoint images, the display control means displays a specific virtual viewpoint image from the plurality of virtual viewpoint images in the selected image display area.
5. The image processing system according to claim 4, characterized in that at least one of the plurality of virtual viewpoints corresponding to the plurality of virtual viewpoint images is determined based on a subject contained in at least one of the plurality of images obtained by capturing images using the plurality of imaging devices.
6. 6. The image processing system according to claim 5, wherein the position of the virtual viewpoint is determined based on the position of a three-dimensional shape representing the subject.
7. The image processing system according to claim 2, characterized in that the position of at least one of the plurality of virtual viewpoints corresponding to the plurality of virtual viewpoint images is determined based on the position of the subject, and the line of sight direction from the virtual viewpoint is determined based on the orientation of the subject.
8. The image processing system according to claim 2, characterized in that the position of at least one of the plurality of virtual viewpoints corresponding to the plurality of virtual viewpoint images is determined based on a position that is a predetermined distance behind the subject, and the line of sight direction from the virtual viewpoint is determined based on the orientation of the subject.
9. The image processing system according to claim 2, characterized in that the position of at least one of the plurality of virtual viewpoints corresponding to the plurality of virtual viewpoint images is determined based on a position on a sphere centered on the subject, and the line of sight direction from the virtual viewpoint is determined based on a direction from the position of the virtual viewpoint toward the subject.
10. the subject is a person, 8. The image processing system according to claim 7, wherein the orientation of the subject is the orientation of the face of the subject.
11. The image processing system according to claim 4, characterized in that, when specific operation information is input to the selected image display area, the display control means switches between the virtual viewpoint image displayed in the selected image display area and a virtual viewpoint image among the plurality of virtual viewpoint images that is different from the virtual viewpoint image displayed in the selected image display area.
12. The image processing system according to claim 11, wherein the specific operation information is operation information relating to at least one of a keyboard typing operation, a mouse clicking operation, a mouse scrolling operation, a touch operation on a display device on which a virtual viewpoint image is displayed, a slide operation, a flick gesture, and a pinch-in / pinch-out operation.
13. The image processing system according to claim 11, characterized in that the display control means superimposes icons corresponding to each of the plurality of virtual viewpoint images on the selected image display area, and when input to select an icon is received, switches between the virtual viewpoint image displayed in the selected image display area and the virtual viewpoint image corresponding to the input icon.
14. 2. The image processing system according to claim 1, wherein the display control means displays an icon representing the virtual viewpoint image superimposed on an image representing the virtual viewpoint image.
15. an identifying step of identifying a virtual viewpoint image associated with a first surface of a three-dimensional digital content, the virtual viewpoint image being generated based on a plurality of images captured by a plurality of imaging devices and a virtual viewpoint, and an image of a viewpoint different from the virtual viewpoint corresponding to the virtual viewpoint image associated with a second surface of the digital content; a display control step of controlling the display of an image corresponding to the virtual viewpoint image and an image corresponding to an image of a viewpoint different from the virtual viewpoint in a display area; an acquisition step of acquiring input information for selecting the display area; and An image processing method characterized in that the display control step displays an image corresponding to the selected display area in a selected image display area, which is a display area different from the selected display area, based on the input information acquired by the acquisition step.
16. A computer program for controlling each unit of the image processing system according to any one of claims 1 to 14 by a computer.
Citation Information
Patent Citations
Electronic apparatus
JP2014044569A
Virtual viewpoint image generation device, virtual viewpoint image generation method, and virtual viewpoint image generation program
JP2015045920A
Information processing apparatus, information processing method, and program
JP2019179080A
Control device, control method and program
JP2020119095A