Image processing system, image processing method, and computer program
The image processing system generates virtual viewpoint images from multiple camera perspectives to create engaging three-dimensional digital content, addressing the lack of attractive content in existing technologies.
Patent Information
- Application Number
- JP2022037555
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-10
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2042-03-10
AI Technical Summary
Existing technologies fail to provide attractive digital content that includes virtual viewpoint images.
An image processing system that generates virtual viewpoint images by acquiring and processing images from multiple cameras with different viewpoints, estimating three-dimensional shapes, and associating these images to create three-dimensional digital content.
Enables the generation of attractive three-dimensional digital content, allowing users to freely view images from various perspectives and enhancing the appeal of the content.
Smart Images

Figure 0007746197000001 
Figure 0007746197000002 
Figure 0007746197000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image processing system, an image processing method, a computer program, and the like. [Background technology]
[0002] A technology that generates a virtual viewpoint image from a specified virtual viewpoint using a plurality of images captured by a plurality of imaging devices has attracted attention. Patent Document 1 describes a method of capturing images of a subject using a plurality of imaging devices installed at different positions, and generating a virtual viewpoint image using a three-dimensional shape of the subject estimated from the captured images. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-45920 Summary of the Invention [Problem to be solved by the invention]
[0004] However, it has not been possible to provide attractive digital content that includes virtual viewpoint images and other images. The present disclosure aims to provide an image processing system that can generate attractive content. [Means for solving the problem]
[0005] An image processing system according to one embodiment of the present disclosure includes: an acquisition means for acquiring a virtual viewpoint image generated based on a plurality of images captured by a plurality of imaging devices and a virtual viewpoint; before provisional a virtual viewpoint image; The aforementioned a generating means for generating a three-dimensional digital content including an image of a virtual viewpoint and an image of a different viewpoint; The generating means generates the virtual viewpoint image. The aforementioned Images from different viewpoints are associated with different surfaces that make up the digital content. 、 the images from different viewpoints are captured images obtained by capturing images using an imaging device different from the plurality of imaging devices, The virtual viewpoint image and the captured image include the same subject. It is characterized by: [Effects of the Invention]
[0006] According to the present disclosure, it is possible to provide an image processing system that can generate attractive content. [Brief explanation of the drawings]
[0007] [Figure 1] 1 is a diagram showing an example of the device configuration of an image processing system 100 according to a first embodiment. [Figure 2] 1 is a diagram showing a hardware configuration of an image processing system 100 according to a first embodiment. [Figure 3] 1 is a flowchart illustrating the operation flow of the image processing system 100 according to the first embodiment. [Figure 4] 4A to 4C are diagrams showing examples of stereoscopic images as content generated by the content generation unit 4 in the first embodiment. [Figure 5] 10 is a flowchart illustrating the operation flow of the image processing system 100 according to the second embodiment. [Figure 6] 10 is a flowchart illustrating the operation flow of the image processing system 100 according to the third embodiment. [Figure 7] This is a continuation of the flowchart in Figure 6. [Figure 8] 8 is a continuation of the flowchart of FIG. 6 and FIG. 7. DETAILED DESCRIPTION OF THE INVENTION
[0008] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. However, the present disclosure is not limited to the following embodiments. In each drawing, the same members or elements are designated by the same reference numerals, and duplicate descriptions will be omitted or simplified.
[0009] (Embodiment 1) The image processing system of the first embodiment generates a virtual viewpoint image seen from a virtual viewpoint based on captured images captured by a plurality of imaging devices (cameras) from different directions, the state of the imaging devices, and a specified virtual viewpoint. The virtual viewpoint image is then displayed on the surface of a virtual three-dimensional image. Note that the imaging devices may have not only cameras but also functional units for performing image processing. Furthermore, the imaging devices may have sensors for acquiring distance information in addition to cameras.
[0010] The multiple cameras capture images of the imaging area from multiple directions. The imaging area may be, for example, an area surrounded by a stadium field at an arbitrary height. The imaging area may correspond to the three-dimensional space from which the three-dimensional shape of the subject is estimated. The three-dimensional space may be the entire imaging area or a part of it. The imaging area may also be a concert venue, an imaging studio, etc.
[0011] The multiple cameras are installed in different positions and directions (attitudes) surrounding the imaging area, and capture images synchronously. Note that the multiple cameras do not need to be installed around the entire circumference of the imaging area, and may be installed in only a part of the imaging area depending on installation location restrictions, etc. The number of cameras is not limited, and for example, if the imaging area is a rugby stadium, several tens to several hundreds of cameras may be installed around the stadium.
[0012] The multiple cameras may also include cameras with different angles of view, such as a telephoto camera and a wide-angle camera. For example, using a telephoto camera to capture high-resolution images of players can improve the resolution of the generated virtual viewpoint image. In ball games, the ball moves over a wide range, so using a wide-angle camera for capture can reduce the number of cameras. Combining the imaging areas of a wide-angle camera and a telephoto camera for capture can improve the degree of freedom in installation position. The cameras are synchronized to a common time, and capture time information is added to each captured image frame.
[0013] A virtual viewpoint image, also called a free viewpoint image, allows a user to monitor an image corresponding to a viewpoint freely (arbitrarily) designated by the user. For example, a virtual viewpoint image includes an image corresponding to a viewpoint selected by the user from a limited number of viewpoint candidates. The virtual viewpoint may be designated by a user operation or automatically by AI based on the results of image analysis, etc. The virtual viewpoint image may be a video or a still image.
[0014] The virtual viewpoint information used to generate a virtual viewpoint image is information including the position and direction (posture) of the virtual viewpoint, as well as the angle of view (focal length), etc. Specifically, the virtual viewpoint information includes parameters representing the three-dimensional position of the virtual viewpoint, parameters representing the direction (line of sight) from the virtual viewpoint in the pan, tilt, and roll directions, focal length information, etc. However, the contents of the virtual viewpoint information are not limited to those described above.
[0015] The virtual viewpoint information may also have parameters for each of a plurality of frames. In other words, the virtual viewpoint information may have parameters corresponding to each of a plurality of frames constituting a moving image of a virtual viewpoint image, and may be information indicating the position and orientation of the virtual viewpoint at each of a plurality of consecutive points in time.
[0016] A virtual viewpoint image is generated, for example, by the following method. First, multiple camera images are acquired by capturing images from different directions using cameras. Next, a foreground image is acquired from the multiple camera images, extracting a foreground area corresponding to a subject such as a person or a ball, and a background image is acquired from the background area other than the foreground area. The foreground image and background image contain texture information (such as color information).
[0017] Then, a foreground model representing the three-dimensional shape of the subject and texture data for coloring the foreground model are generated based on the foreground image. Also, texture data for coloring a background model representing the three-dimensional shape of the background, such as a stadium, is generated based on the background image. Then, the texture data is mapped onto the foreground model and background model, and rendering is performed according to the virtual viewpoint indicated by the virtual viewpoint information, thereby generating a virtual viewpoint image.
[0018] However, the method for generating the virtual viewpoint image is not limited to this, and various methods can be used, such as a method for generating the virtual viewpoint image by projective transformation of the captured image without using a foreground or background model.
[0019] A foreground image is an image in which a subject area (foreground area) is extracted from an image captured by a camera. A subject extracted as a foreground area refers to a dynamic subject (moving object) that moves (its absolute position and shape may change) when images are captured from the same direction in a time series. For example, in a sport, the subject includes people such as players and referees on the field where the sport is being played, and in a ball game, it includes not only people but also the ball. In concerts and entertainment events, foreground subjects include singers, musicians, performers, and presenters.
[0020] A background image is an image of at least an area (background area) different from the foreground subject. Specifically, a background image is an image in which the foreground subject has been removed from the captured image. The background also refers to an object that remains stationary or nearly stationary when images are captured from the same direction in chronological order.
[0021] Such imaging objects include, for example, a stage for a concert, a stadium where an event such as a sport is held, a structure such as a goal used in a ball game, a field, etc. However, the background is at least an area different from the subject, which is the foreground. Note that the imaging objects may include other objects in addition to the subject and background.
[0022] Fig. 1 is a diagram showing an image processing system 100 according to this embodiment. Some of the functional blocks shown in Fig. 1 are realized by causing a computer included in the image processing system 100 to execute a computer program stored in a memory serving as a storage medium. However, some or all of these functions may be realized by hardware. Examples of hardware that can be used include a dedicated circuit (ASIC) and a processor (reconfigurable processor, DSP).
[0023] Furthermore, the individual functional blocks of the image processing system 100 do not have to be built into the same housing, but may be configured as separate devices connected to each other via signal paths. The image processing system 100 is connected to multiple cameras. The image processing system 100 also includes a shape estimation unit 2, an image generation unit 3, a content generation unit 4, a storage unit 5, a display unit 115, an operation unit 116, etc.
[0024] The shape estimation unit 2 is connected to a plurality of cameras 1 and an image generation unit 3, and the display unit 115 is connected to the image generation unit 3. Each functional block may be implemented in a separate device, or all or some of the functional blocks may be implemented in the same device.
[0025] Multiple cameras 1 are placed at different locations around a stage for a concert or other event, a stadium for a sporting event, a structure such as a goal used in a ball game, or a field, and each camera captures images from a different viewpoint. Each camera has an identification number (camera number) to identify the camera. Camera 1 may also include other functions, such as a function to extract a foreground image from captured images, and hardware (circuits, devices, etc.) that realizes those functions. The camera number may be set based on the installation location of camera 1, or may be set based on other criteria.
[0026] The image processing system 100 may be located within the venue where the camera 1 is installed, or may be located outside the venue, for example, at a broadcasting station. The image processing system 100 is connected to the camera 1 via a network. The shape estimation unit 2 acquires images from the multiple cameras 1. Then, the shape estimation unit 2 estimates the three-dimensional shape of the subject based on the images acquired from the multiple cameras 1.
[0027] Specifically, the shape estimation unit 2 generates three-dimensional shape data expressed using a known expression method. The three-dimensional shape data may be point cloud data made up of points, mesh data made up of polygons, or voxel data made up of voxels.
[0028] The image generation unit 3 acquires information indicating the position and posture of the three-dimensional shape data of the subject from the shape estimation unit 2, and can generate a virtual viewpoint image including the two-dimensional shape of the subject that is expressed when the three-dimensional shape of the subject is viewed from a virtual viewpoint. In order to generate the virtual viewpoint image, the image generation unit 3 can also accept designation of virtual viewpoint information (such as the position of the virtual viewpoint and the line of sight from the virtual viewpoint) from the user, and generate the virtual viewpoint image based on the virtual viewpoint information. Here, the image generation unit 3 functions as an acquisition means that generates a virtual viewpoint image based on multiple images acquired from multiple cameras.
[0029] The virtual viewpoint image is sent to the content generation unit 4, and the content generation unit 4 generates, for example, a three-dimensional digital content as described later. The digital content including the virtual viewpoint image generated by the content generation unit 4 is output to the display unit 115.
[0030] The content generation unit 4 can also directly receive images from multiple cameras and supply the image from each camera to the display unit 115. Furthermore, based on an instruction from the operation unit 116, it can also switch on which surface of the virtual stereoscopic image the image from each camera and the virtual viewpoint image are to be displayed.
[0031] The display unit 115 is configured by, for example, a liquid crystal display or an LED, and acquires and displays digital content including virtual viewpoint images from the content generation unit 4. It also displays a GUI (Graphical User Interface) and the like for the user to operate each camera 1.
[0032] The operation unit 116 is composed of a joystick, a jog dial, a touch panel, a keyboard, a mouse, etc., and is used by the user to operate the camera 1, etc. The operation unit 116 is also used by the user to select an image to be displayed on the surface of the digital content (stereoscopic image) generated by the content generation unit 4. Furthermore, the operation unit 116 can be used to specify the position and posture of the virtual viewpoint for generating a virtual viewpoint image in the image generation unit 3.
[0033] The position and orientation of the virtual viewpoint may be directly specified on the screen by a user's operation instruction, or when a predetermined subject is specified on the screen by a user's operation instruction, the predetermined subject may be tracked through image recognition, and virtual viewpoint information from the subject or virtual viewpoint information from positions around the subject in an arc shape centered on the subject may be automatically specified.
[0034] Furthermore, the system may be configured to automatically specify virtual viewpoint information from the subject or from positions around the subject in an arc, based on image recognition of a subject that meets pre-specified conditions specified by a user's operation. In this case, the specified conditions may include, for example, the name of a specific athlete, the shooter, the performer of a fine play, the position of the ball, etc.
[0035] The storage unit 5 includes a memory for storing the digital content generated by the content generation unit 4, the virtual viewpoint images, the camera images, etc. The storage unit 5 may also have a removable recording medium. The removable recording medium may store, for example, a plurality of camera images captured at other venues or other sporting scenes, a virtual viewpoint image generated using those images, or digital content generated by combining those images.
[0036] The storage unit 5 may also be configured to store multiple camera images downloaded from an external server or the like via a network, virtual viewpoint images generated using those images, digital content generated by combining those images, etc. Those camera images, virtual viewpoint images, digital content, etc. may also be created by a third party.
[0037] FIG. 2 is a diagram showing the hardware configuration of the image processing system 100 according to the first embodiment, and the hardware configuration of the image processing system 100 will be described with reference to FIG.
[0038] The image processing system 100 includes a CPU 111, a ROM 112, a RAM 113, an auxiliary storage device 114, a display unit 115, an operation unit 116, a communication I / F 117, and a bus 118. The CPU 111 controls the entire image processing system 100 using computer programs stored in the ROM 112, the RAM 113, the auxiliary storage device 114, etc., thereby realizing each functional block of the image processing system shown in FIG.
[0039] The RAM 113 temporarily stores computer programs and data supplied from the auxiliary storage device 114, and data supplied from the outside via the communication I / F 117. The auxiliary storage device 114 is configured, for example, with a hard disk drive or the like, and stores various data such as image data, audio data, and digital content including virtual viewpoint images from the content generation unit 4.
[0040] As described above, the display unit 115 displays digital content including virtual viewpoint images, a GUI, etc. As described above, the operation unit 116 receives operation inputs from the user and inputs various instructions to the CPU 111. The CPU 111 operates as a display control unit that controls the display unit 115 and an operation control unit that controls the operation unit 116.
[0041] The communication I / F 117 is used for communication with devices external to the image processing system 100 (for example, the camera 1 or an external server). For example, if the image processing system 100 is connected to an external device via a wired connection, a communication cable is connected to the communication I / F 117. If the image processing system 100 has a function for wireless communication with an external device, the communication I / F 117 is equipped with an antenna. The bus 118 connects each unit of the image processing system 100 to transmit information.
[0042] In this embodiment, an example is shown in which the display unit 115 and the operation unit 116 are included inside the image processing system 100, but at least one of the display unit 115 and the operation unit 116 may exist as a separate device outside the image processing system 100. In addition, the image processing system 100 may be in the form of, for example, a PC terminal.
[0043] Fig. 3 is a flowchart for explaining the operation flow of the image processing system 100 of the embodiment 1. Also, Figs. 4(A) to 4(C) are diagrams showing examples of three-dimensional digital content generated by the content generation unit 4 in the embodiment 1. The CPU 111 serving as a computer in the image processing system 100 executes a computer program stored in a memory such as the ROM 112 or the auxiliary storage device 114, thereby performing the operations of the steps in the flowchart of FIG.
[0044] In this embodiment, the image processing system 100 is installed in a broadcasting station or the like, and may produce and broadcast cubic digital content 200 as shown in Fig. 4(A), or may provide it via the Internet. In this case, an NFT can be assigned to the digital content 200.
[0045] That is, in order to increase asset value, scarcity can be achieved by, for example, limiting the amount of content distributed and managing it with a serial number. NFT stands for Non-fungible Token, and is a token issued and circulated on a blockchain. Examples of NFT formats include token standards called ERC-721 and ERC-1155. Tokens are usually stored in association with a wallet managed by the user.
[0046] In step S31, CPU 111 associates the main camera image (first image) with first surface 201 of digital content 200 having a three-dimensional shape, for example, as shown in Fig. 4(A). Note that the main camera image associated with first surface 201 may be displayed for user confirmation. Furthermore, as shown in Fig. 4(A), if the line of sight from a viewpoint virtually viewing the digital content (specifically, the direction perpendicular to the paper surface of Fig. 4(A)) is not parallel to the normal direction of first surface 201, the following display may be performed.
[0047] In other words, the main camera image displayed on first surface 201 may be generated by projective transformation according to the angle of the normal direction of first surface 201 relative to the display surface of the digital content. Here, the main camera image (main image, first image) is an image selected for TV broadcasting or the like from among multiple images obtained from multiple cameras installed at a sports venue. Note that the main image is an image that includes a predetermined subject within its angle of view. Furthermore, the main camera image does not have to be captured by a camera installed at the sports venue.
[0048] For example, the image may be captured by a handheld camera brought by a cameraman, or may be captured by a camera brought by a spectator in the venue or an electronic device such as a smartphone equipped with a camera. The main camera image may be one of the multiple cameras used to generate the virtual viewpoint image, or may be a camera not included in the multiple cameras.
[0049] The user of the broadcasting station or the like sequentially selects which camera's image to broadcast or distribute over the Internet as the main image using the operation unit 116. For example, when broadcasting or distributing the moment a goal is scored, the image from a camera near the goal is often broadcast as the main image.
[0050] In this embodiment, as shown in Figures 4(A) to 4(C), the surface seen on the left side is the first surface, the surface seen on the right side is the second surface, and the surface seen on the upper side is the third surface. However, this is not limited to this. Which surfaces are the first to third surfaces can be arbitrarily set in advance.
[0051] In step S32, the content generation unit 4 associates data such as the name of the player who shot the goal, the name of the team he belongs to, and the final game result of the game in which he scored, as associated data with the third side 203 of the digital content 200. Note that the associated data associated with the third side 203 may be displayed for user confirmation. When an NFT is granted, data indicating its rarity, such as the number of issues, may be displayed as associated data on the third side 203. The number of issues may be determined by a user who generates the digital content using the image generation system, or may be determined automatically by the image generation system.
[0052] In step S33, the image generation unit 3 acquires, from among images from the multiple cameras 1, an image that includes, for example, a goal or a shooter, and whose viewpoint direction differs by a predetermined angle from the viewpoint direction of the camera capturing the main camera image. The predetermined angle is, for example, 90 degrees. In this case, since the positions and attitudes of the multiple cameras are known in advance, the CPU 111 can determine from which camera an image with a viewpoint direction that differs by the predetermined angle from the main camera image as described above can be acquired. Note that, although the expression "viewpoint of an image" is sometimes used below, this refers to the viewpoint of the camera capturing the image or a virtual viewpoint specified for generating the image.
[0053] Alternatively, in step S33 (acquisition step), a virtual viewpoint image from a virtual viewpoint that includes an image-recognized subject and is a predetermined virtual viewpoint (for example, a viewpoint that differs by 90 degrees as described above) may be acquired in the image generation unit 3. In this case, the predetermined virtual viewpoint (a viewpoint that differs by 90 degrees as described above, i.e., a posture that differs) may be specified for the image generation unit 3, and the virtual viewpoint image may be generated and acquired.
[0054] Alternatively, virtual viewpoint images from a plurality of viewpoints may be generated in advance by the image generation unit 3, and the corresponding image may be selected from among them. In this embodiment, an image with a viewpoint that differs by a predetermined angle from the main camera image is an image with a viewpoint that differs by, for example, 90 degrees, but this angle can be set in advance.
[0055] Furthermore, the virtual viewpoint image may be an image corresponding to a virtual viewpoint specified based on the orientation of a subject included in the main camera image (for example, the orientation of the face or body of a person). Note that if there are multiple subjects included in the main camera image, the virtual viewpoint may be set for one of the subjects, or virtual viewpoints may be set for multiple subjects.
[0056] Although the above description has been given of an example in which a viewpoint at a predetermined angle is selected with respect to the main image, it is also possible to select and acquire a virtual viewpoint image from a predetermined viewpoint, such as the subject viewpoint, a viewpoint from behind the subject, or one of the virtual viewpoints located on an arc centered on the subject.
[0057] The subject viewpoint is a virtual viewpoint in which the position of the subject is the position of the virtual viewpoint and the direction of the subject is the line of sight from the virtual viewpoint. For example, when a person is the subject, the subject viewpoint is a viewpoint in which the position of the person's face is the position of the virtual viewpoint and the direction of the person's face is the line of sight from the virtual viewpoint. Alternatively, the line of sight of the person may be the line of sight from the virtual viewpoint.
[0058] A viewpoint from behind a subject is a virtual viewpoint in which a position at a predetermined distance behind the subject is set as the virtual viewpoint, and the direction from that position toward the subject's position is set as the line of sight from the virtual viewpoint. Alternatively, the line of sight from the virtual viewpoint may be determined according to the orientation of the subject. For example, when a person is set as the subject, a viewpoint from behind the subject is a virtual viewpoint in which a position at a predetermined distance from the person's back is set as the line of sight from the virtual viewpoint, and the orientation of the person's face is set as the line of sight from the virtual viewpoint.
[0059] A virtual viewpoint within a position on an arc centered on the subject is a virtual viewpoint that is set at a position on a sphere defined by a predetermined radius centered on the position of the subject, and the direction from that position toward the position of the subject is the line of sight from the virtual viewpoint. For example, when a person is the subject, the virtual viewpoint is a position on a sphere defined by a predetermined radius centered on the position of the person, and the direction from that position toward the position of the subject is the line of sight from the virtual viewpoint.
[0060] Step S33 thus functions as a virtual viewpoint image generating step for acquiring a virtual viewpoint image of a viewpoint having a predetermined relationship with the first image and using it as the second image. Note that the virtual viewpoint image of a viewpoint having a predetermined relationship with the first image is a virtual viewpoint image captured at the same time (image capturing timing) as the viewpoint of the first image. In this embodiment, the viewpoint having a predetermined relationship with the first image is a viewpoint that has a predetermined angular relationship or a predetermined positional relationship with the viewpoint of the first image, as described above.
[0061] Then, in step S34, CPU 111 associates the second image with second side 202 of digital content 200. Note that the second image may be displayed for user confirmation. Note that at this time, the main image associated with first side 201 and the second image associated with the second side are synchronized and controlled so that they are images captured at the same time, as described above.
[0062] In this way, steps S31 to S34 associate the first image with the first surface of the three-dimensional digital content, which will be described later, and associate the virtual viewpoint image of a virtual viewpoint having a predetermined relationship with the first image with the second surface 202. Furthermore, steps S31 to S34 function as a content generation step (content generation means).
[0063] Next, in step S35, CPU 111 determines whether an operation to change the viewpoint of the second image displayed on second screen 202 has been performed via operation unit 116. That is, while watching a sports scene that changes from moment to moment, the user may change the viewpoint of the second image to be displayed on second screen 202, for example, by selecting a camera image with a desired viewpoint from among multiple cameras 1.
[0064] Alternatively, a virtual viewpoint image from a desired viewpoint may be acquired by specifying the desired viewpoint from among the virtual viewpoints to the image generator 3. In step S35, if such a viewpoint change operation is performed, the result becomes Yes, and the process proceeds to step S36.
[0065] In step S36, the CPU 111 selects a viewpoint image after the viewpoint has been changed from among the multiple cameras 1, or acquires a virtual viewpoint image after the viewpoint has been changed from the image generation unit 3. The virtual viewpoint image may be acquired by acquiring a virtual viewpoint image that has been generated in advance, or by generating a new virtual viewpoint image based on the changed viewpoint. The acquired image is then set as the second image, and the process proceeds to step S34 to associate it with the second plane.
[0066] In this state, the first image, the second image, and the associated data are associated with the first to third sides of the digital content 200, respectively, on the display unit 115. Here, the user may confirm, from the display, the state in which the first image, the second image, and the associated data are associated with the first to third sides of the digital content 200, respectively. In this case, the numbers of the sides may also be displayed so that it is clear which sides are the first to third sides.
[0067] If there is no viewpoint change in step S35, the process proceeds to step S37, where the CPU 111 determines whether or not to assign an NFT to the digital content 200. To do this, for example, a GUI is displayed on the display unit 115, asking whether or not to assign an NFT to the digital content 200. If the user selects to assign an NFT, this is determined, and the process proceeds to step S38, where the NFT is assigned to the digital content 200, encryption is performed, and the process proceeds to step S39.
[0068] If the answer in step S37 is No, the process proceeds directly to step S39. The digital content 200 in step S37 may be a three-dimensional image having a shape such as that shown in Figures 4(B) and 4(C). In addition, in the case of a polyhedron, it is not limited to a hexahedron as shown in Figure 4(A) and may be, for example, an octahedron.
[0069] In step S39, CPU 111 determines whether or not to end the flow for generating digital content 200 in Fig. 3. If the user has not operated operation unit 116 to end the flow, the process returns to step S31 and repeats the above processing, and if the flow is to end the flow in Fig. 3. Note that even if the user has not operated operation unit 116 to end the flow, the flow may be automatically ended when a predetermined period of time (e.g., 30 minutes) has elapsed since the last operation of operation unit 116.
[0070] 4(B) and (C) are diagrams showing modified examples of digital content 200, with Fig. 4(B) being the digital content 200 of Fig. 4(A) in the form of a sphere. For example, when viewed from the front of spherical digital content 200, a first image is displayed on first surface 201, which is the left spherical surface, and a second image is displayed on second surface 202, which is the right spherical surface. Furthermore, the above-mentioned associated data is displayed on third surface 203, which is the upper spherical surface.
[0071] Fig. 4(C) is a diagram showing an example in which each plane of the digital content 200 in Fig. 4(A) is curved to a desired curvature. In this way, the digital content in this embodiment may be displayed using a sphere as shown in Fig. 4(B) or (C), or a cube with a spherical surface, to display an image.
[0072] (Embodiment 2) Next, a second embodiment will be described with reference to FIG. Fig. 5 is a flowchart for explaining the operation flow of the image processing system 100 of embodiment 2. Note that the operation of each step in the flowchart of Fig. 5 is performed by the CPU 111 serving as the computer of the image processing system 100 executing a computer program stored in a memory such as the ROM 112 or the auxiliary storage device 114.
[0073] In FIG. 5, steps with the same reference numerals as those in FIG. 3 are the same processes, and the explanation thereof will be omitted. 5, the CPU 111 acquires a camera image from a viewpoint designated by the user or a virtual viewpoint image from a virtual viewpoint designated by the user from the image generation unit 3. The acquired image is then set as the second image. The rest of the process is the same as the flow in FIG. 3.
[0074] In the first embodiment, a second image having a predetermined relationship (different angle) with respect to the main image (first image) is acquired. However, in the second embodiment, the user selects a desired camera or acquires a virtual viewpoint image of a desired subject from a desired viewpoint, and uses it as the second image.
[0075] The camera image or virtual viewpoint image selected by the user in step S51 includes, for example, an image of a bird's-eye view from diagonally above the sports venue or an image of a viewpoint from diagonally below the sports venue, etc. In this way, in the second embodiment, the user can select the virtual viewpoint image to be displayed on the second screen.
[0076] Furthermore, the virtual viewpoint image selected by the user in step S51 may be a virtual viewpoint image of a viewpoint from a position distant from the subject, such as a zoomed-out image. Additionally, previously generated camera images and virtual viewpoint images generated based on them may be stored in the storage unit 5, and these may be read out and displayed on the first to third screens as the first image, second image, and associated data, respectively.
[0077] Between steps S38 and S39, a step may be inserted in which the CPU 111 automatically switches to a default stereoscopic image display when a predetermined period of time (e.g., 30 minutes) has elapsed since the last operation of the operation unit 116. The default stereoscopic image display may display, for example, a main image on the first screen, associated data on the third screen, and a camera image or a virtual viewpoint image from the viewpoint most frequently used in past statistics on the second screen.
[0078] (Embodiment 3) The third embodiment will be described with reference to Figures 6 to 8. Figure 6 is a flowchart for explaining the operation flow of image processing system 100 of the third embodiment, Figure 7 is a flowchart continuing from Figure 6, and Figure 8 is a flowchart continuing from Figures 6 and 7. Note that the operation of each step in the flowcharts of Figures 6 to 8 is performed by CPU 111, which serves as a computer in image processing system 100, executing a computer program stored in a memory such as ROM 112 or auxiliary storage device 114.
[0079] In the third embodiment, when the user selects the number of virtual viewpoints from 1 to 3, the display of the first to third planes of the digital content 200 is automatically switched accordingly. In step S61, the user selects the number of virtual viewpoints from among 1 to 3, and this is accepted by the CPU 111. In step S62, the CPU 111 acquires the selected number of virtual viewpoint images from the image generation unit 3.
[0080] At this time, a representative virtual viewpoint is automatically selected. That is, the scene is analyzed, and the virtual viewpoint that is most frequently used for that scene, for example, based on past statistics, is set as the first virtual viewpoint, the next most frequently used virtual viewpoint is set as the second virtual viewpoint, and the next most frequently used virtual viewpoint is set as the third virtual viewpoint. Note that the second virtual viewpoint may be set in advance to differ, for example, by +90° from the first virtual viewpoint, and the third virtual viewpoint may be set in advance to differ, for example, by -90° from the first virtual viewpoint. Here, +90° and -90° are examples, and the present invention is not limited to these angles.
[0081] In step S63, the CPU 111 determines whether the number of selected virtual viewpoints is 1, and if so, proceeds to step S64. In step S64, the CPU 111 acquires a main image from a main camera among the multiple cameras 1 and associates it with the first level 201 of the digital content 200.
[0082] Then, in step S65, the CPU 111 associates the associated data with the third surface 203 of the digital content 200. The associated data may be the same as the associated data displayed in step S32 of FIG. 3 in the first embodiment, such as the name of the player who shot the ball into the goal. In step S66, the CPU 111 associates the first virtual viewpoint image from the first virtual viewpoint described above with the second surface of the digital content 200, and then proceeds to step S81 in FIG.
[0083] If the determination in step S63 is No, the CPU 111 determines in step S67 whether or not the number of selected virtual viewpoints is two, and if it is two, the process proceeds to step S68. In step S68, the CPU 111 associates the associated data with the third side 203 of the digital content 200. The associated data may be the same as the associated data associated in step S65, such as the name of the player who shot the ball into the goal.
[0084] Then, in step S69, the CPU 111 associates the first virtual viewpoint image from the first virtual viewpoint with the first plane 201 of the digital content 200. Also, the CPU 111 associates the second virtual viewpoint image from the second virtual viewpoint with the second plane 202 of the digital content 200. Thereafter, the process proceeds to step S81 in FIG. 8.
[0085] If it is determined No in step S67, the process proceeds to step S71 in Fig. 7. In step S71, CPU 111 determines whether or not the user has selected to associate associated data with third surface 203. If Yes, the process proceeds to step S72, and if No, the process proceeds to step S73.
[0086] In step S72, CPU 111 associates the virtual viewpoint image from the first virtual viewpoint with first plane 201 of digital content 200, associates the virtual viewpoint image from the second virtual viewpoint with second plane 202, and associates the virtual viewpoint image from the third virtual viewpoint with third plane 203. Then, the process proceeds to step S81 in FIG. 8.
[0087] If the answer is No in step S71, in step S73, the CPU 111 associates the associated data with the third side 203 of the digital content 200. The associated data may be the same as the associated data associated in step S65, such as the name of the player who shot the ball into the goal.
[0088] Then, in step S74, CPU 111 associates the first virtual viewpoint image from the first virtual viewpoint with first plane 201 of digital content 200. Furthermore, in step S75, CPU 111 associates the second virtual viewpoint image from the second virtual viewpoint and the third virtual viewpoint image from the third virtual viewpoint with second plane 202 of digital content 200 so that they can be displayed side by side. That is, second plane 202 is divided into two areas for displaying the second virtual viewpoint image and the third virtual viewpoint image, and a virtual viewpoint image is associated with each area. Thereafter, the process proceeds to step S81 in FIG. 8.
[0089] 8, the CPU 111 determines whether or not to assign an NFT to the digital content 200. To do this, for example, a GUI is displayed on the display unit 115, asking whether or not to assign an NFT to the digital content 200. If the user selects to assign an NFT, the process proceeds to step S82, where the NFT is assigned to the digital content 200, encryption is performed, and the process proceeds to step S83.
[0090] If the answer in step S81 is No, the process proceeds directly to step S83. Note that the digital content 200 in step S81 may be in the form shown in Figures 4(B) and 4(C), as described above. In step S83, CPU 111 determines whether or not to end the flow of FIGS. 6 to 8, and if the user has not operated operation unit 116 to end it, the process proceeds to step S84.
[0091] In step S84, the CPU 111 determines whether the number of virtual viewpoints has changed. If the number has changed, the process returns to step S61. If the number has not changed, the process returns to step S62. If the determination in step S83 is Yes, the flow of FIGS. 6 to 8 ends.
[0092] In the third embodiment, an example has been described in which, when the user selects the number of virtual viewpoints from 1 to 3, the images to be associated with the first to third planes of digital content 200 are automatically switched accordingly. However, the user may also select the number of camera images to be associated with the planes constituting digital content 200 from among images from multiple cameras. Then, a predetermined camera may be automatically selected from multiple cameras 1 accordingly, and the camera images from that camera may be automatically associated with the first to third planes of digital content 200.
[0093] The maximum number of viewpoints does not have to be three. For example, the number of viewpoints may be determined within a range that maximizes the number of surfaces constituting the digital content or the number of surfaces to which images can be associated. Furthermore, if multiple images can be associated with one surface, the maximum number of viewpoints can be further increased.
[0094] Furthermore, between steps S82 and S83, a step may be inserted in which the CPU 111 automatically switches to content consisting of a default stereoscopic image display when a predetermined period (e.g., 30 minutes) has elapsed since the last operation of the operation unit 116. The default stereoscopic image display displays, for example, a main image on the first plane, and a camera image or a virtual viewpoint image from the viewpoint most frequently used according to past statistics on the second plane. The third plane may be, for example, associated data.
[0095] As described above, in the third embodiment, in steps S69, S72, and S95, it is possible to associate a virtual viewpoint image different from the virtual viewpoint image to be displayed on the second surface with the first surface.
[0096] Although the present disclosure has been described in detail above based on a number of embodiments, the present disclosure is not limited to the above embodiments, and various modifications are possible based on the gist of the present disclosure, and are not excluded from the scope of the present disclosure. For example, the above-described embodiments 1 to 3 may be combined as appropriate. Furthermore, a plurality of virtual viewpoint images with different viewpoints may be associated with one surface constituting the digital content.
[0097] Note that a computer program that realizes part or all of the control in this embodiment and the functions of the above-described embodiment may be supplied to an image processing system or the like via a network or various storage media. A computer (or a CPU, MPU, or the like) in the image processing system or the like may then read and execute the program. In this case, the program and the storage medium storing the program constitute the present disclosure. [Explanation of symbols]
[0098] 1 camera 3. Image generation section 4 Content Generation Unit 100 Image Processing System
Claims
1. an acquisition means for acquiring a virtual viewpoint image generated based on a plurality of images captured by a plurality of imaging devices and a virtual viewpoint; a generating means for generating a three-dimensional digital content including the virtual viewpoint image and an image of a viewpoint different from the virtual viewpoint, the generating means associates the virtual viewpoint image and the different viewpoint image with different surfaces constituting the digital content; the images from different viewpoints are captured images obtained by capturing images using an imaging device different from the plurality of imaging devices, An image processing system, wherein the virtual viewpoint image and the captured image include the same subject.
2. 2. The image processing system according to claim 1, wherein the virtual viewpoint image is an image having a predetermined relationship with the captured image.
3. 3. The image processing system according to claim 2, wherein the timing of capturing the image used to generate the virtual viewpoint image is the same as the timing of capturing the captured image.
4. 4. The image processing system according to claim 2, wherein the virtual viewpoint has a predetermined angular relationship or a predetermined positional relationship with the viewpoint of the different image capturing device.
5. 5. The image processing system according to claim 1, wherein the position and orientation of the virtual viewpoint are determined based on at least the position and orientation of the subject.
6. 6. The image processing system according to claim 5, wherein the position of the virtual viewpoint is determined based on the position or orientation of a three-dimensional shape representing the subject.
7. the position of the virtual viewpoint is the position of the subject, 7. The image processing system according to claim 5, wherein the line of sight direction from the virtual viewpoint is determined based on the orientation of the subject.
8. the subject is a person, 8. The image processing system according to claim 7, wherein the orientation of the subject is the orientation of the face of the subject.
9. the position of the virtual viewpoint is a position at a predetermined distance behind the subject, 7. The image processing system according to claim 5, wherein the line of sight from the virtual viewpoint is a direction toward the subject from a position a predetermined distance behind the subject.
10. the position of the virtual viewpoint is a position on an arc of a circle having the position of the subject as its center, 7. The image processing system according to claim 5, wherein the line of sight from the virtual viewpoint is a direction from a position on an arc of a circle centered on the position of the subject toward the subject. Hmm.
11. 11. The image processing system according to claim 1, wherein another virtual viewpoint image is also associated with the surface to which the virtual viewpoint image is associated.
12. 12. The image processing system according to claim 1, wherein the digital content has a three-dimensional shape with a plurality of surfaces, and the virtual viewpoint image and the captured image are displayed on different surfaces.
13. the subject is an athlete, the captured image and the virtual viewpoint image include a specific play of the athlete, 13. The image processing system according to claim 1, wherein the digital content includes information indicating the specific play.
14. 14. The image processing system according to claim 1, wherein the generating unit assigns an NFT (Non-fungible Token) to the digital content.
15. an acquisition step of acquiring a virtual viewpoint image generated based on a plurality of images obtained by capturing images using a plurality of imaging devices and a virtual viewpoint; a generating step of generating a three-dimensional digital content including the virtual viewpoint image and an image of a viewpoint different from the virtual viewpoint, In the generating step, the virtual viewpoint image and the different viewpoint image are associated with different surfaces constituting the digital content; the images from different viewpoints are captured images obtained by capturing images using an imaging device different from the plurality of imaging devices, An image processing method, wherein the virtual viewpoint image and the captured image include the same subject.
16. A computer program for controlling each unit of the image processing system according to any one of claims 1 to 14 by a computer.
Citation Information
Patent Citations
Virtual viewpoint image generation device, virtual viewpoint image generation method, and virtual viewpoint image generation program
JP2015045920A
Approach for displaying 3D objects
JP2017501500A
Rendering content in a 3D environment
JP2019520618A
Information processing apparatus, information processing method, video processing system, and program
JP2021086189A
Video processing apparatus, video processing method, and program
JP2022032491A