Image processing device, image processing method and program

The image processing device facilitates easy tracking and switching of virtual camera positions and orientations by displaying multiple virtual viewpoint images in distinct areas, addressing the challenge of maintaining clear camera views during fast-moving subject tracking.

JP2026042071APending Publication Date: 2026-03-10CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies face challenges in accurately grasping the position and orientation of switchable virtual cameras, particularly when tracking fast-moving subjects like basketball players, making it difficult to maintain a clear view of the virtual camera's position and orientation.

Method used

An image processing device that acquires multiple virtual viewpoint images of the same subject and displays them in separate areas, allowing users to easily control and switch between these views using input devices, with the display control unit managing the display of specific virtual viewpoint images based on user operation.

Benefits of technology

Enables easy grasping of the position and orientation of switchable virtual cameras by providing clear visual cues and intuitive user interaction, enhancing the user's ability to track and switch between different virtual camera perspectives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026042071000001_ABST
    Figure 2026042071000001_ABST
Patent Text Reader

Abstract

To easily grasp the position and attitude of a switchable virtual camera. [Solution] The image generating device 10 acquires three or more virtual viewpoint images including the same subject, and controls the display of a specific virtual viewpoint image from the three or more virtual viewpoint images displayed in a first display area in a second display area larger than the first display area based on user operation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to techniques for controlling the display of virtual viewpoint images. [Background technology]

[0002] A technology that uses multiple cameras installed at different positions to capture images synchronously and then generates images (virtual viewpoint images) from any virtual camera (virtual viewpoint) based on user operation using the captured images has attracted attention. This technology makes it possible to view highlight scenes of a soccer or basketball game from various angles, for example, and can give users a more realistic feeling than with ordinary images.

[0003] Patent Document 1 discloses a technology in which viewpoint information such as the position and orientation of a virtual camera is registered, and the registered viewpoint information is read based on user operation, thereby switching to a virtual camera with the registered position and orientation at any timing. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-114869 Summary of the Invention [Problem to be solved by the invention]

[0005] However, there is a risk that it is difficult to grasp the position and orientation of the switchable virtual camera. For example, when a basketball is the subject of imaging and the switchable virtual camera follows the player, the switchable virtual camera moves in accordance with the player's movements, and therefore it is difficult to grasp the position and orientation of the switchable virtual camera.

[0006] The present disclosure makes it easy to grasp the position and orientation of a switchable virtual camera. [Means for solving the problem]

[0007] The image processing device of the present disclosure is characterized by having an acquisition means for acquiring three or more virtual viewpoint images including the same subject, and a display control means for controlling, based on user operation, the display of a specific virtual viewpoint image from among the three or more virtual viewpoint images displayed in a first display area in a second display area larger than the first display area. [Effects of the Invention]

[0008] According to the present disclosure, the position and orientation of a switchable virtual camera can be easily grasped. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a diagram illustrating an example of the configuration of an image generating device 10. FIG. [Figure 2] FIG. 1 is a diagram showing an example of installation of an imaging system 101. [Figure 3] FIG. 10 is a diagram illustrating tracking information. [Figure 4] FIG. 10 is a diagram illustrating a tracking viewpoint. [Figure 5] 5 is a diagram illustrating a display unit 109 and an input device 510. FIG. [Figure 6] FIG. 1 is a diagram illustrating an example of a world coordinate system. [Figure 7] 1 is a diagram illustrating an example of a hardware configuration of an image generating device 10. FIG. [Figure 8] FIG. 2 is a diagram illustrating the functional configuration of a virtual viewpoint generation unit 105. [Figure 9] FIG. 2 is a diagram illustrating the functional configuration of a display control unit 108. [Figure 10] 10A and 10B are diagrams illustrating display contents displayed on a display unit 109. FIG. [Figure 11] 10 is a flowchart showing a process in which the image generating device 10 displays an operation virtual viewpoint image and a switching virtual viewpoint image on the display unit 109. [Figure 12] 10A and 10B are diagrams illustrating the display order of switching virtual viewpoint images by the display control unit 108. FIG. [Figure 13] 10 is a flowchart showing the flow of processing in which the display control unit 108 switches between a selected switching virtual viewpoint image and an operation virtual viewpoint image. [Figure 14] 10A and 10B are diagrams showing a display screen displayed by the display unit 109 in the process of switching between the selected switching virtual viewpoint image and the operation virtual viewpoint image. DETAILED DESCRIPTION OF THE INVENTION

[0010] The present invention will be described in detail below based on preferred embodiments thereof with reference to the accompanying drawings. Note that the configurations shown in the following embodiments are merely examples, and the present disclosure is not limited to the illustrated configurations.

[0011] The image processing system is a system that generates a virtual viewpoint image that represents a scene from a specified virtual viewpoint based on multiple images captured by multiple imaging devices and a specified virtual viewpoint. The virtual viewpoint image in this embodiment is also called a free viewpoint video, but is not limited to an image corresponding to a viewpoint freely (arbitrarily) specified by a user. For example, the virtual viewpoint image also includes an image corresponding to a viewpoint selected by a user from multiple candidates. Furthermore, although the present embodiment will mainly describe a case where the virtual viewpoint image is a video, the virtual viewpoint image may also be a still image.

[0012] The viewpoint information used to generate a virtual viewpoint image is information indicating the position and orientation (line of sight direction) of the virtual viewpoint. Specifically, the viewpoint information is a parameter set including a parameter indicating the three-dimensional position of the virtual viewpoint and a parameter indicating the orientation of the virtual viewpoint in the pan, tilt, and roll directions. Note that the content of the viewpoint information is not limited to the above. For example, the parameter set serving as viewpoint information may include a parameter indicating the size of the field of view (angle of view) of the virtual viewpoint. Furthermore, the viewpoint information may have multiple parameter sets. For example, the viewpoint information may have multiple parameter sets corresponding to multiple frames constituting a moving image of the virtual viewpoint image, and may be information indicating the position and orientation of the virtual viewpoint at each of multiple consecutive time points.

[0013] In this embodiment, the multiple image capturing devices are cameras each having an independent housing and capable of capturing images from a single viewpoint. However, this is not limiting, and two or more image capturing devices may be configured in the same housing. For example, a single camera equipped with multiple lens groups and multiple sensors and capable of capturing images from multiple viewpoints may be installed as the multiple image capturing devices.

[0014] 1 is a configuration diagram of an image generating device 10. The image generating device 10 includes a three-dimensional model generating unit 102, a three-dimensional model tracking unit 103, a storage unit 104, a virtual viewpoint generating unit 105, a virtual viewpoint image generating unit 106, an input unit 107, a display control unit 108, and a display unit 109. Note that the configuration is not limited to the above, and multiple units may be included in different devices. For example, the three-dimensional model tracking unit 103, the virtual viewpoint generating unit 105, the input unit 107, the display control unit 108, and the display unit 109 may be included in different image processing devices.

[0015] 2, the imaging system 101 has multiple physical cameras 201 installed at different positions surrounding an imaging area 202, and captures images in a time-synchronized manner. Multiple images captured synchronously from multiple viewpoints are transmitted to the 3D model generation unit 102. The imaging area 202 may be an imaging studio where imaging is performed to generate a virtual viewpoint image, a stadium where a sports competition is held, or a stage where a concert or a play is held. The imaging system 101 may be a device that captures not only video but also audio and other sensor information.

[0016] The 3D model generation unit 102 extracts the subject as a foreground from the multiple captured images sent from the imaging system 101 and generates a 3D model of the subject from the extracted foreground image. One method for extracting the foreground is to use background difference information. For example, a background image is captured in advance, with no foreground present, and the difference between the image with the foreground and the background image is calculated. If the difference value is greater than a threshold, the pixel position is determined to be the foreground. There are various other methods for extracting the foreground, such as methods using image features related to the subject or machine learning, but the present proposal does not limit the foreground extraction method. The 3D model may be generated using a volume intersection method or depth data obtained from stereo image processing. The present invention does not limit the method for generating the 3D model. The generated 3D model is transmitted to the 3D model tracking unit 103 and the storage unit 104.

[0017] The three-dimensional model tracking unit 103 assigns an identifier to each three-dimensional model generated by the three-dimensional model generation unit 102, and assigns position information of the three-dimensional model as tracking information, and transmits the information to the storage unit 104. The identifier includes unique ID information for identifying each subject and attribute information for identifying the subject's attributes, such as the team to which each subject belongs. As shown in FIG. 3, the tracking information includes position information of the three-dimensional model for each time (time code) and is stored linked to the identifier. The position of the subject is approximated, for example, by the center of gravity of a bounding box surrounding the subject. Note that the position information may be linked to a wireless tag (e.g., GPS) attached to the subject and the three-dimensional model. In this embodiment, the three-dimensional model is tracked using a wireless tag, but this is not limiting. For example, a representative position of the subject may be determined using a portion of the three-dimensional model of the subject, and tracking information may be generated from the range of movement of the representative position of the subject in consecutive frames.

[0018] The storage unit 104 saves and accumulates a group of data (material data) used to generate a virtual viewpoint image and tracking information. The material data specifically includes captured images and three-dimensional models received from the three-dimensional model generation unit 102, and camera parameters of each imaging device. A background model and a background texture image are stored in advance in the storage unit 104 as data used to generate the background of the virtual viewpoint image. The background model may be data of a previously captured studio set or a stadium field, or data of a fictional space generated using CG or the like. In this embodiment, it is assumed that a three-dimensional model of a basketball goal is stored as the background model. The material data group and tracking information stored in the storage unit 104 are transmitted to the virtual viewpoint generation unit 105, the virtual viewpoint image generation unit 106, and the display control unit 108 in accordance with processing by the image generation processing device 10.

[0019] The virtual viewpoint generation unit 105 generates camera parameters of a virtual camera that serves as a viewpoint (tracking viewpoint) for tracking the three-dimensional model, based on the tracking information received from the storage unit 104. The camera parameters of the virtual camera include parameters for specifying the position and orientation, focal length, angle of view, and time. Note that the camera parameters may include parameters that define other elements, or may not include some of the parameters described above. The generated camera parameters are transmitted to the virtual viewpoint image generation unit 106. For example, as shown in FIG. 4, the camera parameters are controlled so that the virtual camera 401 is positioned 3 meters away from the player 402 to be tracked and on a line connecting the player 402 to be tracked and the goal 403. Note that the position of the camera parameters of the tracking viewpoint is not limited to this. For example, the virtual camera 401 may be positioned a predetermined distance away from the line connecting the player 402 to be tracked and the goal 403. It is sufficient that the target to be tracked and the object of interest are included within the angle of view of the virtual camera. The method for determining the camera parameters of the tracking viewpoint is not limited to this. In this embodiment, the image generating device 10 has multiple virtual viewpoint generating units 105, each generating camera parameters for a virtual camera that tracks a different three-dimensional model of a subject. The camera parameters of the virtual cameras generated by the multiple virtual viewpoint generating units 105 are transmitted to different virtual viewpoint image generating units 106. However, this is not limiting, and one virtual viewpoint generating unit 105 may generate camera parameters for multiple virtual cameras, or the camera parameters of multiple virtual cameras may be transmitted to one virtual viewpoint image generating unit 106.

[0020] The virtual viewpoint image generation unit 106 acquires material data from the storage unit 104 and camera parameters of the virtual camera from the virtual viewpoint generation unit 105, and generates a virtual viewpoint image corresponding to the virtual camera. Model-based rendering, for example, can be used as a method for generating the virtual viewpoint image. This process allows for the generation of a virtual viewpoint image viewed from the position and orientation of the virtual camera. The method for generating the virtual viewpoint image is not limited to this. Unless otherwise specified, the term "image" in this disclosure will be described as including the concepts of both moving images and still images. The generated virtual viewpoint image is transmitted to the display control unit 108.

[0021] 5A and 5B are diagrams illustrating the display unit 109 and the input device 510. Fig. 5A shows the layout of the operation virtual viewpoint image and the switching virtual viewpoint image displayed on the display unit 109, and Fig. 5B shows an example of the input device 510 connected to the input unit 107.

[0022] As shown in the example of FIG. 5(b), the input unit 107 acquires input information based on user operation from the input device 510. The input device 510 has a stick 511a, a stick 511b, a seesaw switch 512, and a group of buttons 513. The user operates these to change the camera parameters of the virtual camera. The stick 511a and the stick 511b each have an operation axis with three degrees of freedom. The stick 511a controls the position of the virtual camera, and the stick 511b controls the attitude of the virtual camera. Furthermore, the focal length and angle of view of the virtual camera are changed by tilting the seesaw switch 512 to the plus or minus side. The group of buttons 513 has arrow buttons indicating up, down, left, and right directions, and the user switches between the operation virtual viewpoint and the switching virtual viewpoint. Note that the configuration of the input device 510 is not limited to this and may be, for example, a tablet terminal, a keyboard, a mouse, or the like. The input information includes information for changing the position and attitude of the virtual camera, the focal length, and the angle of view, and information for switching between the operation virtual viewpoint and the switching virtual viewpoint.

[0023] The display control unit 108 sets virtual viewpoint images to be displayed in an operation virtual viewpoint display area 501 (first display area) and a switching virtual viewpoint display area 502 (second display area). As shown in FIG. 5(a), the operation virtual viewpoint display area 501 and the switching virtual viewpoint display area 502 are arranged on the display screen of the display unit 109. The operation virtual viewpoint display area 501 displays a virtual viewpoint image of an operation virtual viewpoint that accepts a user operation to change the position and attitude of a virtual camera. The switching virtual viewpoint display area 502 displays a virtual viewpoint image of a virtual viewpoint (switching virtual viewpoint) that can be switched as the operation virtual viewpoint. The order of the virtual viewpoint images to be displayed as the switching virtual viewpoint is determined based on position information of the three-dimensional model of the tracking target at each viewpoint. A plurality of switching virtual viewpoint display areas 502 are provided, and three are provided in the example of FIG. 5(a). Note that a plurality of switching virtual viewpoint images may be displayed side by side in one display area. When the user selects one of the switching virtual viewpoints via the input unit 107, the selected switching virtual viewpoint is set as the operation virtual viewpoint and displayed in the operation virtual viewpoint display area 501. The display control unit 108 also sets an image to be displayed in the third display area 503. In this embodiment, an overhead image is displayed in the third display area 503, but this is not limiting. For example, a virtual advertisement, the score of a match, or the like may be displayed.

[0024] The display unit 109 displays the operation virtual viewpoint image and the switching virtual viewpoint image set by the display control unit 108 .

[0025] FIG. 6 is a diagram showing a world coordinate system (x, y, z) in the imaging space of the virtual camera. The world coordinate system is used to represent the camera parameters of the virtual camera and the position of the subject. Here, the subject is a tangible object that exists in the imaging space, such as a field 601, a ball 602, or a player 402. In the world coordinate system, the center of the field 601 is set as the origin (0, 0, 0). Furthermore, the x-axis is set in the direction of the long sides of the field 601, the y-axis is set in the direction of the short sides of the field 601, and the z-axis is set in the vertical direction relative to the field 601. However, the method for setting the world coordinate system is not limited to this.

[0026] 7 is a diagram showing an example of the hardware configuration of the image generating device 10. The image generating device 10 has a CPU 701, a RAM 702, a ROM 703, an auxiliary storage device 704, and a communication I / F 705.

[0027] The CPU 701 controls the entire image generation device 10 using computer programs and data stored in the RAM 702 and the ROM 703, thereby realizing each function of the image generation device 10 shown in FIG. 1 . The image generation device 10 may also have one or more dedicated hardware components separate from the CPU 701, and at least some of the processing performed by the CPU 701 may be performed by the dedicated hardware components. Examples of the dedicated hardware components include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and a digital signal processor (DSP). The RAM 702 temporarily stores programs and data supplied from the auxiliary storage device 704 and data supplied from an external device via the communication I / F 705. The ROM 703 stores programs that do not require modification. The auxiliary storage device 704 may be, for example, a hard disk drive, and stores various data such as image data and audio data. The communication I / F 705 is used for communication between the image generation device 10 and external devices. For example, if the image generation device 10 is connected to an external device via a wired connection, a communication cable is connected to the communication I / F 705. If the image generation device 10 has a function for wireless communication with an external device, the communication I / F 705 includes an antenna.

[0028] 8 is a block diagram showing an example of the functional configuration of the virtual viewpoint generation unit 105. The virtual viewpoint generation unit 105 has a three-dimensional model designation unit 801, an object of interest setting unit 802, a virtual viewpoint setting unit 803, and a camera parameter calculation unit 804.

[0029] The three-dimensional model designation unit 801 designates a three-dimensional model for generating a tracking viewpoint from among the three-dimensional models tracked by the three-dimensional model tracking unit 103. The three-dimensional model for generating a tracking viewpoint is designated by the user via the input unit 107. Alternatively, the designation may be based on attribute information included in the tracking information. For example, a three-dimensional model having attribute information "Player (Team A)" is designated as the three-dimensional model for generating a tracking viewpoint. The three-dimensional model designation unit 801 is provided in each virtual viewpoint generation unit 105, and each designates a different three-dimensional model.

[0030] The attention object setting unit 802 sets an attention object for determining camera parameters of the tracking viewpoint. The attention object is a three-dimensional model that serves as a reference for the position and orientation of the virtual camera and is set by the user via the input unit 107. The user does not necessarily need to set an attention object; if not set, the three-dimensional model of the subject set for each sport as the default becomes the attention object. For example, in basketball, the goal is set as the default attention object. Note that the attention object is assumed to be a stationary object, but is not limited to this. For example, in races such as horse racing and bicycle racing, the leading player may be set as the attention object. In this case, a tracking viewpoint, described below, can be generated in which the player being focused on faces the leading player and follows the leading player. Each tracking viewpoint is a viewpoint that includes the subject to be tracked and the attention object within the field of view of the virtual camera. The attention object is a three-dimensional model common to all tracking viewpoints and is a three-dimensional model other than the subject to be tracked at each tracking viewpoint. In other words, the virtual viewpoints generated by the plurality of virtual viewpoint generators 105 each include a different subject while including the same subject within the angle of view.

[0031] The virtual viewpoint setting unit 803 defines conditions for determining the position and orientation of the virtual camera at the tracking viewpoint. The conditions for the position and orientation of the virtual camera include the distance between the target subject and the virtual camera, and the relationship between the target subject and the target object and the position and orientation of the virtual camera. For example, the virtual camera position is defined so that the virtual camera is located 5 m away from the target subject position, at a height z=2.0 m, on a line connecting the virtual camera, the target subject, and the target object in this order. The orientation of the virtual camera is defined so that the center of the target subject is always on the optical axis of the virtual camera. Note that the conditions for the position and orientation of the virtual camera at the tracking viewpoint are not limited to this. It is sufficient that the target subject and the target object are within the angle of view of the virtual camera. Alternatively, it is sufficient that the distance between the target subject and the virtual camera is smaller than the distance between the target object and the virtual camera. For example, the virtual camera position may be determined by rotating the line connecting the virtual camera, the target subject, and the target object in this order by a predetermined angle around the z axis around the target subject. The orientation of the virtual camera may be set so that the center of the subject to be tracked and the center of the target object are on the optical axis.

[0032] The camera parameter calculation unit 804 acquires position information of the subject to be tracked and the three-dimensional model of the object of interest, and calculates the camera parameters of the virtual camera at the tracking viewpoint based on the conditions defined by the virtual viewpoint setting unit 803.

[0033] Fig. 10 is a diagram illustrating the display content displayed on the display unit 109. In the example of Fig. 10, the object of interest is the goal 403, and the viewpoints following the player 402 to be followed are displayed in the operation virtual viewpoint display area 501 and the switching virtual viewpoint display area 502. Here, the subscript a indicates the virtual viewpoint image displayed in the operation virtual viewpoint display area 501, and b, c, and d indicate the virtual viewpoint images displayed in the switching virtual viewpoint display area 502. Meanwhile, both the virtual viewpoint image displayed in the operation virtual viewpoint display area 501 and the virtual viewpoint image displayed in the switching virtual viewpoint display area 502 include the same object of interest, the goal 403.

[0034] FIG. 9 is a block diagram illustrating an example of the functional configuration of the display control unit 108.

[0035] The display viewpoint control unit 901 acquires information on which subject is to be tracked and position information of the three-dimensional model of the subject to be tracked from the virtual viewpoint image generation unit 106. Then, based on the relationship between the three-dimensional model of the subject to be tracked in the virtual viewpoint set as the operation virtual viewpoint and the subject to be tracked set in the other virtual viewpoint, it determines the tracking viewpoint to be displayed in the switching virtual viewpoint display area 502. The display viewpoint control unit 901 sets the virtual viewpoint determined to be displayed as the switching virtual viewpoint. A method for determining the virtual viewpoint to be displayed will be described in detail later.

[0036] The display order control unit 902 determines the order in which the virtual viewpoint images of the virtual viewpoint set as the switching virtual viewpoint by the display viewpoint control unit 901 are displayed in the switching virtual viewpoint display area 502 based on the position information of the three-dimensional model of each tracking target. The method for determining the order will be described in detail later. Information on the order of the tracking viewpoints is transmitted to the display switching unit 903.

[0037] The display switching unit 903 arranges the virtual viewpoint images displayed in the switching virtual viewpoint display area 502 in the order determined by the display order control unit 902. Furthermore, upon receiving a user operation to switch the switching virtual viewpoint image to the operation virtual viewpoint image, the display switching unit 903 sets the switching virtual viewpoint selected by the user as the operation virtual viewpoint. Furthermore, the display switching unit 903 sets a screen layout, such as the size and display position of images to be displayed on the display unit 109. Based on the set layout, the screen configurations of the operation virtual viewpoint display area 501 and the switching virtual viewpoint display area 502 configured on the display unit 109 are determined. Note that, in this embodiment, the operation virtual viewpoint display area 501 is set larger than the switching virtual viewpoint display area 502. Furthermore, when other display images, such as an overhead image showing the entire field, are displayed on the display unit, the display switching unit 903 also sets the layout of the other display images. When an overhead image is displayed, one of the virtual viewpoint image generation units 106 generates a virtual viewpoint image of the overhead image.

[0038] When the marker control unit 904 acquires a user operation to select a switching virtual viewpoint image to be displayed in the switching virtual viewpoint display area 502, it highlights the subject corresponding to the selected switching virtual viewpoint image. Specifically, it superimposes a selection marker on the virtual viewpoint image displayed in the operation virtual viewpoint display area 501 at the position of the three-dimensional model of the subject to be tracked in the selected virtual viewpoint image. The selection marker moves in accordance with the movement of the player based on the tracked position information.

[0039] FIG. 12 is a diagram illustrating a process for determining the display viewpoints and the display order to be displayed on the display unit 109. When capturing and tracking two subjects as imaging targets, the number of virtual viewpoints is small. Therefore, even if two switching virtual viewpoint images are displayed in the switching virtual viewpoint display area 502, the user can easily find the desired switching virtual viewpoint image. On the other hand, when capturing a sport involving a large number of players, such as basketball or baseball, the virtual viewpoints are set according to the number of players, which may make it difficult for the user to easily identify the desired virtual viewpoint. Therefore, in this embodiment, an example is described in which three or more virtual viewpoint images are displayed as switching virtual viewpoint images from among six or more generated virtual viewpoint images. Note that the number of switching virtual viewpoint images to be displayed is not limited to five and may be changed depending on the size of the display device, the size of the switching virtual viewpoint image, and the like. First, among the multiple virtual viewpoint images, the switching virtual viewpoint image is determined based on the distance between the three-dimensional model of the subject to be tracked in the operation virtual viewpoint image and the three-dimensional model of the subject to be tracked in the switching virtual viewpoint image. In this case, the display viewpoint control unit 901 calculates the distance between the 3D model of the subject to be followed in the operation virtual viewpoint image and the 3D model of the subject to be followed in the switching virtual viewpoint image. Then, it determines to display a predetermined number of virtual viewpoint images as switching virtual viewpoints in order of closest distance. The method of determining the switching virtual viewpoint image to be displayed is not limited to this. It may also be determined to display a following viewpoint targeting a 3D model of a player on the same team (with the same attributes) as the 3D model of the subject to be followed in the operation virtual viewpoint. In this case, the display viewpoint control unit 901 refers to attribute information included in the identifier of each 3D model tracked from the storage unit 104. Then, it determines to display a following viewpoint that tracks a 3D model having attributes matching those of the 3D model of the subject to be followed in the operation virtual viewpoint. In the example of FIG. 12, in the switching virtual viewpoint display area 502, players on the same team as the 3D model set as the subject to be followed in the operation virtual viewpoint display area 501 are displayed in a shaded manner.The method of displaying subjects having the same attribute information is not limited to this, and for example, the contours of the three-dimensional models of the subjects may be emphasized, or subjects not having the same attribute information may be made semi-transparent. In this embodiment, it is assumed that the virtual viewpoint image to be displayed as the operation virtual viewpoint image is specified by a user operation.

[0040] The order of the switching virtual viewpoint images to be displayed next may be determined based on, for example, the distance between the 3D model of the target object at each of the following viewpoints and the target object. In this case, the display viewpoint control unit 901 calculates the distance between each 3D model and the target object from the acquired position information of each 3D model, and determines the order in ascending order of distance. In the example of FIG. 12, the switching virtual viewpoint display area 502 displays the switching virtual viewpoint images from left to right in ascending order of distance. The order of the switching virtual viewpoint images may also be determined based on information obtained from outside the image generation device 10. For example, the field goal success rate of each player may be acquired, and the switching virtual viewpoint images corresponding to the players may be arranged in descending order of success rate. Alternatively, the position information of each player may be acquired, and the switching virtual viewpoint images may be arranged according to the field goal success rate in the area where the player is standing.

[0041] FIG. 11 is a flowchart showing the process in which the image generating device 10 displays an operation virtual viewpoint image and a switching virtual viewpoint image on the display unit 109. This process is executed for the update interval of the operation virtual viewpoint image. In this embodiment, it is assumed that the virtual viewpoint images displayed as the operation virtual viewpoint image and the switching virtual viewpoint image are images of 60 frames per second. Therefore, this process (S1101 to S1109) is executed for each frame. However, this is not limitative, and this process may be executed for each plurality of frames. Alternatively, the update interval may be different between the operation virtual viewpoint image and the switching virtual viewpoint image. Furthermore, it is assumed that the initial value (virtual viewpoint) of the operation virtual viewpoint image is determined in advance.

[0042] In step S1101, the 3D model generation unit 102 acquires raw data for a virtual viewpoint image. Specifically, a 3D model of the subject is generated based on parameters such as multiple images received from the imaging system 101 and the position and orientation of each imaging device. However, this is not limiting, and the 3D model of the subject may be acquired from another device. In this case, an identifier including the position information and attribute information of the 3D model is also acquired, and step S1102 is skipped.

[0043] In step S1102, the three-dimensional model tracking unit 103 acquires tracking information for each subject. It tracks the position information of each three-dimensional model generated by the three-dimensional model generation unit 102. Then, it assigns an identifier to each three-dimensional model and transmits it to the storage unit 104 together with the position information (tracking information).

[0044] In step S1103, the virtual viewpoint generation unit 105 generates a virtual viewpoint (tracking viewpoint). Specifically, each of the multiple virtual viewpoint generation units 105 generates camera parameters that become a tracking viewpoint for a specified subject based on the position information of the three-dimensional model received from the accumulation unit 104. In this embodiment, camera parameters are generated such that the target to be tracked and the target object are included within the angle of view of the virtual camera. As a result, the virtual viewpoint images of all the tracking viewpoints each include a different subject, but the same target object.

[0045] In step S1104, the virtual viewpoint image generation unit 106 acquires material data from the storage unit 104 and camera parameters from the virtual viewpoint generation unit 105. Then, it renders a three-dimensional model of the subject as seen from the position and attitude of the virtual camera to generate a virtual viewpoint image. Note that one set of the virtual viewpoint generation unit 105 and the virtual viewpoint image 106 is provided for each tracking target, so that the same number of virtual viewpoint images as the number of tracking targets are generated.

[0046] In step S1105, the display viewpoint control unit 901 determines whether or not the user has performed an operation to select one of the switching virtual viewpoint images displayed in the switching virtual viewpoint display area 502 via the input unit 107. If an operation to select one of the switching virtual viewpoint images has been acquired, the process proceeds to step S1106; if not, the process proceeds to step S1107.

[0047] In step S1106, the display switching unit 903 sets the virtual viewpoint of the switching virtual viewpoint image selected in step S1105 as the operation virtual viewpoint.

[0048] In step S1107, the display viewpoint control unit 901 determines the tracking viewpoint to be displayed in the switching virtual viewpoint display area 502 based on the relationship between the three-dimensional model of the subject to be tracked in the operation virtual viewpoint and the three-dimensional model of the subject to be tracked in the other virtual viewpoints. Based on the determination result, a predetermined number of tracking viewpoints are set as switching virtual viewpoints from among all the tracking viewpoints.

[0049] In step S1108, the display order control unit 902 determines the order in which the tracking viewpoints set as the switching virtual viewpoint in step S1107 are to be displayed in the switching virtual viewpoint display area 502 based on the position information of the three-dimensional models of the subjects to be tracked.

[0050] In step S1109, the display switching unit 903 arranges the virtual viewpoint image of the following viewpoint set as the operation virtual viewpoint in the operation virtual viewpoint display area 501. The display switching unit 903 arranges the virtual viewpoint image of the following viewpoint set as the switching virtual viewpoint in the switching virtual viewpoint display area 502.

[0051] By repeating the above process for each frame, the switching virtual viewpoint image and the operation virtual viewpoint image each contain different subjects but the same object of interest, making it easier for the user to grasp the relationship between the positions and attitudes of the operation virtual viewpoint and the switching virtual viewpoint.

[0052] 13 is a flowchart showing the flow of processing for switching the selected switching virtual viewpoint as the operation virtual viewpoint. This processing is also executed for the update interval of the operation virtual viewpoint image, similar to the processing shown in FIG. 11, and is executed for each frame in this embodiment.

[0053] In step S1301, the display switching unit 903 determines whether or not the user has entered an operation to select one of the switching virtual viewpoint images via the input unit 107. If a selection operation has been entered, the process proceeds to step S1302. If not, the process proceeds to step S1307.

[0054] In step S1302, the marker control unit 904 determines whether a selection marker is displayed in the selected switching virtual viewpoint image. If a selection marker is displayed, the process proceeds to step S1303, and if a selection marker is not displayed, the process proceeds to step S1304. The selection marker is a circular icon such as the one indicated by 1403 in FIG. 14.

[0055] In step S1303, the display switching unit 903 sets the virtual viewpoint corresponding to the selected switching virtual viewpoint image as the operation virtual viewpoint, and also cancels the selection marker displayed on the 3D model of the subject to be tracked in the selected switching virtual viewpoint image and the thick frame displayed in the selected switching virtual viewpoint image.

[0056] In step S1304, the marker control unit 904 superimposes a selection marker on the position of the three-dimensional model of the subject to be tracked in the switching virtual viewpoint image selected by the user on the virtual viewpoint image displayed in the operation virtual viewpoint display area 501. Furthermore, a thick frame 1402 is displayed on the selected switching virtual viewpoint image displayed in the switching virtual viewpoint display area 502.

[0057] In step S1305, the marker control unit 904 determines whether a selection marker is displayed on the three-dimensional model of the subject to be tracked in the switching virtual viewpoint image different from the switching virtual viewpoint image selected by the user, which was acquired in step S1301. If a selection marker is displayed, the process proceeds to step S1306, and if a selection marker is not displayed, the process proceeds to step S1307.

[0058] In step S1306, the marker control unit 904 cancels (deletes) the selection marker displayed on the three-dimensional model of the subject to be tracked in the switching virtual viewpoint image different from the switching virtual viewpoint image selected by the user. Furthermore, if a thick frame 1402 is displayed in the switching virtual viewpoint image different from the switching virtual viewpoint image selected by the user, the thick frame is canceled (deleted).

[0059] In step S1307, the display switching unit 903 displays the switching virtual viewpoint image and the operation virtual viewpoint image on the display unit 109.

[0060] The above process allows the user to easily identify the subject to be tracked by the switching virtual viewpoint image. In particular, when there are many subjects to be tracked, the number of switching virtual viewpoint images increases, which may make it difficult for the user to identify the desired switching virtual viewpoint image. In such a case, the above process allows the user to easily identify which switching virtual viewpoint image corresponds to each subject in the operation virtual viewpoint image. While the present embodiment displays a selection marker, this is not limiting. Alternatively, the color of the 3D model of the target subject may be changed or its outline may be emphasized. As a modified example, when the user selects a subject in the operation virtual viewpoint image, the corresponding switching virtual viewpoint image may be highlighted, or both the subject in the operation virtual viewpoint image and the switching virtual viewpoint image may be highlighted. In this modified example, the user can intuitively identify the position of the subject corresponding to the switching virtual viewpoint image, which makes it easier to identify the position and orientation of the virtual camera in the switching virtual viewpoint image.

[0061] FIG. 14 is a diagram showing a display screen displayed in the process of switching between the selected switching virtual viewpoint image and the operation virtual viewpoint image.

[0062] In the example of FIG. 14(a), the user has selected the first switching virtual viewpoint image from the left among the switching virtual viewpoint images displayed in the switching virtual viewpoint display area 502, and it is surrounded by a thick frame indicating that it is in a selected state. When one of the switching virtual viewpoint images is in a selected state, a selection marker 1401 is superimposed on the position of the three-dimensional model of the subject to be tracked in the selected switching virtual viewpoint image in the operation virtual viewpoint display area 501. The selection marker 1401 is displayed at the feet of player 402b (z=0m) based on the tracked position information. Here, the subscript a of player 402 indicates the tracking viewpoint displayed in the operation virtual viewpoint display area 501, and b indicates the player displayed in the selected switching virtual viewpoint display area 502.

[0063] Fig. 14(a) shows the result of a user performing an operation to switch the selected switching virtual viewpoint to the operation virtual viewpoint via the input unit 107. The switching virtual viewpoint that was selected in Fig. 14(a) is switched to the operation virtual viewpoint, which is displayed in the operation virtual viewpoint display area 501. When the operation virtual viewpoint is switched, the display viewpoint control unit 901 sets a new following viewpoint to be displayed in the switching virtual viewpoint display area, and the following viewpoints are sorted in the order set by the display order control unit 902. Furthermore, the selection marker 1401 displayed in the operation virtual viewpoint image and the thick frame displayed in the switching virtual viewpoint image, which were displayed in Fig. 14(a), are canceled.

[0064] In this embodiment, switching virtual viewpoint images that can be switched at any timing are displayed as candidates for the operation virtual viewpoint from which the user performs operation. Among the multiple virtual viewpoint images, multiple switching virtual viewpoint images are determined based on the relationship with the three-dimensional model being tracked in the operation virtual viewpoint image. The display order of the switching virtual viewpoint images is then determined based on position information of the three-dimensional model of the subject to be tracked in each virtual viewpoint image. Because these operation virtual viewpoint images and switching virtual viewpoint images are tracking viewpoints that include the same object of interest, the position and orientation of the switchable virtual camera can be easily grasped.

[0065] In this embodiment, when a switching virtual viewpoint image displayed in the switching virtual viewpoint display area 502 is selected, the corresponding virtual viewpoint is set as the operation virtual viewpoint. In other words, the selected switching virtual viewpoint image is replaced with the operation virtual viewpoint image. However, this is not limiting, and the selected switching virtual viewpoint image may be enlarged and displayed as the operation virtual viewpoint image. In this case, the switching virtual viewpoint image enlarged and displayed as the operation virtual viewpoint image is highlighted. This modification is effective when the number of tracking viewpoints is small.

[0066] In this embodiment, each virtual viewpoint image is an image viewed from a tracking viewpoint that tracks a specific subject. Therefore, in each virtual viewpoint image, the corresponding subject may be displayed in a manner that allows it to be distinguished from other subjects, so that the subject corresponding to each virtual viewpoint image can be easily identified. For example, the subject corresponding to the virtual viewpoint image may be highlighted by changing its color or by displaying a circular icon at the position of the subject. In this case, a different highlighting method from the selection marker in FIG. 14 is used. [Explanation of symbols]

[0067] 102 3D model generation unit 103 3D model tracking unit 105 Storage Unit 107 Virtual viewpoint generation unit 111 Display control unit

Claims

1. an acquisition means for acquiring three or more virtual viewpoint images including the same subject; a display control means for controlling the display of a specific virtual viewpoint image, among the three or more virtual viewpoint images displayed in the first display area, in a second display area larger than the first display area based on a user operation; 1. An image processing device comprising:

2. 2. The image processing apparatus according to claim 1, wherein the three or more virtual viewpoint images correspond to different subjects from the same subject.

3. The image processing device according to claim 2 , wherein the three or more virtual viewpoint images each include a corresponding subject and the same subject.

4. 3. The image processing device according to claim 2, wherein the three or more virtual viewpoint images are displayed so that the corresponding subject can be distinguished from other subjects.

5. The image processing apparatus according to claim 1 , wherein the acquisition unit acquires a user operation for selecting the specific virtual viewpoint image from among the three or more virtual viewpoint images.

6. The image processing device according to claim 1 , wherein the first display area and the second display area are included in different display devices.

7. The image processing device according to claim 1 , wherein the first display area and the second display area are included in the same display device.

8. The image processing device described in claim 1, characterized in that when a first user operation is acquired, the display control means controls to highlight the specific virtual viewpoint image displayed in the first display area or the subject corresponding to the specific virtual viewpoint image included in the virtual viewpoint image displayed in the second display area, and when a second user operation is acquired, the display control means controls to display the specific virtual viewpoint image in the second display area.

9. The image processing device described in claim 8, characterized in that the first user operation is an operation of selecting a specific virtual viewpoint image from the three or more virtual viewpoint images displayed in the first display area or an operation of selecting a subject corresponding to the specific virtual viewpoint image included in the virtual viewpoint images displayed in the second display area.

10. The image processing device according to claim 8, characterized in that the second user operation is an operation of selecting the highlighted specific virtual viewpoint image displayed in the first display area or an operation of selecting a highlighted subject corresponding to the specific virtual viewpoint image included in the virtual viewpoint image displayed in the second display area.

11. 2. The image processing device according to claim 1, wherein the same subject is a stationary object.

12. The image processing device according to claim 11, wherein the same subject is a goal.

13. an acquisition step of acquiring three or more virtual viewpoint images including the same subject; a display control step of controlling the display of a specific virtual viewpoint image, among the three or more virtual viewpoint images displayed in the first display area, in a second display area larger than the first display area based on a user operation; An image processing method comprising:

14. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Display control device and display control method

    JP2019114869A