Image processing apparatus, image processing method, and program
Patent Information
- Application Number
- JP2023004979
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-01-17
- Publication Date
- 2026-01-16
AI Technical Summary
Existing methods for generating virtual viewpoint images in sports broadcasts fail to clearly depict players' facial expressions and line of sight due to their small size, making it difficult for viewers to understand their positions and formations on the field.
An image processing device that acquires and generates virtual viewpoint images using three-dimensional shape data of objects, enlarging and transforming foreground models to ensure clear visibility of players' facial expressions and line of sight, particularly in overhead views.
Enables easy checking of player positions and formations by displaying players in larger sizes, allowing viewers to see their facial expressions and line of sight clearly.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a technique for generating a virtual viewpoint image. [Background technology]
[0002] There is a technology to generate a virtual viewpoint image, which is an image seen from a virtual viewpoint. The virtual viewpoint image allows a viewer to view a scene such as a sports event from various angles, and therefore provides a higher sense of realism than a normal captured image.
[0003] Incidentally, in sports broadcasts, images captured so as to include the entire field may be used so that viewers can confirm the positions and formations of players on the field on which the sport is being played.
[0004] Patent Document 1 discloses a method of showing viewers the positions of players on the field by arranging figures of players on a figure representing the entire field based on an image captured so as to include the entire field. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2010-183302 A Summary of the Invention [Problem to be solved by the invention]
[0006] When players on the field are represented by figures to show the positions of the players on the entire field as in Patent Document 1, viewers cannot see the facial expressions or line of sight of the players.
[0007] It is also possible to use a virtual viewpoint image that includes the entire field to show the positions of the players on the entire field, but in such a virtual viewpoint image, the players are displayed small, making it difficult for the viewer to check the players' facial expressions and line of sight. [Means for solving the problem]
[0008] The image processing device of the present disclosure is characterized by having a first acquisition means for acquiring information of a virtual viewpoint for generating a virtual viewpoint image, which is an image of an object included in an imaging area of an imaging device viewed from a virtual viewpoint, a second acquisition means for acquiring three-dimensional shape data of the object generated based on an image captured by the imaging device, where at least a portion of the object is enlarged more than the three-dimensional shape data generated based on the image captured by the imaging device, and a generation means for generating the virtual viewpoint image based on the three-dimensional shape data acquired by the second acquisition means when the field of view represented by the virtual viewpoint information is a bird's-eye view. Effect of the Invention
[0009] According to the present disclosure, it is possible to generate a virtual overhead viewpoint image that allows a user to easily confirm an object. [Brief description of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram showing an example of a hardware configuration of an image processing apparatus. [Diagram 2] FIG. 4 is a diagram showing an example of an imaging region. [Diagram 3] FIG. 2 is a block diagram illustrating the functional configuration of the image processing apparatus. [Figure 4] FIG. 4 is a diagram showing an example of an operation unit. [Diagram 5] 11A and 11B are diagrams showing an example of a format of a parameter set of virtual viewpoint information and transformation information. [Figure 6] FIG. 13 is a diagram for explaining a method for enlarging a foreground model. [Figure 7] FIG. 13 is a diagram for explaining a method of dividing a foreground model. [Figure 8] 11 is a flowchart illustrating a process of generating a virtual viewpoint image. [Figure 9] 11 is a flowchart illustrating a process of generating a virtual viewpoint image. [Figure 10] 1A and 1B are diagrams for explaining an example of a virtual viewpoint image generated by an image processing device. [Figure 11] 13 is an example of a virtual viewpoint image when rendered using orthographic projection. [Figure 12] 11A and 11B are diagrams for explaining a method of determining whether or not the field of view is an overhead view based on the position of a virtual viewpoint. [Figure 13] 11 is a flowchart illustrating a process of generating a virtual viewpoint image. [Figure 14] 11 is a flowchart illustrating a process of generating a virtual viewpoint image. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] Hereinafter, the technology of the present disclosure will be described in detail based on the embodiments with reference to the accompanying drawings. Note that the configurations shown in the following embodiments are merely examples, and the technology of the present disclosure is not limited to the configurations shown in the drawings.
[0012] <Embodiment 1> [Hardware configuration] FIG. 1 is a diagram showing an example of the hardware configuration of an image processing device 100 that generates a virtual viewpoint image. The virtual viewpoint image is an image that represents a view from a virtual viewpoint different from the viewpoint of an actual imaging device. The virtual viewpoint image is generated using a plurality of captured images obtained by capturing an imaging area in which an object is present from a plurality of viewpoints in a time-synchronized manner by installing a plurality of imaging devices 110 (see FIG. 3) at different positions. Note that the virtual viewpoint image may be a moving image or a still image. In the following embodiment, the virtual viewpoint image will be described as a moving image.
[0013] As shown in FIG. 1, the image processing device 100 includes a CPU 111, a ROM 112, a RAM 113, an auxiliary storage device 114, a display unit 115, an operation unit 116, a communication I / F 117, and a bus 118.
[0014] The CPU 111 realizes each function of the image processing device 100 by controlling the entire image processing device 100 using computer programs or data stored in the ROM 112 and the RAM 113. Note that the image processing device 100 may have one or more dedicated hardware pieces different from the CPU 111, and at least a part of the processing by the CPU 111 may be executed by the dedicated hardware pieces. Examples of the dedicated hardware pieces include an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), and a DSP (Digital Signal Processor).
[0015] ROM 112 stores programs that do not require modification, etc. RAM 113 temporarily stores programs and data supplied from auxiliary storage device 114, and data supplied from the outside via communication I / F 117, etc. Auxiliary storage device 114 is composed of, for example, a hard disk drive, etc., and stores various data such as image data, audio data, etc.
[0016] The display unit 115 is composed of, for example, a liquid crystal display, an LED, etc., and displays a GUI (Graphical User Interface) etc. for the user to operate the image processing device 100. The operation unit 116 is composed of, for example, a keyboard, a mouse, a joystick, a touch panel, etc., and inputs various instructions to the CPU 111 in response to operations by the user. The CPU 201 also operates as a display control unit that controls the display unit 115 and an operation control unit that controls the operation unit 116.
[0017] The communication I / F 117 is used for communication with an external device of the image processing device 100. For example, when the image processing device 100 is connected to the external device by wire, a communication cable is connected to the communication I / F 117. When the image processing device 100 has a function of wireless communication with the external device, the communication I / F 117 includes an antenna. The bus 118 connects each part of the image processing device 100 to transmit information.
[0018] In this embodiment, the display unit 115 and the operation unit 116 are assumed to exist inside the image processing device 100, but at least one of the display unit 115 and the operation unit 116 may exist outside the image processing device 100 as a separate device.
[0019] [About the imaging device] The image processing device 100 is connected to a plurality of imaging devices 110 (see FIG. 3). The plurality of imaging devices 110 are arranged to capture images of an imaging area from a plurality of positions. The imaging area may be, for example, a field of a stadium where sports such as soccer and baseball are played, or a stage of a venue where concerts and entertainment are held. The plurality of imaging devices 110 are installed at different positions so as to surround such an imaging area, and capture images in a time-synchronized manner. Note that the plurality of imaging devices 110 do not need to be installed all around the imaging area, and may be installed only in a partial direction of the imaging area depending on restrictions on the installation location, etc. The plurality of imaging devices 110 may also include other imaging devices with different functions, such as a telephoto camera and a wide-angle camera.
[0020] Fig. 2 is a diagram showing an example of an imaging area. As shown in Fig. 2, the imaging area is a field 200 on which a sport is played. A coordinate system is set in the imaging area. The set coordinate system is used to specify three-dimensional positions of the multiple imaging devices 110, the virtual viewpoint, the foreground model, and the like. For example, the center position 201 on the field 200 is set as the origin, and an x-axis is set in the left-right direction and a y-axis is set in the up-down direction. In addition, although not shown, a z-axis is set in the height direction.
[0021] [Function configuration] 3 is a diagram for explaining the functional configuration of the image processing device 100. The image processing device 100 has an image processing unit 310 and a control unit 320. The image processing unit 310 has a captured image acquisition unit 311, a foreground / background separation unit 312, a background model holding unit 313, a foreground model generation unit 314, a foreground model holding unit 315, a foreground model deformation unit 316, and a virtual viewpoint image generation unit 317.
[0022] The captured image acquisition section 311 acquires a plurality of captured images obtained at respective times by the plurality of imaging devices 110 capturing images in a time-synchronized manner.
[0023] The foreground / background separation unit 312 separates each captured image obtained by the multiple imaging devices 110 into a foreground image and a background image. The foreground / background separation unit 312 separates the captured image into a foreground image and a background image by using, for example, a background difference method.
[0024] A foreground image is an image obtained by extracting the area of a foreground object (foreground area) from a captured image. A foreground object extracted as a foreground area refers to a dynamic object (moving body) that moves (whose position or shape may change) when images are captured from the same direction in chronological order. Foreground objects include, for example, players, referees, and other people on the field where a sport is being played, or a ball if the sport is a ball game. Or, they can be singers, musicians, performers, or presenters in concerts and entertainment.
[0025] A background image is an image that represents an area (background area) different from the foreground object in a captured image. Specifically, a background image is an image in a state where the foreground object has been removed from the captured image. The background refers to an imaged object that remains stationary or nearly stationary when images are captured from the same direction in chronological order. Such imaged objects include, for example, stadiums where sports are played, venues where concerts are played, and structures and fields such as goals used in ball games. The background may be at least an area different from the foreground object. The imaged objects of the multiple image capture devices 110 may include other objects in addition to the foreground object and the background.
[0026] The background model holding unit 313 holds a background model, which is data representing the three-dimensional shape of an object that serves as the background of a virtual viewpoint image, such as a stadium or a venue. The background model is generated, for example, by performing three-dimensional measurements of the stadium or venue that serves as the background in advance. The format of the background model is, for example, a mesh model. The background model holding unit 313 also holds texture data for coloring the background model. The texture data for coloring the background model is generated based on the background image obtained by the foreground / background separation unit 312 separating the background from the captured image.
[0027] The foreground model generating unit 314 generates data (three-dimensional shape data) representing the three-dimensional shape of a foreground object using a foreground image obtained by the foreground / background separating unit 312 separating the captured image. The three-dimensional shape data of the foreground object is called a foreground model. The foreground model generating unit 314 generates the foreground model using the foreground image based on, for example, a volume intersection method. The foreground model is described as being in the form of a mesh model that represents a three-dimensional shape by connecting polygonal faces such as triangles, but the form of the foreground model is not limited to a mesh model. The foreground model generated by the foreground model generating unit 314 is three-dimensional shape data generated so that the ratio of the size to the background is the same as the actual ratio.
[0028] Foreground model holding unit 315 holds the foreground model generated by foreground model generation unit 314. In addition, foreground model holding unit 315 also holds texture data for coloring the foreground model. The texture data for coloring the foreground model is generated based on the foreground image obtained by foreground / background separation unit 312 separating the foreground / background from the captured image.
[0029] Foreground model deformation unit 316 acquires the foreground model to be processed from foreground model storage unit 315, and acquires a deformed foreground model by deforming the foreground model according to instructions from control unit 320. The deformed foreground model is output to virtual viewpoint image generation unit 317. The deformation of the foreground model includes enlargement and division. The details of the deformation process of the foreground model will be described later. Foreground model deformation unit 316 may acquire a foreground model that has been deformed by an external device or the like.
[0030] The virtual viewpoint image generating unit 317 maps the corresponding texture data to the foreground model and the background model, respectively, and performs rendering in accordance with the virtual viewpoint information specified by the control unit 320 to generate a virtual viewpoint image. The generated virtual viewpoint image is displayed on the display unit 115, for example.
[0031] The virtual viewpoint information includes at least the position of the virtual viewpoint, the line of sight direction from the virtual viewpoint, and the focal length. If the virtual viewpoint is replaced with a virtual camera (virtual camera), the position of the virtual viewpoint corresponds to the position of the virtual camera, and the line of sight direction from the virtual viewpoint corresponds to the orientation of the virtual camera. Moreover, the virtual viewpoint image corresponds to a captured image obtained by virtually capturing an image with the virtual camera.
[0032] The virtual viewpoint information is information of a parameter set including a parameter representing a three-dimensional position of a virtual viewpoint, a parameter representing a pan, tilt, and roll direction representing a line of sight from the virtual viewpoint, and a parameter representing a focal length. The virtual viewpoint information may have a plurality of parameter sets. For example, the virtual viewpoint information may include a plurality of parameter sets corresponding to a plurality of frames constituting a moving image of a virtual viewpoint image, and may be information indicating the position of the virtual viewpoint and the line of sight from the virtual viewpoint at each of a plurality of consecutive time points.
[0033] As shown in FIG. 1, control unit 320 of image processing device 100 has a virtual viewpoint designation unit 321 , a foreground model transformation designation unit 322 , a field of view determination unit 323 , and a foreground model transformation control unit 324 .
[0034] The virtual viewpoint designation unit 321 generates virtual viewpoint information for generating a virtual viewpoint image. The user designates the position, line of sight, and focal length of the virtual viewpoint using the operation unit 116, and the virtual viewpoint designation unit 321 generates virtual viewpoint information including the position, line of sight, and focal length of the virtual viewpoint designated by the user.
[0035] 4 is a diagram showing an example of the operation unit 116. The joysticks 401 and 402 are configured to allow the user to operate three axes (up / down, left / right, twist). The three-dimensional position (x, y, z) of the virtual viewpoint is specified by operating each axis (up / down, left / right, twist) of the joystick 401. The line of sight (pan, tilt, roll) from the virtual viewpoint is specified by operating each axis (up / down, left / right, twist) of the joystick 402. The seesaw switch 403 is used to specify the focal length of the virtual viewpoint.
[0036] Foreground model deformation designation unit 322 generates deformation information that is information used when foreground model deformation unit 316 deforms the foreground model. The deformation information includes values of the magnification ratio and the division ratio. The user designates the values of the magnification ratio and the division ratio using operation unit 116, and foreground model deformation designation unit 322 generates deformation information that includes the values of the magnification ratio and the division ratio designated by the user. The virtual viewpoint information and the deformation information are output to image processing unit 310 as a parameter set.
[0037] 5 is a diagram showing an example of the format of a parameter set including virtual viewpoint information and transformation information that is generated by the control unit 320 and transmitted to the image processing unit 310. A parameter set is composed of, for example, a combination of nestable keys (items) and values corresponding to the keys.
[0038] In Figure 5, viewpoint is a key corresponding to virtual viewpoint information. viewpoint includes the keys position, rotation, and zoom, and holds values corresponding to each key. position is a key corresponding to the three-dimensional position of the virtual viewpoint. rotation is a key corresponding to the line of sight from the virtual viewpoint. zoom is a key corresponding to the focal length. Figure 5 shows that -40.0, 0.0, 2.0 are stored as the three-dimensional position values of the virtual viewpoint, 180.0, 0.0, 0.0 are stored as the line of sight direction values, and 24 is stored as the focal length value.
[0039] On the other hand, foreground is a key that corresponds to transformation information. foreground contains the keys scale and proportion, and holds values corresponding to each key. scale is a key that corresponds to the magnification ratio. proportion is a key that corresponds to the division ratio.
[0040] The visual field determination unit 323 determines whether the visual field (angle of view) represented by the virtual visual field information generated by the virtual visual field designation unit 321 is a bird's-eye view or a normal view. The visual field represented by the virtual visual field information refers to a range included in the virtual visual field image generated by the virtual visual field information. If the virtual visual field is considered as a virtual camera, the visual field represented by the virtual visual field information can be said to be the imaging range of the virtual camera.
[0041] The bird's-eye view is a view that includes the entire imaging area of the multiple imaging devices 110. When the imaging area is the field 200 on which a sport is played as shown in Fig. 2, the bird's-eye view is a view that includes the entire field 200 and allows the viewer to grasp the positions and formations of the players on the field 200. A view narrower than the bird's-eye view, i.e., a view other than the bird's-eye view, is referred to as a normal view.
[0042] Foreground model deformation control unit 324 controls whether to deform the foreground model according to the field of view represented by the virtual viewpoint information. Specifically, when the field of view represented by the virtual viewpoint information is a bird's-eye view, foreground model deformation control unit 324 controls so that the foreground model can be deformed.
[0043] For example, foreground model deformation control unit 324 sets the deformation of the foreground model to be valid when the field of view is an overhead view, and sets the deformation of the foreground model to be invalid when the field of view is a normal view. When the deformation of the foreground model is set to be valid, deformation information including the values of the magnification ratio and division ratio specified by the user is output to foreground model deformation unit 316. Then, control is performed so that foreground model deformation unit 316 deforms the foreground model with the magnification ratio and division ratio specified by the user.
[0044] On the other hand, when the setting is made to disable the deformation of the foreground model, even if the user specifies a magnification ratio and a division ratio, the values of the magnification ratio and the division ratio are corrected to 1, and deformation information including the corrected values of the magnification ratio and the division ratio is transmitted to foreground model deformation unit 316. As will be described later, a magnification ratio value of 1 and a division ratio value of 1 mean that the foreground model will not be deformed. For this reason, when the setting is made to disable the deformation of the foreground model, foreground model deformation unit 316 is controlled so as not to deform the foreground model.
[0045] In addition, when the virtual viewpoint information includes multiple parameter sets, the virtual viewpoint information may include parameters that invalidate the transition of the view from the bird's-eye view to the normal view when the foreground model is transformed.
[0046] Each functional unit in the image processing device 100 in Fig. 3 is realized by the CPU 111 executing a predetermined program, but is not limited to this. Other hardware such as a GPU (Graphics Processing Unit) or an FPGA (Field Programmable Gate Array) for accelerating calculations may also be used. Each functional unit may be realized by cooperation between software and hardware such as a dedicated IC, or some or all of the functions may be realized only by hardware.
[0047] [About transforming the foreground model (enlarging and dividing)] Fig. 6 is a diagram for explaining a method of enlargement processing, which is one type of foreground model deformation processing performed by foreground model deformation unit 316. In Fig. 6, foreground model 601 is a foreground model before enlargement generated by foreground model generation unit 314. Foreground model 604 is a foreground model after enlargement obtained by foreground model deformation unit 316 performing deformation processing on foreground model 601.
[0048] Foreground model deformation unit 316 performs processing to enlarge foreground model 601. When enlarging the foreground model, foreground model deformation unit 316 obtains a value of the enlargement ratio of the foreground model from foreground model deformation specification unit 322. The value of the enlargement ratio is represented as scale. The enlargement ratio value can have a value of scale ≧ 1. When scale = 1 is obtained as the enlargement ratio value, it means that the foreground model is not enlarged.
[0049] Foreground model deformation unit 316 derives bounding box 602, which is a rectangular parallelepiped that circumscribes foreground model 601 to be enlarged. Center 603 of the bottom surface of derived bounding box 602 becomes the reference point when foreground model 601 is enlarged.
[0050] Foreground model deformation unit 316 fixes the position of the reference point and enlarges foreground model 601 in accordance with the enlargement ratio. Specifically, foreground model deformation unit 316 generates enlarged foreground model 604 by converting the three-dimensional coordinates of each vertex constituting a mesh for expressing foreground model 601 using the following equation (1). v2=(v1-vb)×scale+vb...Equation (1)
[0051] v1 is the coordinate of the vertex that composes the mesh before expansion. v2 is the coordinate of the corresponding vertex after expansion. vb is the coordinate of the reference point.
[0052] The following relationship holds between foreground model 601 before enlargement and foreground model 604 after enlargement. The values of width, depth, and height of bounding box 602 surrounding foreground model 601 before enlargement are L1, W1, and H1, respectively. The values of width, depth, and height of bounding box 605 surrounding foreground model 604 after enlargement are L2, W2, and H2, respectively. In this case, the relationship L2 / L1=W2 / W1=H2 / H1=scale holds. The three-dimensional coordinate values of center 603 of the bottom surface of bounding box 602 before enlargement and center 606 of the bottom surface of bounding box after enlargement are equal.
[0053] FIG. 7 is a diagram for explaining one of the deformation processes of the foreground model performed by the foreground model deformation unit 316, which is a splitting method. In FIG. 7, the foreground model 701 is the foreground model before splitting, and the foreground model 706 is the foreground model after splitting.
[0054] When the foreground model deformation unit 316 splits the foreground model, it obtains the value of the splitting ratio of the foreground model from the foreground model deformation specification unit 322. Let the value of the splitting ratio be represented as proportion. The range of values that the splitting ratio can take is 0 < proportion ≤ 1. When proportion = 1 is obtained as the value of the splitting ratio, it means that the splitting of the foreground model is not performed.
[0055] The foreground model deformation unit 316 derives a bounding box 702 that circumscribes the foreground model 701 to be split, and derives the value zd of the z - coordinate at the position of the splitting plane 705 corresponding to the splitting ratio in the derived bounding box 702.
[0056] As shown in FIG. 7, let the value of the z - coordinate at the position of the upper surface 703 of the bounding box 702 be zt, and the value of the z - coordinate at the position of the bottom surface 704 be zb. The splitting ratio indicates the ratio in the z - direction of the part that remains as a result of the splitting in the foreground model before splitting. Therefore, the splitting ratio is represented by the following formula (2). proportion=(zt - zd) / (zt - zb) Formula (2)
[0057] Therefore, the value zd of the z - coordinate representing the position of the splitting plane 705 in the z - direction is derived by the following formula (3). zd = zt×(1 - proportion)+zb×proportion Formula (3)
[0058] The foreground model deformation unit 316 performs processing so that the foreground model 701 is split by the derived splitting plane 705. Specifically, the foreground model deformation unit 316 converts the three - dimensional coordinates of each vertex of the surface constituting the mesh of the foreground model 701 using the following formula (4). v2 = v1 - vd (However, if the z coordinate of v1 is smaller than zd, it is deleted.) Equation (4)
[0059] v1 is the coordinate of the position of a vertex of the face that constitutes the mesh before division. v2 is the coordinate of the position of the corresponding vertex after division. vd is a coordinate to support calculations, and ((0,0,zd-zb)=vd. The meaning in parentheses in equation (4) means that, of the faces that constitute the mesh of foreground model 701, faces that have a vertex with a z coordinate value smaller than zd are deleted.
[0060] The following relationship holds between foreground model 701 before division and foreground model 706 after division: The position of bottom surface 704 of bounding box 702 surrounding foreground model 701 before division and the position of bottom surface 709 of bounding box 707 surrounding foreground model 706 after division are equal.
[0061] The values of the magnification ratio and the division ratio are, for example, values designated by the user by operating the operation unit 116. A method of designating the magnification ratio and the division ratio using the operation unit 116 in FIG. 4 will be described. Seesaw switch 404a in FIG. 4 is a seesaw switch for designating the magnification ratio of the foreground model. Seesaw switch 404b is a seesaw switch for designating the division ratio of the foreground model. When the upper side of seesaw switch 404a, 404b is pressed, the value of the magnification ratio or the division ratio displayed on the screen increases, and when the lower side is pressed, the value decreases. When the determination button on the screen is pressed when the desired values of the magnification ratio and the division ratio are reached by operating seesaw switch 404a, 404b, the values of the magnification ratio and the division ratio displayed on the screen are acquired by foreground model transformation designation unit 322.
[0062] Note that a button may be provided in place of the seesaw switch 404a on the operation unit 116, and the magnification ratio may be specified using the button. For example, a button may be provided that, when pressed by the user, specifies a predetermined value (e.g., 7.0) as the magnification ratio. Also, a button may be provided in place of the seesaw switch 404b, and the division ratio may be specified using the button. For example, when pressed by the user, a predetermined value (e.g., 0.25) may be specified as the division ratio.
[0063] It can be said that the operation unit 116 has two types of zoom switches. The first zoom switch is a seesaw switch 403, and the focal length of the virtual viewpoint, that is, the overall magnification ratio including the foreground and background, is specified by operating the seesaw switch 403. The second zoom switch is a seesaw switch 404a, and the magnification ratio of the foreground model is specified by operating the seesaw switch 404a. Note that when the seesaw switch 404a is operated, only the foreground model is enlarged, and the background model is not enlarged and remains as it is.
[0064] [How the field of view is determined] An example of a method for determining whether the field of view represented by the virtual viewpoint information is a bird's-eye view by the field of view determination unit 323 will be described below. In order to determine whether the field of view represented by the virtual viewpoint information is a bird's-eye view, positions within the imaging area that will be included when the field of view is a bird's-eye view are specified in advance. When the imaging area is a field 200 as shown in FIG. 2, it is preferable that a plurality of positions are specified that surround the field 200. For example, the positions of the four corners of the field (positions 202 to 205 in FIG. 2) are specified.
[0065] Then, it is determined whether each of the specified positions is included in the field of view represented by the virtual viewpoint information. If all of the specified positions 202 to 205 are included in the field of view represented by the virtual viewpoint information, the field of view determination unit 323 determines the field of view to be a bird's-eye view. If any one of the specified positions 202 to 205 is not included in the field of view represented by the virtual viewpoint information, the field of view determination unit 323 determines the field of view to be a normal field of view.
[0066] A determination as to whether the designated positions 202-205 are included in the field of view represented by the virtual viewpoint information is made, for example, as follows. A predetermined position is designated in advance within the imaging area. For example, the center position of the field (position 201 in FIG. 2) is designated as the predetermined position. Then, the resolution of the virtual viewpoint image at the designated predetermined position 201 is calculated based on the virtual viewpoint information using the following formula (5). The resolution is the imaging range per pixel, and is expressed in units of mm / pixel. Resolution = L × δ / f (5)
[0067] L (mm) is the distance from the virtual viewpoint to the predetermined position 201, and is found from the coordinates of the position of the virtual viewpoint and the predetermined position 201. f (mm) is the focal length indicated by the virtual viewpoint information. δ (mm / pixel) is the pixel size of the sensor when the virtual viewpoint is considered as a virtual camera. δ is specified in advance. A larger resolution value means a wider field of view. Therefore, if the obtained resolution is larger than a predetermined value, the field of view determination unit 323 can determine it to be a bird's-eye view, and if it is smaller than the predetermined value, it can determine it to be a normal field of view.
[0068] The predetermined value to be compared with the resolution to determine whether the field of view is an overhead view is obtained, for example, by the following method. The size of the field 200 is 100m x 70m, and the size of the generated virtual viewpoint image is 1980 pixels x 1080 pixels. In order to include the entire horizontal direction of the field 200 in the virtual viewpoint image, a resolution of 100m / 1980 ≒ 50mm / pixel is required, so 50mm / pixel is set as the predetermined value. Alternatively, the predetermined value may be 70m / 1080 pixels ≒ 65mm / pixel obtained using the vertical length of the field and the vertical pixels so that the entire vertical direction of the field is included in the virtual viewpoint image.
[0069] Alternatively, whether or not a certain area is a bird's-eye view may be determined based on the substantial size of the virtual viewpoint image on the screen. This is because, for example, if the virtual viewpoint is located at a position far enough away that the entire field of a stadium is reflected, the virtual viewpoint image can be considered to correspond to an overhead image (having an overhead view) even if the entire field is not reflected in the virtual viewpoint image. In this case, for example, first, the three-dimensional coordinates (x, y, z) of the specified positions 202 to 205 are converted into two-dimensional coordinates (u, v) corresponding to the virtual viewpoint image using a camera matrix determined from the virtual viewpoint information. Then, if the converted two-dimensional coordinates (u, v) fall within the size range of the virtual viewpoint image, the specified positions 202 to 205 are determined to be included in the view represented by the virtual viewpoint information. For example, when the size of the virtual viewpoint image is 1980 pixels x 1080 pixels, if the converted values are 0≦u<1980 and 0≦v<1080, it can be determined that the specified positions 202-205 are included in the field of view represented by the virtual viewpoint information. Alternatively, if the area of a rectangle is calculated from the four converted two-dimensional coordinates and it is determined to be smaller than 1980×1080 (=2,138,400), it can be determined that the specified positions 202-205 are included in the field of view represented by the virtual viewpoint information.
[0070] [flowchart] 8 and 9 are flowcharts for explaining the process of generating one frame of a virtual viewpoint image by the image processing device 100 of this embodiment. The series of processes shown in the flowcharts of Fig. 8 and Fig. 9 are performed by the CPU of the image processing device 100 expanding the program code stored in the ROM into the RAM and executing it. In addition, some or all of the functions of the steps in Fig. 8 and Fig. 9 may be realized by hardware such as ASIC or electronic circuits. Note that the symbol "S" in the explanation of each process means a step in the flowchart, and the same applies to the subsequent flowcharts.
[0071] FIG. 8 is a flowchart for explaining the generation process of one frame of a virtual viewpoint image to be processed when the deformation of the foreground model is set to be invalid.
[0072] In S801, foreground model deformation control unit 324 sets the deformation of the foreground model to be invalid. When the deformation of the foreground model is set to be invalid, even if the user specifies deformation of the foreground model via operation unit 116, the specification is controlled not to be accepted. In other words, even if the user specifies a magnification rate greater than 1, the user's specification is not reflected in the foreground model.
[0073] In S802, the virtual viewpoint designation unit 321 acquires virtual viewpoint information of the frame to be processed. The virtual viewpoint information of the frame to be processed is information generated based on the position of the virtual viewpoint of the frame to be processed, the line of sight direction from the virtual viewpoint, the focal length, and the like, which are designated by the user via the operation unit 116.
[0074] In S803, the field of view determination unit 323 determines the field of view represented by the virtual viewpoint information of the frame to be processed acquired in S802.
[0075] In S804, the control unit 320 determines whether the field of view determined in S803 is an overhead field of view, and branches the process according to the result of the determination. If the control unit 320 determines that the field of view represented by the virtual viewpoint information of the frame to be processed is an overhead field of view (YES in S804), the process proceeds to S806. If the control unit 320 determines that the field of view represented by the virtual viewpoint information of the frame to be processed is a normal field of view (NO in S804), the process proceeds to S805.
[0076] Since the deformation of the foreground model is set to be invalid in S801, the process of generating a virtual viewpoint image corresponding to the frame to be processed is performed using the foreground model that has not been deformed in S805 and S806. The generation of the virtual viewpoint image in S805 and S806 is performed as follows.
[0077] Foreground model deformation control unit 324 outputs deformation information in which the magnification rate value is 1 and the division ratio value is 1 to foreground model deformation unit 316. In addition, foreground model deformation control unit 324 outputs virtual viewpoint information of the frame to be processed to virtual viewpoint image generation unit 317.
[0078] Foreground model deformation unit 316 acquires a foreground model generated based on a captured image corresponding to a frame to be processed from foreground model holding unit 315. That is, a foreground model that has not been subjected to deformation processing and that corresponds to a frame to be processed is acquired. In addition, foreground model deformation unit 316 acquires deformation information output from foreground model deformation control unit 324. The acquired deformation information includes that the magnification rate value is 1 and the division ratio value is 1. Therefore, it is determined based on the deformation information that the foreground model will not be deformed. Therefore, foreground model deformation unit 316 outputs the acquired foreground model to virtual viewpoint image generation unit 317 without deforming it.
[0079] The virtual viewpoint image generation unit 317 acquires a background model from the background model holding unit 313, and acquires an untransformed foreground model of the processing target frame from the foreground model transformation unit 316. The virtual viewpoint image generation unit 317 then renders the acquired background model and foreground model based on the virtual viewpoint information of the processing target frame to generate a virtual viewpoint image of the processing target frame. When the generation of the virtual viewpoint image of the processing target frame is completed in S805, the process proceeds to S809. When the generation of the virtual viewpoint image of the processing target frame is completed in S806, the process proceeds to S807.
[0080] In S807 and S809, the control unit 320 determines whether an instruction to end the virtual viewpoint image has been received from the user. For example, when the user issues an instruction to end using the operation unit 116, the instruction is received by the control unit 320. If the control unit 320 determines in S807 or S809 that an instruction to end from the user has been received (YES in S807, YES in S809), the processing of this flowchart ends.
[0081] If the control unit 320 determines in S809 that the user has not given an end instruction (NO in S809), the process returns to S802 to generate a virtual viewpoint image of the next frame. That is, since the transition to S809 means that the normal field of view has been determined in S804, a process of generating a virtual viewpoint image of the next frame is performed with the foreground model deformation set to be invalid. When a virtual viewpoint image of a moving image is generated in this manner, the process of FIG. 8 or FIG. 9 is repeated at, for example, 60 frames / second to generate a virtual viewpoint image of a moving image made up of a plurality of frames.
[0082] On the other hand, if the control unit 320 determines in S807 that it has not received an end instruction from the user (NO in S807), the process proceeds to S808. That is, the transition to S807 means that the bird's-eye view has been determined in S804, and in this embodiment, in the case of the bird's-eye view, the process proceeds to S808 in order to generate the virtual viewpoint image of the next frame with the setting that enables the deformation of the foreground model.
[0083] 9 is a flowchart for explaining a process for generating one frame of a virtual viewpoint image to be processed when the deformation of the foreground model is set to be valid. The process of S808 will be described in detail with reference to FIG.
[0084] In S901, foreground model transformation control unit 324 performs a setting to enable the transformation of the foreground model.
[0085] S902 is a step similar to S802, in which the virtual viewpoint designation unit 321 acquires virtual viewpoint information of the frame to be processed.
[0086] In S903, foreground model transformation designation unit 322 obtains transformation information for the frame to be processed, which includes the magnification ratio value and the division ratio value designated by the user through operation of operation unit 116.
[0087] S904 is a step similar to S803, and the field of view determination unit 323 determines the field of view represented by the virtual viewpoint information of the frame to be processed acquired in S902.
[0088] In S905, the control unit 320 determines whether the field of view determined in S904 is an overhead field of view, and branches the process according to the result of the determination. When the control unit 320 determines that the field of view represented by the virtual viewpoint information of the frame to be processed is an overhead field of view (YES in S905), the process proceeds to S906.
[0089] In S906, a process of deforming a foreground model corresponding to a frame to be processed is performed. First, in S906, foreground model deformation control unit 324 outputs the deformation information acquired in S903 to foreground model deformation unit 316. Foreground model deformation unit 316 acquires a foreground model of a frame to be processed from foreground model holding unit 315, and deforms the acquired foreground model so as to enlarge or enlarge and divide the foreground model according to the magnification ratio value and the division ratio value included in the deformation information. The deformed foreground model is output to virtual viewpoint image generation unit 317. Then, the process proceeds to S907.
[0090] In S907, a process is performed to generate a virtual viewpoint image corresponding to the frame to be processed using the transformed foreground model. First, in S907, foreground model transformation control unit 324 outputs virtual viewpoint information of the frame to be processed to virtual viewpoint image generation unit 317. Virtual viewpoint image generation unit 317 acquires a background model from background model holding unit 313, and acquires a transformed foreground model from foreground model transformation unit 316. Then, virtual viewpoint image generation unit 317 renders the background model and the transformed foreground model based on the virtual viewpoint information of the frame to be processed, thereby generating a virtual viewpoint image of the frame to be processed. When the virtual viewpoint image of the frame to be processed is generated in S907, the process proceeds to S908.
[0091] On the other hand, if the control unit 320 determines in S905 that the visual field represented by the virtual viewpoint information of the frame to be processed is a normal visual field (NO in S905), the control unit 320 advances the process to S909.
[0092] In S909, control unit 320 branches the process depending on whether or not the user has specified a transformation of the foreground model. When at least one of the magnification rate value and the division ratio value included in the transformation information acquired in S903 is other than 1, control unit 320 determines that the user has specified a transformation. When control unit 320 determines in S909 that the user has specified a transformation of the foreground model (YES in S909), control unit 320 advances the process to S910.
[0093] In S910, the foreground model deformation control unit 324 performs a process of invalidating the virtual viewpoint information acquired in S902. Specifically, the virtual viewpoint information of the processing frame is replaced with the value of the virtual viewpoint information included in the parameter set of the previous frame. The field of view represented by the virtual viewpoint information of the previous frame is the field of view determined to be the bird's-eye view. Therefore, by the process of S910, the virtual viewpoint information of the processing target frame is replaced with the virtual viewpoint information representing the bird's-eye view. Then, the process proceeds to S906 described above, and in S906 to S907, a virtual viewpoint image of the bird's-eye view is generated using the foreground model deformed with the magnification ratio and division ratio designated by the user.
[0094] Thus, in this embodiment, when the setting is made such that the deformation of the foreground model is enabled, the transition from the overhead view to the normal view is not performed during the period in which the deformation of the foreground model is specified by the user.
[0095] On the other hand, if it is determined in S909 that the transformation of the foreground model has not been specified by the user (S909: NO), the control unit 320 advances the process to S911.
[0096] S911 is a step similar to S805, and a virtual viewpoint image of a normal field of view is generated using a foreground model that has not been transformed, in the procedure described in S805. When the virtual viewpoint image of the frame to be processed is generated in S911, the process proceeds to S912.
[0097] In S908 and S912, the control unit 320 determines whether an instruction to end the virtual viewpoint image has been received from the user. If the control unit 320 determines in S908 or S912 that an instruction to end the virtual viewpoint image has been received from the user (YES in S908, YES in S912), the process of this flowchart ends.
[0098] If the control unit 320 determines in S908 that the user has not given an end instruction (NO in S908), the process returns to S902 to generate a virtual viewpoint image for the next frame. That is, since the bird's-eye view was determined in S905, the process of generating a virtual viewpoint image for the next frame is performed with the foreground model deformation set to be valid.
[0099] On the other hand, if the control unit 320 determines in S912 that the user has not given an instruction to end the process (NO in S912), the process proceeds to S913 to generate a virtual viewpoint image for the next frame with the foreground model deformation disabled. In S913, the process of the flowchart in FIG. 8 is performed.
[0100] In addition, when a virtual viewpoint image is generated using an enlarged foreground model, the foreground model may overlap. Therefore, the virtual viewpoint image generating unit 317 may detect whether an overlap of the foreground model occurs when rendering the foreground model, and may display a warning on the display unit 115 or the like when the overlap is detected.
[0101] Furthermore, when the virtual viewpoint image generating unit 317 detects an overlap of a foreground model during rendering of the foreground model, the virtual viewpoint image generating unit 317 may adjust the position of the foreground model to eliminate the overlap. For example, when the virtual viewpoint image generating unit 317 detects an overlap of a foreground model, the virtual viewpoint image generating unit 317 may adjust the overlapped foreground model by shifting it by a distance specified in advance (for example, 1 meter in the positive direction of the x-axis) before rendering.
[0102] Furthermore, when virtual viewpoint image generation unit 317 detects an overlap of a foreground model during rendering of the foreground model, it may perform processing to display a list of foreground models for which overlap is detected on display unit 115 so that the user can select one. Then, when the user selects a foreground model from the list via operation unit 116, virtual viewpoint image generation unit 317 may perform rendering so that the selected foreground model is brought to the foreground.
[0103] [Generated virtual viewpoint image] FIG. 10 is a diagram for explaining an example of a virtual viewpoint image generated by the image processing device 100. FIG. 10(a) is a virtual viewpoint image of a normal field of view, which is an example of a virtual viewpoint image generated using a foreground model that has not been modified. FIG. 10(b) is a diagram showing a comparative example of a virtual viewpoint image of a bird's-eye view, which is an example of a virtual viewpoint image seen from a virtual viewpoint representing a bird's-eye view, which has been generated using a foreground model that has not been modified. In this way, in a bird's-eye view virtual viewpoint image, if a modification is not performed to enlarge the foreground model, the foreground object is displayed small, and the image may become one in which the viewer cannot identify the player, check the player's line of sight, or check the player's facial expression.
[0104] Fig. 10(c) is a diagram of a bird's-eye view virtual viewpoint image generated using a foreground model magnified by 7 times. In the virtual viewpoint image of Fig. 10(c), the players, which are foreground objects, are displayed larger than in Fig. 10(b). This allows the viewer to identify the players, check their line of sight, and check their facial expressions.
[0105] FIG. 10(d) is an example of a virtual overhead viewpoint image generated using a foreground model that represents some parts of a player and is enlarged at a magnification rate of 30 times and divided at a division ratio of 16%. If the foreground model is enlarged by increasing the magnification rate so that the direction of the player's gaze and facial expression can be more clearly seen, the foreground models will overlap. In such a case, by dividing the model while enlarging it, the magnification rate can be increased while suppressing the overlap. As a result, an image can be generated that makes it easy for the viewer to confirm the direction of the gaze and facial expression of each player.
[0106] In this way, it is preferable that the division is performed on the enlarged foreground model. For this reason, for example, the division ratio may be controlled so that it can be set to less than 1 only when the user specifies a magnification ratio greater than 1.
[0107] [Variation 1] The virtual viewpoint image shown in Fig. 10 is an example of a virtual viewpoint image generated by rendering a foreground model and a background model using perspective projection. When rendering using perspective projection, objects farther from the virtual viewpoint are rendered smaller than objects closer to the virtual viewpoint, generating a virtual viewpoint image that looks natural. On the other hand, in order for viewers to check the positions and formations of players, a virtual viewpoint image in which players are displayed at equal sizes across the entire field may be desired.
[0108] Therefore, when the determined field of view is a bird's-eye view, the virtual viewpoint image generation unit 317 may generate the virtual viewpoint image by rendering using a method different from the method used when generating a virtual viewpoint image viewed from a normal field of view.
[0109] For example, if a normal field of view is determined, the virtual viewpoint image generation unit 317 generates a virtual viewpoint image by rendering the foreground model and the background model with perspective projection. On the other hand, if a bird's-eye view is determined, the virtual viewpoint image generation unit 317 generates a virtual viewpoint image by rendering the deformed foreground model with orthographic projection and the background model with perspective projection. Alternatively, if a bird's-eye view is determined, the virtual viewpoint image generation unit 317 may generate a virtual viewpoint image by rendering the deformed foreground model and the background model with orthographic projection.
[0110] Fig. 11 is an example of a virtual viewpoint image when the foreground model used in Fig. 10(c) is rendered by orthographic projection. By rendering the foreground model by orthographic projection, players farther from the virtual viewpoint are prevented from appearing smaller than players closer to the virtual viewpoint, even in a bird's-eye view. This makes it possible to generate an image that makes it easy for viewers to confirm the positions and formation of players farther from the virtual viewpoint.
[0111] [Variation 2] In a bird's-eye view, it is desirable that the size of the foreground object (player) is as uniform as possible across the entire screen. However, in perspective projection, the size of the player changes in proportion to the distance from the virtual viewpoint to the player. Therefore, the field of view determination unit 323 may determine the field of view to be bird's-eye when the distance from the virtual viewpoint to a pre-specified position on the field is longer than a predetermined value. By moving away from the field to a certain extent, the variation in the position from the virtual viewpoint to the players on the field decreases, and the variation in the size of the players in the virtual viewpoint image decreases.
[0112] Fig. 12 is a diagram for explaining Modification 2. Fig. 12(a) is a diagram showing the positional relationship between a field, which is an imaging area, and virtual viewpoints 1201 and 1202. It is assumed that the virtual viewpoint 1201 is located 8 meters from the sideline of the field, and the virtual viewpoint 1202 is located 100 meters from the sideline of the field.
[0113] FIG. 12(b) is a diagram of a virtual viewpoint image seen from a virtual viewpoint 1201. FIG. 12(c) is a diagram of a virtual viewpoint image seen from a virtual viewpoint 1202. Both the virtual viewpoint images of FIG. 12(b) and (c) include the four corners of the field. For this reason, in the method described in the main text of the first embodiment, the field of view indicated by the virtual viewpoint 1201 is also determined to be a bird's-eye view. However, as in the image of FIG. 12(b), which is a virtual viewpoint image seen from the virtual viewpoint 1201, the size of the players varies in the virtual viewpoint image seen from the virtual viewpoint 1201. For this reason, in this modified example, the field of view from the virtual viewpoint 1201 is not determined to be a bird's-eye view. On the other hand, in this modified example, the field of view from the virtual viewpoint 1202 corresponding to FIG. 12(c) is determined to be a bird's-eye view.
[0114] When the virtual viewpoint is used as a reference, the size of a distant player and a nearby player differs by 3.8 times in Fig. 12(b) and by 1.3 times in Fig. 12(c), and the variation in the players, which are foreground objects, is greater in Fig. 12(b). For example, if the variation in size of the foreground object (player) is allowed to be up to 1.3 times, the field of view determination unit 323 may determine that the field of view from the virtual viewpoint is a bird's-eye view when the position of the virtual viewpoint is 100 meters away from the sideline of the field.
[0115] [Variation 3] The foreground model holding unit 315 may hold a model of the player's avatar corresponding to the foreground model, and in the case of a bird's-eye view, a virtual viewpoint image may be generated by rendering the player's avatar model instead of the foreground model. For example, a process is performed to generate a bird's-eye view virtual viewpoint image using a player's avatar model in which at least a part of the model is enlarged compared to the model generated from the captured image. In addition, a process is performed to change the line of sight of the avatar model so that it becomes the same as the player's actual line of sight. When changing the line of sight of the avatar model, the line of sight may be determined using the foreground model of the corresponding player.
[0116] For example, by using an avatar model with a large uniform number, it is possible to generate a bird's-eye view virtual viewpoint image that makes it easier for viewers to identify players. Also, by enlarging the avatar's eyes, it is possible to generate a bird's-eye view virtual viewpoint image that makes it easier for viewers to confirm the player's line of sight.
[0117] As described above, according to this embodiment, it is possible to seamlessly switch between a normal view and a bird's-eye view on one screen. Also, it is possible to display objects in a virtual viewpoint image of the bird's-eye view large so that the user can easily confirm them.
[0118] <Embodiment 2> In the first embodiment, it has been described that when a setting is made to enable transformation and the user instructs transformation of the foreground model, the foreground model is transformed. In this embodiment, a configuration is described in which foreground model transformation control unit 324 instructs whether to transform the model, regardless of the user's instruction. This embodiment is described mainly with a focus on the differences from the first embodiment. Portions not specifically mentioned have the same configuration and processing as the first embodiment and the modified example of the first embodiment.
[0119] FIG. 13 is a flowchart illustrating the processing procedure of the image processing device 100 that corresponds to a normal field of view.
[0120] S1301 is a step similar to S802 and S902, and the virtual viewpoint designation unit 321 acquires virtual viewpoint information of the frame to be processed.
[0121] S1302 is a step similar to S803 and S904, and the field of view determination unit 323 determines the field of view represented by the virtual viewpoint information of the processing target frame acquired in S1301.
[0122] Foreground model deformation control unit 324 of this embodiment is configured to output preset deformation information of the foreground model to foreground model deformation unit 316 according to the field of view represented by the virtual viewpoint information. Specifically, when it is determined that the field of view represented by the virtual viewpoint information of the frame to be processed is a bird's-eye view, foreground model deformation control unit 324 outputs deformation information for performing deformation to foreground model deformation unit 316. For example, deformation information in which the magnification rate value is 7.0 and the division ratio value is 1 is output.
[0123] When the field of view represented by the virtual viewpoint information of the frame to be processed is determined to be a normal field of view, foreground model deformation control unit 324 outputs to foreground model deformation unit 316 deformation information in which the magnification rate value is 1 and the division ratio value is 1. In other words, when the field of view represented by the virtual viewpoint information is determined to be a normal field of view, deformation information that does not deform the foreground model is output.
[0124] In S1303, the control unit 320 determines whether the field of view represented by the virtual viewpoint information has become a bird's-eye view based on the field of view determined in S1302, and branches the process according to the result of the determination. If the control unit 320 determines that the field of view represented by the virtual viewpoint information of the frame to be processed remains the normal field of view as in the previous frame (NO in S1303), the process proceeds to S1304.
[0125] S1304 is a step similar to S805, and a virtual viewpoint image of a normal field of view is generated using a foreground model that has not been transformed, in the procedure described in S805. When the virtual viewpoint image to be processed is generated in S1304, the process proceeds to S1305.
[0126] In S1305, the control unit 320 determines whether an instruction to end the virtual viewpoint image has been received from the user. If the control unit 320 determines that an instruction to end the virtual viewpoint image has been received from the user (YES in S1305), the control unit 320 ends the processing of this flowchart. If the control unit 320 determines that an instruction to end the virtual viewpoint image has not been received from the user (NO in S1305), the control unit 320 returns the processing to S1301 to generate a virtual viewpoint image of the next frame. Then, the processing to generate a virtual viewpoint image of the next frame is performed.
[0127] On the other hand, if the control unit 320 determines that the visual field represented by the virtual viewpoint information has become a bird's-eye view (YES in S1303), the process proceeds to S1404 in Fig. 14. The process of S1404 will be described later.
[0128] In this way, in the flowchart of Fig. 13, a virtual viewpoint image of a normal field of view is generated in S1301 to S1305 using a foreground model that has not been deformed. If the field of view is changed to an overhead view, the determination is YES in S1303 and the process proceeds to S1404 in Fig. 14.
[0129] FIG. 14 is a flowchart for explaining the processing procedure of the image processing device 100 corresponding to the bird's-eye view.
[0130] S1401 is a step similar to S1301, in which the virtual viewpoint designation unit 321 acquires virtual viewpoint information of the frame to be processed.
[0131] S1402 is a step similar to S1302, and the field of view determination unit 323 determines the field of view represented by the virtual viewpoint information of the processing target frame acquired in S1401.
[0132] In S1403, the control unit 320 determines whether the field of view represented by the virtual viewpoint information has become a normal field of view based on the field of view determined in S1402, and branches the process according to the result of the determination. If the control unit 320 determines that the field of view represented by the virtual viewpoint information of the frame to be processed remains the same as the previous frame, which is an overhead field of view (NO in S1403), the process proceeds to S1404.
[0133] In S1404, foreground model deformation unit 316 acquires a foreground model corresponding to the frame to be processed from foreground model holding unit 315. Then, foreground model deformation unit 316 deforms the foreground model of the frame to be processed using the enlargement ratio and division ratio included in the deformation information output from foreground model deformation control unit 324. The deformed foreground model is output to virtual viewpoint image generation unit 317.
[0134] S1405 is a step similar to S907, and a virtual viewpoint image of an overhead view is generated using the transformed foreground model in the procedure described in S907. When the virtual viewpoint image of the processing target frame is generated in S1405, the process proceeds to S1406.
[0135] In S1406, the control unit 320 determines whether an instruction to end the virtual viewpoint image has been received from the user. If the control unit 320 determines that an instruction to end the virtual viewpoint image has been received from the user (YES in S1406), the control unit 320 ends the processing of this flowchart. If the control unit 320 determines that an instruction to end the virtual viewpoint image has not been received from the user (NO in S1406), the control unit 320 returns the processing to S1401 to generate a virtual viewpoint image of the next frame. Then, the processing to generate a virtual viewpoint image of the next frame is performed.
[0136] On the other hand, if the control unit 320 determines in S1403 that the field of view represented by the virtual viewpoint information has been changed to a normal field of view (YES in S1403), the process proceeds to S1304 in Fig. 13. In S1304, a virtual viewpoint image of the processing target frame is generated using a foreground model that has not been transformed as described above.
[0137] 14, a virtual viewpoint image of an overhead view is generated in steps S1401 to S1306 using the deformed foreground model. If the view is changed to a normal view, the result of step S1403 is YES, and the process proceeds to step S1304 in FIG.
[0138] According to the present embodiment described above, it is possible to smoothly transition between a normal field of view and a bird's-eye view on a single screen, and it is possible to display the foreground model large in the bird's-eye view image so that the viewer can easily confirm the images in each field of view.
[0139] <Other embodiments> In the above-described embodiment, the image processing device 100 has been described as generating a three-dimensional model of the foreground and generating a virtual viewpoint image, but the functions included in the image processing device 100 may be realized by one or more devices different from the image processing device 100. For example, each process of extracting the foreground from a captured image, generating a three-dimensional model, and generating a virtual viewpoint image may be performed by a different device.
[0140] The object of the technology of the present disclosure can also be achieved as follows: A storage medium on which program code of software for realizing the functions of the above-described embodiments is recorded is supplied to a system or device. A computer (or a CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.
[0141] In this case, the program code itself read out from the storage medium will realize the functions of the above-described embodiment, and the storage medium storing the program code will constitute the present invention.
[0142] Examples of storage media for supplying the program code include flexible disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, magnetic tapes, non-volatile memory cards, and ROMs.
[0143] The present invention also includes cases where the functions of the above-described embodiments are realized by the following processing: Based on instructions of the program code read by the computer, an operating system (OS) running on the computer performs all or part of the actual processing.
[0144] Furthermore, the functions of the above-mentioned embodiments may be realized by the following process. First, the program code read from the storage medium is written into a memory provided in a function expansion board inserted into a computer or a function expansion unit connected to the computer. Next, based on the instructions of the program code, a CPU or the like provided in the function expansion board or function expansion unit performs some or all of the actual processing.
[0145] The present disclosure can also be realized by a process in which a program for implementing one or more functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more functions.
[0146] The disclosure of the above-described embodiment includes the following configurations.
[0147] (Configuration 1) a first acquisition means for acquiring information on a virtual viewpoint for generating a virtual viewpoint image, which is an image of an object included in an imaging area of an imaging device viewed from the virtual viewpoint; Three-dimensional shape data generated based on an image captured by the imaging device, a second acquisition means for acquiring three-dimensional shape data of the object, the three-dimensional shape data being enlarged in size in which at least a portion of the object is larger than the three-dimensional shape data generated based on the image captured by the imaging device; a generating means for generating the virtual viewpoint image based on the three-dimensional shape data acquired by the second acquiring means when the field of view represented by the virtual viewpoint information is a bird's-eye view; 13. An image processing device comprising:
[0148] (Configuration 2) The imaging area is a field where a sport is played, and the object is a player present on the field. 2. The image processing device according to claim 1,
[0149] (Configuration 3) The present invention further includes a determination unit for determining whether the field of view represented by the information of the virtual viewpoint is the field of view of the overhead view. 3. The image processing device according to configuration 1 or 2.
[0150] (Configuration 4) The determining means is When a plurality of positions in the imaging region are included in the visual field represented by the information of the virtual visual field, the visual field represented by the information of the virtual visual field is determined to be the visual field of the bird's-eye view. 4. The image processing device according to configuration 3.
[0151] (Configuration 5) The determining means is When the resolution at a predetermined position in the imaging region derived based on the information of the virtual viewpoint is greater than a predetermined value, the field of view represented by the information of the virtual viewpoint is determined to be the bird's-eye view. 4. The image processing device according to configuration 3.
[0152] (Configuration 6) The determining means is When the distance from a predetermined position in the imaging region to the virtual viewpoint is greater than a predetermined value, the field of view represented by the information of the virtual viewpoint is determined to be the bird's-eye view. 4. The image processing device according to configuration 3.
[0153] (Configuration 7) The generating means includes: When the determining means determines that the visual field is the bird's-eye view, the virtual visual field image is generated based on the enlarged three-dimensional shape data of the object; When the determining means determines a view other than the bird's-eye view, the virtual viewpoint image is generated based on the unenlarged three-dimensional shape data of the object. 7. The image processing device according to any one of configurations 3 to 6,
[0154] (Configuration 8) The method further includes a model generating unit for generating three-dimensional shape data representing a three-dimensional shape of the object without magnification based on the image captured by the imaging device. 8. The image processing device according to any one of configurations 1 to 7,
[0155] (Configuration 9) the enlarged three-dimensional shape data of the object is three-dimensional shape data of the object that is enlarged in its entirety compared to a three-dimensional shape represented by three-dimensional shape data generated based on an image captured by the imaging device, The second acquisition means includes: A transformation process is at least performed on the unenlarged three-dimensional shape data of the object so that the object is enlarged based on a specified enlargement ratio, thereby generating three-dimensional shape data of the object in which the entirety is enlarged. 9. The image processing device according to configuration 8.
[0156] (Configuration 10) the three-dimensional shape data of the enlarged object is three-dimensional shape data representing a portion of the enlarged object, The second acquisition means includes: processing the unenlarged three-dimensional shape data of the object so as to enlarge the object based on a specified enlargement ratio; Furthermore, a transformation process is performed to divide the enlarged object at a specified division ratio, and three-dimensional shape data representing a portion of the enlarged object is generated. 9. The image processing device according to configuration 8.
[0157] (Configuration 11) The method further includes a modification control means for switching between a first setting for enabling the modification process and a second setting for disabling the modification process, The generating means includes: When the first setting is selected, an instruction for the transformation process is received from a user, and the virtual viewpoint image is generated based on three-dimensional shape data of the enlarged object; In the case of the second setting, an instruction for the transformation process by a user is not accepted, and the virtual viewpoint image is generated based on the three-dimensional shape data of the object that has not been enlarged. 11. The image processing device according to configuration 9 or 10.
[0158] (Configuration 12) The deformation control means When the field of view represented by the information on the virtual viewpoint is the bird's-eye view, the setting is switched to the first setting, and when the field of view represented by the information on the virtual viewpoint is a field of view other than the bird's-eye view, the setting is switched to the second setting. 12. The image processing device according to claim 11,
[0159] (Configuration 13) The second acquisition means includes: The transformation process is performed so that the three-dimensional shape is enlarged according to the specified enlargement ratio, with the center of the bottom surface of a bounding box that encloses the three-dimensional shape represented by the unenlarged three-dimensional shape data of the object as a reference point. 13. The image processing device according to any one of configurations 9 to 12.
[0160] (Configuration 14) Further, the device has an operation unit for a user to specify the operation mode, The operation unit includes at least a first zoom switch for specifying a focal length from the virtual viewpoint, and a second zoom switch for specifying the magnification ratio. 14. The image processing device according to any one of configurations 9 to 13.
[0161] (Configuration 15) The generating means includes: The virtual viewpoint image is generated by rendering the enlarged three-dimensional shape data of the object using orthographic projection. 15. The image processing device according to any one of configurations 1 to 14.
[0162] (Configuration 16) The virtual viewpoint image is a moving image. The generating means includes: When generating a frame including the enlarged object, the virtual viewpoint image is generated so that a transition from the bird's-eye view to a view other than the bird's-eye view does not occur. 16. The image processing device according to any one of configurations 1 to 15.
[0163] (Configuration 17) The generating means includes: The virtual viewpoint image is generated by further using the three-dimensional shape data of the non-magnified background. 17. The image processing device according to any one of configurations 1 to 16,
[0164] (Configuration 18) The second acquisition means includes: obtaining three-dimensional shape data of an avatar representing the object with at least a portion of the object being enlarged; The generating means includes: When the field of view represented by the virtual viewpoint information is the bird's-eye view, a virtual viewpoint image including the avatar is generated using three-dimensional shape data of the avatar. 18. The image processing device according to any one of configurations 1 to 17.
[0165] (Configuration 19) a first acquisition step of acquiring information on a virtual viewpoint for generating a virtual viewpoint image, which is an image of an object included in an imaging area of an imaging device viewed from the virtual viewpoint; Three-dimensional shape data generated based on an image captured by the imaging device, a second acquisition step of acquiring three-dimensional shape data of the object, the three-dimensional shape data being enlarged in such a manner that at least a portion of the object is larger than the three-dimensional shape data generated based on the image captured by the imaging device; a generating step of generating the virtual viewpoint image based on the three-dimensional shape data acquired in the second acquiring step when the field of view represented by the virtual viewpoint information is a bird's-eye view; 13. An image processing method comprising:
[0166] (Configuration 20) A program for causing a computer to execute each of the means of the image processing device according to any one of configurations 1 to 18. [Explanation of symbols]
[0167] 100 Image processing device 314 Foreground Model Generation Unit 316 Foreground model transformation part 317 Virtual viewpoint image generation unit
Claims
1. a first acquisition means for acquiring information on a virtual viewpoint for generating a virtual viewpoint image, which is an image of an object included in an imaging area of an imaging device viewed from the virtual viewpoint; a second acquisition means for acquiring three-dimensional shape data of the object in which at least a portion of the three-dimensional shape data generated based on the image captured by the imaging device is enlarged; a generating means for generating the virtual viewpoint image based on the three-dimensional shape data acquired by the second acquiring means when the field of view represented by the virtual viewpoint information is a bird's-eye view; 1. An image processing device comprising:
2. The imaging area is a field where a game is played, and the object is a player present on the field.
2. The image processing device according to claim 1, wherein:
3. The present invention further includes a determination unit for determining whether the field of view represented by the virtual viewpoint information is the bird's-eye view.
2. The image processing device according to claim 1, wherein:
4. The determining means When the field of view represented by the information on the virtual viewpoint includes a plurality of positions in the imaging area, the field of view represented by the information on the virtual viewpoint is determined to be the bird's-eye view.
4. The image processing device according to claim 3.
5. The determining means When the resolution at a predetermined position in the imaging region derived based on the information of the virtual viewpoint is greater than a predetermined value, the field of view represented by the information of the virtual viewpoint is determined to be the bird's-eye view.
4. The image processing device according to claim 3.
6. The determining means When the distance from a predetermined position in the imaging area to the virtual viewpoint is greater than a predetermined value, the field of view represented by the information of the virtual viewpoint is determined to be the bird's-eye view.
4. The image processing device according to claim 3.
7. The generating means If the determining means determines that the field of view is a bird's-eye view, the virtual viewpoint image is generated based on the enlarged three-dimensional shape data of the object; When the determining means determines that the view is other than the bird's-eye view, the virtual viewpoint image is generated based on the unenlarged three-dimensional shape data of the object.
4. The image processing device according to claim 3.
8. The method further comprises a model generation means for generating three-dimensional shape data representing the unenlarged three-dimensional shape of the object based on the image captured by the imaging device.
2. The image processing device according to claim 1, wherein:
9. the enlarged three-dimensional shape data of the object is three-dimensional shape data of the object that is enlarged overall compared to the three-dimensional shape represented by the three-dimensional shape data generated based on the image captured by the imaging device, The second acquisition means At least a transformation process is performed on the three-dimensional shape data of the object that has not been enlarged so that the object is enlarged based on a specified enlargement ratio, and three-dimensional shape data of the object that has been entirely enlarged is generated.
9. The image processing device according to claim 8,
10. the three-dimensional shape data of the enlarged object is three-dimensional shape data representing a portion of the enlarged object, The second acquisition means The three-dimensional shape data of the object that has not been enlarged is processed so that the object is enlarged based on a specified enlargement ratio, and further a transformation process is performed to divide the enlarged object at a specified division ratio, thereby generating three-dimensional shape data representing a portion of the enlarged object.
9. The image processing device according to claim 8,
11. The method further includes a transformation control means for switching between a first setting for enabling the transformation process and a second setting for disabling the transformation process, The generating means If the first setting is selected, an instruction for the transformation process from a user is received, and the virtual viewpoint image is generated based on three-dimensional shape data of the enlarged object; In the case of the second setting, the instruction for the transformation process from the user is not accepted, and the virtual viewpoint image is generated based on the three-dimensional shape data of the object that has not been enlarged.
10. The image processing device according to claim 9,
12. The deformation control means When the field of view represented by the information on the virtual viewpoint is the bird's-eye view, the setting is switched to the first setting, and when the field of view represented by the information on the virtual viewpoint is a field of view other than the bird's-eye view, the setting is switched to the second setting.
12. The image processing device according to claim 11.
13. The second acquisition means The transformation process is performed so that the three-dimensional shape is enlarged according to the specified enlargement ratio, using the center of the bottom of a bounding box that encloses the three-dimensional shape represented by the three-dimensional shape data of the object that has not been enlarged as a reference point.
10. The image processing device according to claim 9,
14. further comprising an operation unit for a user to specify; The operation unit includes at least a first zoom switch for specifying a focal length from the virtual viewpoint and a second zoom switch for specifying the magnification ratio.
10. The image processing device according to claim 9,
15. The generating means The virtual viewpoint image is generated by rendering the enlarged three-dimensional shape data of the object using orthographic projection.
2. The image processing device according to claim 1, wherein:
16. The virtual viewpoint image is a moving image. The generating means When generating a frame including the enlarged object, the virtual viewpoint image is generated so that a transition from the bird's-eye view to a view other than the bird's-eye view does not occur.
2. The image processing device according to claim 1, wherein:
17. The generating means The virtual viewpoint image is generated by further using the three-dimensional shape data of the non-magnified background.
2. The image processing device according to claim 1, wherein:
18. The second acquisition means obtaining three-dimensional shape data of an avatar representing the object with at least a portion of the object enlarged; The generating means When the field of view represented by the virtual viewpoint information is the bird's-eye view, a virtual viewpoint image including the avatar is generated using three-dimensional shape data of the avatar.
2. The image processing device according to claim 1, wherein:
19. a first acquisition step of acquiring information about a virtual viewpoint for generating a virtual viewpoint image, which is an image of an object included in an imaging area of an imaging device viewed from the virtual viewpoint; a second acquisition step of acquiring three-dimensional shape data of the object in which at least a portion of three-dimensional shape data generated based on the image captured by the imaging device is enlarged; a generating step of generating the virtual viewpoint image based on the three-dimensional shape data acquired in the second acquiring step when the field of view represented by the virtual viewpoint information is a bird's-eye view; An image processing method comprising:
20. A program for causing a computer to execute each of the means of the image processing apparatus according to any one of claims 1 to 18.