Image processing apparatus, image processing method, storage medium, and program product

By generating object models and setting base point positions, and combining the texture information of virtual objects, the problem of low resolution at positions far from the base point in existing technologies is solved, achieving high-definition VR image display, allowing users to clearly identify details in the virtual space.

CN121600147APending Publication Date: 2026-03-03CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511157897.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-20
Filing Date
2025-08-19
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

In directional images generated by existing technologies, objects located far from the base point have low resolution and are difficult to visually identify clearly.

Method used

By generating object models, the positions of virtual objects in virtual space are determined, base point positions are set, and high-definition VR images are generated using image generation components, including texture information of virtual objects, rendering the three-dimensional shape and color information of real objects and background objects, and generating VR images corresponding to the base points.

Benefits of technology

The resolution of objects located far from the base point has been improved, making them clearly visible in VR images, allowing users to better identify details in virtual space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600147A_ABST
    Figure CN121600147A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing apparatus, an image processing method, a storage medium, and a program product. An image processing apparatus according to the present disclosure generates a VR image that allows the generation of a directional image in which an image corresponding to an object present at a position away from a base point can be represented with high definition, obtains an object model, and generates a VR image in which the image corresponding to the object present at a position away from the base point; the object model is generated based on a plurality of captured images obtained by capturing images from a plurality of positions and indicates a three-dimensional shape of an object present in the captured area; determining the position of a virtual object in the virtual space; obtaining the texture of the virtual object; setting the position of a base point in the virtual space; and generating a VR image corresponding to the base point based on the position of the base point, the object model, the virtual object, and the texture of the virtual object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to image processing techniques for generating VR (virtual reality) images based on multiple captured images obtained from cameras at multiple locations. Background Technology

[0002] There exists a technique that generates an image (hereinafter referred to as a "virtual viewpoint image") corresponding to the field of view from any virtual viewpoint (hereinafter referred to as a "virtual viewpoint") by using multiple captured images obtained by simultaneously capturing images from multiple camera devices. The multiple camera devices are arranged at different locations. For description purposes, the aforementioned multiple captured images are hereinafter referred to as "multi-viewpoint images." Additionally, there exists a technique that generates or obtains an image corresponding to the field of view in any direction from a predetermined position (hereinafter referred to as a "base point") by cutting out a portion of the image region from a panoramic image or omnidirectional image that corresponds to the surrounding field of view within a maximum 360-degree range from that predetermined position (hereinafter referred to as a "base point"). For description purposes, the image is hereinafter referred to as a "directional image."

[0003] Japanese Patent Application Publication No. 2020-68513 (hereinafter referred to as "Patent Document 1") discloses a technique for generating images, such as panoramic images or omnidirectional images, corresponding to a base point by combining multiple virtual viewpoint images. Specifically, the technique disclosed in Patent Document 1 uses multi-viewpoint images to generate virtual viewpoint images corresponding to multiple corresponding virtual viewpoints set in the same position and having different viewing directions, and combines the generated multiple virtual viewpoint images. In addition, Patent Document 1 discloses a technique that distributes the image generated by combining virtual viewpoint images and corresponding to a base point to a user terminal, and uses the image corresponding to the base point to generate a directional image corresponding to the direction specified by the user in the user terminal. Summary of the Invention

[0004] The inventors have discovered that in directional images generated by the technology disclosed in Patent Document 1, the resolution of the representation of objects located further away from the base point decreases even more, and therefore, in some cases, it is difficult to visually identify objects with high accuracy in the displayed directional image. For example, in a game where the camera is focused on a baseball field and the base point is located near the catcher, the representation of the scoreboard above the center field screen has low resolution in the directional image corresponding to the direction from the base point to the pitcher.

[0005] According to one aspect of this disclosure, techniques are provided for generating VR images that allow for the generation of directional images. These directional images can represent, in high definition, images corresponding to objects existing at locations far from a base point.

[0006] An image processing apparatus according to the present disclosure includes: a model acquisition unit for acquiring an object model generated based on a plurality of captured images obtained from multiple locations, the object model indicating the three-dimensional shape of an object existing in a camera area; a determination unit for determining the position of a virtual object in a virtual space; a texture acquisition unit for acquiring the texture of the virtual object; a setting unit for setting the position of a base point in the virtual space; and an image generation unit for generating a VR image corresponding to the base point based on the position of the base point, the object model, the virtual object, and the texture of the virtual object.

[0007] The features of this disclosure will become apparent from the following description of embodiments with reference to the accompanying drawings. The following description of the embodiments is given by way of example. Attached Figure Description

[0008] Figure 1A and Figure 1B These are examples of the configurations used to describe the image processing system according to Embodiment 1;

[0009] Figure 2 This is a block diagram illustrating an example of the hardware configuration of the image processing apparatus according to Embodiment 1;

[0010] Figure 3 This is a block diagram illustrating an example of the functional configuration of the image processing device according to Embodiment 1;

[0011] Figure 4 This is a diagram illustrating an example of the virtual object generation process performed by the virtual object generation unit according to Embodiment 1;

[0012] Figure 5 This is a diagram illustrating an example of a VR image generated by an image generation unit according to Embodiment 1;

[0013] Figure 6 This is a flowchart illustrating an example of the processing flow of the image processing apparatus according to Embodiment 1; and

[0014] Figure 7 This is a flowchart illustrating an example of the virtual object generation process performed by the virtual object generation unit according to Embodiment 1. Detailed Implementation

[0015] In the following, this disclosure is described in detail with reference to the accompanying drawings and preferred embodiments. The configurations shown in the following embodiments are merely exemplary, and this disclosure is not limited to the configurations illustrated schematically. Incidentally, the same reference numerals are assigned to and describe the same constituent parts.

[0016] It should be noted that in the following embodiments, the virtual viewpoint image is an image generated based on the position and orientation of a virtual camera device (hereinafter referred to as the "virtual camera") defined at least by the position of the virtual viewpoint and the viewing direction at the virtual viewpoint. Virtual viewpoint images are also referred to as free viewpoint images or arbitrary viewpoint images, etc.

[0017] Furthermore, VR images are images that can undergo VR display as described below. VR images include omnidirectional images and panoramic images, etc., with a video image range (effective video image range) wider than the display range that a display unit can display at one time. Additionally, VR images are not limited to still images and include moving images. VR images have a maximum video image range (effective video image range) corresponding to a field of view of 360 degrees in the left-right direction and 360 degrees in the up-down direction. Furthermore, VR images also include images with a field of view wider than that captured by a normal camera device or a video image range wider than the display range that a display unit can display at one time, although the video image range (effective video image range) corresponds to a field of view less than 360 degrees in the left-right direction and less than 360 degrees in the up-down direction. For example, setting the display mode of the display device (a display device capable of displaying VR images) to "VR view" allows VR images to undergo VR display. Displaying a portion of a VR image with a 360-degree field of view and changing the orientation of the display device in the left-right direction (horizontal rotation direction) allows for movement of the display range and enables viewing of seamless omnidirectional video images in the left-right direction.

[0018] Here, VR display (VR view) is a display method (display mode) that allows for changing the display range, within which a video image of a VR image is displayed within the field of view boundary corresponding to the orientation of the display device. VR displays include "monocular VR displays (monocular VR views)," which perform a transformation (applying distortion correction) mapping a VR image to a virtual sphere and displaying one image. Additionally, VR displays include "binocular VR displays (binocular VR views)," which perform a transformation mapping VR images for the left and right eyes to corresponding virtual spheres and display the VR images side-by-side in the left and right regions. "Binocular VR displays" utilize VR images for the left and right eyes, with parallax between them, thus allowing stereoscopic vision.

[0019] For example, when a user wears a display device such as an HMD (Head-Mounted Display), any VR display shows a video image within a field of view corresponding to the user's facial orientation. For instance, suppose the VR image is displayed at a certain point in time within a field of view centered at 0 degrees horizontally (a specific orientation, such as north) and 90 degrees vertically (90 degrees relative to the zenith, i.e., horizontal). If the orientation of the display device is reversed from the front and back (e.g., the display face is changed from south to north), the display range changes to a video image of the same VR image within a field of view centered at 180 degrees horizontally (the opposite orientation, such as south) and 90 degrees vertically. That is, when the user is wearing an HMD, if the user turns from north to south, this also changes the video image displayed on the HMD from a north-facing video image to a south-facing video image.

[0020] [Example 1]

[0021] In this embodiment, an aspect will be described in which virtual objects representing real-world objects are arranged near a base point in a virtual three-dimensional space (hereinafter referred to as "virtual space") using multi-view images, and a VR image corresponding to the base point is generated. Here, the real-world object is a physically existing object. Furthermore, the virtual object is an object model corresponding to a virtual object generated using CG (computer graphics) technology, etc., that is not physically existing.

[0022] <Configuration of Image Processing System>

[0023] Figure 1A and Figure 1B These are figures illustrating examples of the configuration of the image processing system 1 according to Embodiment 1. Specifically, Figure 1A This is a block diagram illustrating an example configuration of an image processing system 1. The image processing system 1 includes an image processing device 100, multiple sensor systems 11-1 to 11-n (n being an integer greater than 2), an input device 12, a display device 13, a distribution server 14, and one or more user terminals 15. Specifically, where it is not necessary to distinguish between sensor systems 11-1 to 11-n for description, the sensor systems 11-1 to 11-n will be referred to as sensor system 11 without distinction.

[0024] Each sensor system 11 includes one or more camera devices. Each camera device includes a digital still camera or a digital video camera, etc. Images captured by the camera devices included in the respective sensor systems 11 (multi-view images) are sent to an image processing device 100, and the image processing device 100 obtains these images (multi-view images). Specifically, the multiple camera devices included in the multiple sensor systems 11 are arranged at different locations and capture images synchronously with each other. It should be noted that the multiple images included in the multi-view images can be the captured images themselves obtained by the respective camera devices, or images obtained by performing image processing (such as processing to extract a predetermined region from the captured images). The multi-view images obtained by synchronously capturing images by the multiple camera devices included in the multiple sensor systems 11, together with the imaging parameters of the respective camera devices included in the multiple sensor systems 11, are sent to the image processing device 100.

[0025] Camera parameters include data such as extrinsic parameters, intrinsic parameters, and the image size of the captured image. Extrinsic parameters are those indicating the position and orientation of the camera device, and include parameters represented by rotation matrices and position vectors. Intrinsic parameters are those indicating information specific to the camera device, and include parameters indicating focal length, the center position of the captured image, and distortion of the optical system such as lenses. Image size is represented by the number of pixels in the horizontal and vertical directions of the captured image. Figure 1B This is a diagram illustrating an example arrangement of multiple sensor systems 11. The multiple sensor systems 11 are arranged around a camera area 120. The following description is given under the assumption that the camera area 120 is a baseball field where a baseball game is played, and that multiple sensor systems 11, such as 100 sensor systems 11, are arranged around the baseball field. Needless to say, the camera area 120 is not limited to a baseball field. For example, the camera area 120 could be an indoor basketball court, etc.

[0026] Image processing device 100 includes a personal computer or server device, etc., and uses the acquired multi-view images to generate VR images such as omnidirectional images or panoramic images. Specifically, image processing device 100 first generates an object model using the acquired multi-view images. The object model indicates the three-dimensional shape of a real-world object present in the camera area 120. It is assumed that the object model generated by image processing device 100 includes information indicating the shape of the object and information indicating the color of the shape. The object model generated by image processing device 100 and corresponding to the real-world object is stored in association with the timecode used when capturing the multi-view images. Figure 1A and Figure 1B Databases not listed in the table.

[0027] A timecode is information that uniquely identifies the time when a camera device captured an image (or a frame in the case of a film). This information is represented, for example, in a format such as "date:hour:minute:second.frame number". Subsequently, the image processing device 100 arranges virtual objects in a virtual space corresponding to the camera area 120. The virtual objects do not exist in the camera area 120. Then, the image processing device 100 together draws (also called "renders") the virtual objects and the generated object models corresponding to the real-world objects to generate a VR image, such as a panoramic image. The generated VR image is output to the distribution server 14 and distributed to the user terminal 15 via the distribution server 14.

[0028] More specifically, the image processing device 100 first acquires multi-view images and corresponding camera parameters. Then, based on the acquired multi-view images and camera parameters, the image processing device 100 generates shape information indicating the three-dimensional shape of an object used as the foreground (hereinafter referred to as the "foreground object") and color information (also referred to as "texture information") related to the three-dimensional shape. It should be noted that shape information is also referred to as "geometric information." Examples of foreground objects include moving bodies, such as natural persons and spheres present in the camera area 120. For example, the Visual Hull technique is used to generate the shape information related to the foreground object. For example, when generating shape information using the Visual Hull technique, the shape information is obtained as a three-dimensional point cloud, which is a set of points represented using three-dimensional coordinates. It should be noted that the method for generating shape information related to the foreground object using captured images is not limited to the Visual Hull technique. Furthermore, the method for representing shape information related to the foreground object is not limited to a three-dimensional point cloud, and the shape information can be represented using polygonal meshes or voxels, etc.

[0029] Color information related to the foreground object is generated in the method described below. For example, the image processing device 100 uses any point in a 3D point cloud as a point of interest, and uses the color value of the projection destination pixel when the point of interest is projected onto a captured image obtained by a camera device, wherein the camera device is capable of capturing images of points in a camera region 120 corresponding to the point of interest. The image processing device 100 sequentially performs this processing using all or some points included in the 3D point cloud as points of interest, thereby generating color information related to the foreground object. It should be noted that the method for generating color information related to the foreground object is not limited to the method described above. The foreground object is rendered by projecting each point in the 3D point cloud onto the image sensing surface of a virtual camera.

[0030] Subsequently, the image processing device 100 generates color information (texture information) related to the background object based on the multi-view image, camera parameters, and shape information related to the object used as the background (hereinafter referred to as the "background object"). In the case where the main subject of the camera is a baseball game, the background object includes objects within the baseball field, such as fences, center field screens, scoreboards, advertising boards, stands with spectator seating, and the field itself, excluding foreground objects. Furthermore, the shape information related to the background object is information indicating the three-dimensional shape of the background object.

[0031] For example, shape information related to a background object is pre-created and stored in a storage device. The image processing device 100 obtains the shape information by reading it from the storage device and generates color information corresponding to the obtained shape information related to the background object. It should be noted that the image processing device 100 can obtain the shape information related to the background object by generating it based on multi-view images and camera parameters. The following description assumes that the shape information related to the background object is data represented using a polygonal mesh and that the color information is data represented using a texture image.

[0032] A texture image is an image to be UV-mapped to the three-dimensional shape of a background object indicated by shape information related to the background object. The texture image is generated in the method described below. For example, when the corresponding vertices of multiple polygons included in a polygonal mesh corresponding to the background object are projected onto a corresponding captured image, the image processing device 100 first cuts out an image region including pixels corresponding to the corresponding vertices as a partial image. Specifically, when the corresponding vertices of multiple polygons are projected onto captured images obtained by a camera device capable of capturing at least a portion of the background object, the image processing device 100 cuts out an image region including pixels corresponding to the corresponding vertices as a partial image. Subsequently, the image processing device 100 generates a texture image by combining the multiple partial images cut from the corresponding captured images. It should be noted that the method for generating color information related to the background object is not limited to the method described above.

[0033] The background object is rendered by virtually photographing the background object model using a virtual camera. The background object model is obtained by UV mapping a texture image onto a polygonal mesh corresponding to the background object, and corresponds to the background object. The image processing device 100 arranges the virtual objects in a virtual space. Details of the method for arranging the virtual objects will be described below. For example, the position of the virtual objects is determined based on user operations on the input device 12 as described below.

[0034] Subsequently, the image processing device 100 renders the image, including the object model corresponding to the real object and the virtual object within the view of a virtual camera virtually positioned at a base point, thereby generating a VR image. Specifically, the VR image generated by the image processing device 100 is obtained by virtually capturing an image of the entire circumference of a base point set in virtual space using a virtual camera. The position of the base point is determined, for example, based on user operations on the input device 12 as described below. Specifically, the VR image can be in any format, as long as it is the format requested by the distribution server 14. Examples of formats include isometric cylindrical projection and cubemap, etc. The VR image generated by the image processing device 100 is output to the distribution server 14.

[0035] Input device 12 receives operations performed by a user (hereinafter referred to as "operator") of image processing device 100 and sends a signal corresponding to the operation to image processing device 100. Specifically, for example, the operator inputs information by operating input device 12 to specify the position of the base point of the VR image and the position of the virtual objects. Input device 12 sends a signal corresponding to the input to image processing device 100. Image processing device 100 determines the position of the base point of the VR image and the position of the virtual objects based on the signals received from input device 12, and sets the base point of the VR image and arranges the virtual objects. Input device 12 includes input units such as a joystick, touch panel, keyboard, or mouse, or two or more of these input units. The operator inputs the position of the base point of the VR image and the position of the virtual objects by operating the input units. The position of the base point and the position of the virtual objects are specified using, for example, three-dimensional coordinates in virtual space.

[0036] As an example, the following description assumes that the input device 12 includes a keyboard as an input unit, and the operator uses the keyboard to input the three-dimensional coordinate values ​​of the base point position and the position of the virtual object. Furthermore, the description will be given under the assumption that the values ​​of the corresponding components (x, y, z) of a pre-set three-dimensional coordinate system in the virtual space are input as three-dimensional coordinate values. It should be noted that the determination of the base point position and the position of the virtual object in the VR image by the image processing device 100 is not limited to the method based on input made by the operator through the input device 12. For example, the image processing device 100 can read from a storage device a file that has been pre-set and stores three-dimensional coordinate values ​​corresponding to the base point position and the position of the virtual object in the VR image, and perform this determination.

[0037] Display device 13 includes a liquid crystal display (LCD) and displays VR images generated by image processing device 100, as well as an operation GUI (graphical user interface), as information required for the operator to perform operations. The operator views the VR images displayed on display device 13 and inputs the positions of base points and virtual objects via input device 12. For example, in the case of a baseball game, the operator inputs the coordinates of virtual positions corresponding to the pitcher's or nearby positions, the catcher's or nearby positions, etc., via input device 12. Distribution server 14 includes a personal computer or server equipment, obtains VR images output from image processing device 100, and distributes the VR images to one or more user terminals 15.

[0038] Each user terminal 15 includes a personal computer, smartphone, tablet terminal, or HMD, and each generates an image corresponding to any viewing direction based on the VR image received from the distribution server 14. Furthermore, the user terminal 15 causes a display device included in or connected to the user terminal 15 to display the generated image. That is, the image displayed by the display device is a portion of the image area in the VR image. The user of the user terminal 15 (hereinafter referred to as "user") specifies any position in the VR image using a direction from a base point, and allows the display device to display the image area in the VR image corresponding to that position. In other words, each user terminal 15 specifies any direction different from each other, and each display device is allowed to display a portion of the image area in the VR image.

[0039] When each user terminal 15 is a portable or wearable device such as a smartphone, tablet, or HMD, the user terminal 15 can use, for example, a gyroscope sensor or a geomagnetic sensor included in the user terminal 15 to detect changes in the orientation of the display surface of the display device. In this case, the user terminal 15 can change the orientation from the base point, i.e., the image area in the VR image displayed by the display device, based on the detected change in orientation. It should be noted that the specification of the orientation from the base point is not limited to the aforementioned gyroscope sensor, etc. The orientation from the base point can be specified by the user operating an input device included in the user terminal 15 or an input device connected to the user terminal.

[0040] <Hardware Configuration of Image Processing Equipment>

[0041] Will use Figure 2 To describe the hardware configuration of the image processing device 100. Figure 2This is a block diagram illustrating an example of the hardware configuration of an image processing apparatus 100 according to Embodiment 1. The image processing apparatus 100 includes a CPU 201, RAM 202, ROM 203, a communication unit 204, an input / output unit 205, and a GPU 206 as hardware configuration. The CPU 201 is a processor that uses RAM 202 as working memory to execute programs stored in ROM 203 and to control the various units included in the image processing apparatus 100 as a whole. The CPU 201 executes various programs to implement the following description. Figure 3 The functions of the various units included as functional configurations in the image processing device 100 are illustrated herein. RAM 202 temporarily stores computer programs read from ROM 203 and data such as results obtained during computation. ROM 203 holds computer programs and data that do not need to be changed. The following description is given under the assumption that ROM 203 holds shape information related to the background object.

[0042] Communication unit 204 is a communication interface such as Ethernet(R) or USB, and is used for data communication with external devices. For example, communication unit 204 sends VR images to distribution server 14 via Ethernet. Input / output unit 205 inputs / outputs data through input and output interfaces. Input / output unit 205 receives signals from input device 12, such as those related to the position of the base point of the VR image and the position of the virtual objects. Additionally, input / output unit 205 outputs signals, such as those related to the VR image and the operation GUI, to display device 13. GPU 206 is a calculator or processor dedicated to image processing. GPU 206 performs image processing to generate virtual viewpoint images or VR images, etc., based on multi-viewpoint images input from multiple sensor systems 11.

[0043] <Functional Configuration of Image Processing Equipment>

[0044] Figure 3This is a block diagram illustrating an example of the functional configuration of the image processing apparatus 100 according to Embodiment 1. The image processing apparatus 100 includes an acquisition unit 300, a model generation unit 301, a virtual object generation unit 302, a base point setting unit 303, a virtual object position determination unit 304, a virtual object arrangement unit 305, an image generation unit 306, and an output unit 307 as its functional configuration. Each unit included in the image processing apparatus 100 is implemented by executing programs stored in ROM 203 using RAM 202 as its working memory via CPU 201. The acquisition unit 300 acquires multi-view images output from multiple sensor systems 11 and imaging parameters of the imaging devices included in each sensor system 11 via communication unit 204. The multi-view images and imaging parameters acquired by the acquisition unit 300 are stored in RAM 202 and used for processing by the model generation unit 301, the virtual object generation unit 302, the image generation unit 306, etc.

[0045] The model generation unit 301 generates an object model corresponding to the foreground object and an object model (background object model) corresponding to the background object. Specifically, the model generation unit 301 uses the multi-view image and camera parameters obtained by the acquisition unit 300 to generate the object model corresponding to the foreground object. The method for generating the object model corresponding to the foreground object has already been described above, so its description will be omitted.

[0046] Additionally, the model generation unit 301 uses shape information related to the background object stored in ROM 203, along with multi-view images and camera parameters obtained by the acquisition unit 300, to generate a background object model. Specifically, the model generation unit 301 first uses the shape information, multi-view images, and camera parameters to generate a texture image that will undergo UV mapping to a polygonal mesh corresponding to the background object. Subsequently, the model generation unit 301 subjects the generated texture image to UV mapping. The background object model is generated by UV mapping to a polygonal mesh corresponding to the background object. The method for generating a texture image that will undergo UV mapping to a polygonal mesh has already been described above, so its description will be omitted. The object model generated by the model generation unit 301 and corresponding to the foreground object, and the background object model generated by the model generation unit 301 and corresponding to the background object, are stored in RAM 202 and used for processing by the virtual object generation unit 302, image generation unit 306, etc.

[0047] The base point setting unit 303 sets the base point of the VR image. Specifically, for example, based on a signal related to the position of the base point of the VR image output from the input device 12 or information related to that position pre-stored in the ROM 203, the base point setting unit 303 determines the position of the base point of the VR image in the virtual space. Information related to the position of the base point of the VR image determined by the base point setting unit 303, such as the three-dimensional coordinates of that position in the virtual space, is stored in the RAM 202 as a setting value related to the position of the base point.

[0048] The virtual object position determination unit 304 determines the position of the virtual object in the virtual space. Specifically, for example, the virtual object position determination unit 304 determines the position of the virtual object based on a signal related to the position of the virtual object output from the input device 12 or information related to that position pre-stored in the ROM 203, etc. The method for determining the position of the virtual object is not limited to this, and the virtual object position determination unit 304 may, for example, determine the position of the virtual object based on the position of the base point of the VR image set by the base point setting unit 303. Specifically, for example, the virtual object position determination unit 304 determines the position near the base point of the VR image as the position of the virtual object.

[0049] Alternatively, for example, the position of the virtual object can be determined based on the position of the base point of the VR image set by the base point setting unit 303 and the position of the object model generated by the model generation unit 301 that corresponds to the foreground object. Specifically, for example, the virtual object position determination unit 304 determines the position of the virtual object based on the positional relationship between the position of the base point of the VR image and the position of the object model corresponding to the foreground object. More specifically, for example, the virtual object position determination unit 304 determines the position where the virtual object does not occlude the object model when it is viewed from the base point of the VR image as the position for arranging the virtual object. Information such as the three-dimensional coordinates of the position in virtual space related to the position of the virtual object determined by the virtual object position determination unit 304 is stored in RAM 202.

[0050] The virtual object generation unit 302 generates virtual objects. The details of the virtual object generation process performed by the virtual object generation unit 302 will be described below. The virtual objects generated by the virtual object generation unit 302 are stored in RAM 202. The virtual object placement unit 305 places the virtual objects at the positions determined by the virtual object position determination unit 304.

[0051] The process of generating virtual objects will be described. For example, in the case of a baseball game where the camera is focused on the field, the virtual object might be a virtual object representing the shape of a panel such as a scoreboard placed above the center field screen. For example, shape information indicating the three-dimensional shape of the virtual object is prepared in advance and stored in a ROM 203 or similar storage device. The virtual object generation unit 302 uses a virtual viewpoint image obtained by virtually photographing a background object model, etc., using a virtual camera as a texture image, and applies the texture image to a pre-prepared polygonal mesh of the panel shape to generate the virtual object. Here, the virtual camera described is a different type of virtual camera used to generate VR images. The virtual camera photographs an object model corresponding to a real-world object of interest, such as a scoreboard, at high resolution. The following will use… Figure 4 This will describe in more detail the methods used to generate virtual objects.

[0052] Virtual objects generated in this way are positioned near the base point of the VR image, which can have the following effect. Specifically, representations of real-world objects located far from the camera area 120 corresponding to the base point may appear large in the VR image. Representations of real-world objects typically appear small, have low resolution, and are difficult to visually recognize in VR images. As a result, representations of real-world objects can be clearly displayed on the display device of the user terminal 15 or displayed in a way that is visually recognizable to the user.

[0053] It should be noted that in this embodiment, the description is given under the following assumptions: the virtual object generation unit 302 uses a virtual viewpoint image as a texture image to fit a polygonal mesh of the panel shape as described above. However, this is not limiting. For example, the virtual object generation unit 302 may use an image obtained by cutting out an image region including a representation of an object of interest, such as a scoreboard, from a texture image of the shape of a real object to be UV-mapped to, such as a background object, as a texture image. Alternatively, for example, the virtual object generation unit 302 may use images obtained by cutting out respective image regions including representations of objects of interest, such as scoreboards, from the captured images included in the multi-viewpoint images obtained by the acquisition unit 300 as texture images.

[0054] Furthermore, in this embodiment, the description will be given under the assumption that the virtual object has a panel shape as described above, but the shape of the virtual object is not limited to this. For example, the virtual object generation unit 302 can extract shape information indicating the three-dimensional shape of a background object of interest, such as a scoreboard located in the distance, from the shape information related to the background object, and use this shape information as the shape information related to the virtual object. This allows the shape of the virtual object to be the same as or substantially the same as the shape of the background object of interest, such as a scoreboard.

[0055] Image generation unit 306 generates VR images. Specifically, image generation unit 306 generates VR images from base points in virtual space that correspond to object models of foreground and background objects and the appearance of virtual objects. The method for generating VR images has already been described above, so its description will be omitted. The VR images generated by image generation unit 306 are stored in RAM 202. Output unit 307 outputs the VR images generated by image generation unit 306 to distribution server 14 via communication unit 204.

[0056] <Dual Object Generation and Processing>

[0057] Reference Figure 4 Describe the process of generating virtual objects. Figure 4 This is a diagram illustrating an example of the virtual object generation process performed by the virtual object generation unit 302 according to Embodiment 1. As an example, Figure 4 This example illustrates an overview of the camera subject viewed from a distance, in the context of a baseball game. Figure 4 In the example, the gray rectangles represent the players competing. Figure 4 In this context, object 403 is a scoreboard positioned above a center field screen, which is a background object and part of the baseball field 402. Object 403 is the real-world object of interest in this embodiment, corresponding to the virtual object. That is, the virtual object has, for example, a panel shape representing the shape of object 403.

[0058] Figure 4The illustrated imaging device 401 indicates the position and orientation of a virtual imaging device, i.e., a virtual camera in a virtual space. The imaging device 401 does not actually exist in the imaging area 120. The position 404 is a position corresponding to the position of the reference point set by the reference point setting unit 303 in the virtual space. The virtual object 405 indicates the shape and position of a virtual object arranged in the virtual space by the virtual object arrangement unit 305, and does not actually exist in the imaging area 120. The imaging device 401, as a virtual camera, generates a virtual viewpoint image corresponding to the representation of the object 403 by virtually imaging an object model corresponding to the object 403 in the virtual space. The virtual object generation unit 302 generates the virtual object 405 by attaching the generated virtual viewpoint image as a texture image to the panel-shaped virtual object 405. The generated virtual object is arranged at a position in the virtual space corresponding to a position near the position 404 where the catcher of the baseball field 402, which is set as the reference point, exists.

[0059] <Example of Generation of VR Image>

[0060] Figure 5 FIG. is an example diagram illustrating a VR image 501 generated by the image generation unit 306 according to Embodiment 1. The VR image 501 includes a representation 502 of the object 403 and a representation 503 of the virtual object 405, and the VR image 501 is displayed on the display device 13. The image processing device 100 displays a reference point position designation area 504 and a virtual object position designation area 505 by superimposing the reference point position designation area 504 and the virtual object position designation area 505 on the VR image 501. The reference point position designation area 504 is used for an operator to input three-dimensional coordinate values of the reference point. The virtual object position designation area 505 is used for the operator to input three-dimensional coordinate values for arranging the virtual object. The operator uses the input device 12 to input the three-dimensional coordinate values into the reference point position designation area 504 and the virtual object position designation area 505. The image processing device 100 receives a signal based on the input, sets the reference point of the VR image, and determines the position for arranging the virtual object.

[0061] A background object model corresponding to the object 403 exists at a position in the virtual space that is away from the position 404 of the reference point. Therefore, the resolution of the representation 502 of the object 403 in the VR image becomes lower. That is, when the user sets the viewing direction from the reference point to the direction where the background object model corresponding to the object 403 is located in the user terminal 15, in the image displayed on the display device of the user terminal 15, the resolution of the representation of the object 403 also becomes lower. Therefore, it is difficult for the user to visually recognize information such as the score of the game included in the image or information related to the players.

[0062] On the other hand, a virtual viewpoint image obtained by virtually photographing the background object model corresponding to object 403 using a virtual camera is used as a texture image and attached to the virtual object 405 near the base point 404. Therefore, in the VR image, the resolution of the representation 503 of the virtual object 405 becomes increasingly larger and higher. That is, the resolution of the representation of the virtual object 405 in the image displayed on the user terminal 15 also becomes higher. This allows the user to visually identify information such as match scores or player-related information included in the image.

[0063] The VR image generated by the image processing device 100 is distributed to multiple user terminals 15 via the distribution server 14. Each user terminal 15 allows a user to specify any viewing direction to select an image area from the VR image corresponding to that direction, and allows the display device of the user terminal 15 to display that image area. This allows each user terminal 15, for example, in the case of a baseball game, to use their own user terminal 15 to receive information beneficial to watching the baseball game presented on object 403 while watching the game.

[0064] Operation of Image Processing Equipment

[0065] Reference Figure 6 and Figure 7 To describe the operation of the image processing device 100. Figure 6 This is a flowchart illustrating an example of the processing flow of the image processing apparatus 100 according to Embodiment 1. Figure 6 The processing in the illustrated flowchart is implemented by the CPU 201 loading the control program stored in ROM 203 into RAM 202 and executing the control program. It should be noted that the processing in the flowchart is repeated whenever multi-view images and camera parameters are output from the multiple sensor systems 11. Furthermore, each processing step (process) is indicated below by appending "S" to the beginning of the reference numerals. First, in S600, the acquisition unit 300 acquires the multi-view images and camera parameters output from the multiple sensor systems 11. The acquired multi-view images and camera parameters are stored in RAM 202 and used for processing by the model generation unit 301, etc.

[0066] Next, in S601, the model generation unit 301 generates object models corresponding to real foreground and background objects based on the multi-view image. The generated object models are stored in RAM 202 and used for processing by the virtual object generation unit 302, image generation unit 306, etc. Next, in S602, the virtual object generation unit 302 generates virtual objects based on the object models generated in S601 that correspond to background objects, etc. The generated virtual objects are stored in RAM 202 and used for processing by the virtual object placement unit 305, image generation unit 306, etc. Next, in S603, the base point setting unit 303 sets the position of the base point of the VR image based on the signal sent from the input device 12. The setting value related to the position of the base point is, for example, a three-dimensional coordinate value in virtual space indicating the position of the base point. The setting value that indicates the position of the base point of the VR image is stored in RAM 202 and used for processing by the image generation unit 306, etc.

[0067] Next, in S604, the virtual object position determination unit 304 determines the position where each of the virtual objects generated in S602 is placed based on the signal sent from the input device 12. The position of each virtual object is determined, for example, as its three-dimensional coordinates in the virtual space where the virtual objects are placed. It should be noted that the direction of placing the virtual objects is determined such that the normal of the surface of the virtual object that is attached to the textured image faces the base point. Information related to the determined position of the virtual objects is stored in RAM 202 and used for processing by the virtual object placement unit 305, etc. Next, in S605, the virtual object placement unit 305 places the virtual objects generated in S602 at the positions in the virtual space determined in S604.

[0068] Next, in S606, the image generation unit 306 generates a VR image based on the position of the base point set in S603, including the object model generated in S601 and the virtual objects generated in S602 and arranged in S605. The generated VR image is stored in RAM 202 and used for processing by the output unit 307, etc. Then, in S607, the output unit 307 outputs the VR image generated in S606 to the distribution server 14. After S607, the image processing device 100... Figure 6 The processing in the illustrated flowchart has ended.

[0069] Furthermore, the above description has been given under the assumption that the processes from S600 to S605 are executed sequentially, but this is not limiting. For example, as long as the process in S600 is executed before the processes in S601 and S602, and the process in S604 is executed before the process in S605, then the processes from S600 to S605 can be executed in any order, and two or more processes from S600 to S605 can be executed in parallel.

[0070] <Dual Object Generation and Processing>

[0071] Reference Figure 7 This describes the process of virtual object generation by the virtual object generation unit 302 in S602. Figure 7 This is a flowchart illustrating an example of the virtual object generation process performed by the virtual object generation unit 302 according to Embodiment 1. Figure 7 This is a flowchart illustrating an example of the processing flow in S602. The following describes an example of a virtual object generation process where a virtual object is generated using a virtual viewpoint image obtained by virtually photographing a background object model corresponding to a background object using a virtual camera. [The following text has been used...] Figure 4 It describes the virtual viewpoint image. Figure 7 The processes in the illustrated flowchart are executed by the virtual object generation unit 302 after the processes in S601.

[0072] Following S601, in S701, the virtual object generation unit 302 determines the position and orientation of a virtual camera capable of capturing images of an object model corresponding to a real-world object of interest, such as object 403, based on signals from the input device 12. In this embodiment, the description will assume, as described above, that the operator uses the input device 12 to specify the position of the object model corresponding to the object of interest in virtual space; however, the method for determining the position and orientation of the virtual camera is not limited to this. For example, the virtual object generation unit 302 can determine the position and orientation of the virtual camera by reading a file or data pre-included with position-related information (such as three-dimensional coordinate values) from the ROM 203.

[0073] Next, in S702, the virtual object generation unit 302 arranges a virtual camera with the position and orientation determined in S701, and generates a virtual viewpoint image by virtually photographing the object model corresponding to the object of interest using the virtual camera. Next, in S703, the virtual object generation unit 302 generates a virtual object by fitting the virtual viewpoint image generated in S702 to a shape such as a panel shape represented by a polygonal mesh of the virtual object. Here, the virtual object generation unit 302 can fit an image obtained by masking a predetermined image region in the virtual viewpoint image to the shape of the virtual object. The predetermined image region is, for example, an image region including information unrelated to the intent of the camera subject, such as a baseball game. The predetermined image region is an image region including information that is not suitable for placement near a base point.

[0074] Furthermore, the virtual object generation unit 302 can attach an image obtained by cutting out only a predetermined image region from the virtual viewpoint image to the shape of a virtual object. Additionally, the virtual object generation unit 302 can attach an image obtained by combining two or more images obtained by cutting out two or more image regions from the virtual viewpoint image to the shape of a virtual object. Furthermore, the virtual object generation unit 302 can attach an image obtained by changing the image size or tone of the virtual viewpoint image to the shape of a virtual object. Additionally, the virtual object generation unit 302 can attach an image obtained by changing the transparency of the virtual viewpoint image to the shape of a virtual object. Even when virtual objects are arranged, the image is attached to the shape of the virtual object, thereby making the representation of objects behind which are occluded by virtual objects in the view from the base point visually identifiable. After S703, the virtual object generation unit 302 makes... Figure 7 The process in the illustrated flowchart (i.e., the process in S602) ends.

[0075] As described above, the image processing device 100 is configured to generate an object model corresponding to a real-world object using multi-view images, and to generate a virtual object based on the generated object model or multi-view images. Furthermore, the image processing device 100 is configured to arrange the generated virtual object near a base point, and to generate a VR image corresponding to the view from the base point while the virtual object is arranged. This configuration of the image processing device 100 allows the generation of VR images that can represent the object with high definition. Because the object exists at a location far from the base point, it is difficult to visually recognize the object's representation by cutting out from the VR image. As a result, the user can clearly view the object's representation displayed on the user terminal 15.

[0076] [Modification of Example 1]

[0077] In Embodiment 1, as an example, aspects of the image processing device 100 generating object models have been described; however, the image processing device 100 can be configured to obtain objects generated by... Figure 1A and Figure 1B The object model is generated by an external device not illustrated in the example. Furthermore, in Embodiment 1, as an example, the aspect of the image processing device 100 generating images such as virtual viewpoint images to be applied as texture images to virtual objects has been described, but this is not limiting. For example, the image processing device 100 can be configured to obtain images generated by external devices not illustrated in the example. Figure 1A and Figure 1B Texture images generated by external devices not shown in the example are then applied to virtual objects.

[0078] Furthermore, in Embodiment 1, as an example, the aspect of the image processing device 100 applying virtual viewpoint images, etc., as texture images to virtual objects has been described, but this is not limiting. For example, the image processing device 100 can be configured to obtain... Figure 1A and Figure 1B An external device, not illustrated, applies a texture image to a virtual object and places the resulting virtual object at a predetermined location.

[0079] Furthermore, in Embodiment 1, as an example, the aspect of attaching a single texture image to a virtual object has been described; however, multiple texture images can be attached to a single virtual object. Specifically, for example, multiple images obtained by cutting out multiple image regions from a virtual viewpoint image, multiple images cut out from a captured image, or combinations thereof, can be attached to the shape of the virtual object while the attachment position changes. Additionally, for example, multiple images or combinations thereof cut out from multiple virtual viewpoint images obtained from multiple virtual cameras or from multiple captured images included in a multi-viewpoint image can be attached to the shape of the virtual object while the attachment position changes. With such a configuration, for example, in the presence of multiple objects of interest, the user terminal 15 can generate VR images that can simultaneously display representations of multiple objects.

[0080] Furthermore, in Embodiment 1, as an example, the arrangement of a single virtual object in virtual space has been described; however, multiple virtual objects can be arranged in virtual space. In this case, texture images that are different from each other or the same texture image can be overlaid onto the corresponding virtual object. Overlaying different texture images onto the corresponding virtual object allows, for example, the arrangement of the corresponding virtual object based on the positional relationship between the multiple objects when multiple objects of interest exist. Additionally, overlaying the same texture image onto the corresponding virtual object allows the generation of a VR image in which, for example, even if the user changes the viewing direction of the image displayed on the user terminal 15, the representation of the object of interest will always be displayed.

[0081] Furthermore, in Embodiment 1, as an example, the aspect of generating a single VR image for a base point has been described. However, multiple VR images, for example, differing from each other in resolution or image size, can be generated for a single base point. With such a configuration, the user can select the VR image distributed from the distribution server 14 by considering factors such as the rendering processing capabilities of the user terminal 15, the status of the communication line between the distribution server 14 and the user terminal 15, or the amount of data received when receiving VR images.

[0082] Furthermore, in Embodiment 1, as an example, the aspect of setting one base point has been described; however, multiple base points can be set, and one or more VR images can be generated for each of the multiple set base points. According to this configuration, a user can select from multiple VR images corresponding to multiple base points a VR image corresponding to the position of the viewpoint that the user wants the display device of user terminal 15 to display. Specifically, for example, in the case where the camera is focused on a baseball game, one user allows their own or her own user terminal 15's display device to display an image with the catcher's position as the viewpoint. At this time, another user allows their own or her own user terminal 15's display device to display an image with the pitcher's position as the viewpoint.

[0083] Furthermore, in Embodiment 1, when the position of the base point is set before generating the object model, the model generation unit 301 can generate an object model corresponding to the foreground object based on the position of the set base point. Specifically, for example, the model generation unit 301 can skip or simplify the generation of a portion of the object model that is not visible when viewing the object model from the set base point position, so that a highly accurate object model is generated only when viewing the object model from the base point position. This configuration allows for a reduction in the amount of computation required to generate an object model corresponding to the foreground object.

[0084] Furthermore, in Embodiment 1, as an example, aspects of generating VR images (in which an image obtained by a monocular camera is displayed on the display device of the user terminal 15) have been described, but this is not limiting. Specifically, for example, the image processing device 100 can generate two VR images from which two images that allow stereoscopic vision can be cut out on the display device of the user terminal 15. In this case, for example, the image processing device 100 sets two base points that cause appropriate parallax and generates VR images corresponding to the two correspondingly set base points. In addition, in order to allow the cutting out of images that allow stereoscopic vision even if the user changes the viewing direction of the image displayed on the user terminal 15, the image processing device 100 can generate multiple VR images corresponding to the positions obtained by rotating the two base points, the centers of which are located at the midpoint between the two base points.

[0085] Other embodiments

[0086] The embodiments of the present invention can also be implemented by the following method: providing software (including computer program products of computer programs) that performs the functions of the above embodiments to a system or device via a network or various storage media, and the computer (central processing unit (CPU) or microprocessor unit (MPU) of the system or device) reads and executes the computer program.

[0087] According to this disclosure, VR images that allow the generation of directional images can be generated. These directional images can represent, in high definition, images corresponding to objects existing at locations far from a base point.

[0088] While this disclosure has been described with reference to embodiments, it should be understood that this disclosure is not limited to the disclosed embodiments. The scope of the appended claims should be given the broadest interpretation to cover all such modifications and equivalent structures and functions.

Claims

1. An image processing apparatus, comprising: A model acquisition component is used to acquire an object model based on multiple captured images obtained from cameras at multiple locations, the object model indicating the three-dimensional shape of an object present in the captured area; Determining components are used to determine the position of virtual objects in virtual space; A texture acquisition component is used to acquire the texture of the virtual object; The setting component is used to set the position of the base point in the virtual space; as well as An image generation component is used to generate a VR image corresponding to the base point based on the position of the base point, the object model, the virtual object, and the texture of the virtual object.

2. The image processing apparatus according to claim 1, wherein, The determining component determines the position of the virtual object in the virtual space based on the position of the base point.

3. The image processing apparatus according to claim 1, wherein, The determining component determines the location near the base point as the location of the virtual object in the virtual space.

4. The image processing apparatus according to claim 1, wherein, The determining component determines the position of the virtual object in the virtual space based on the position of the base point and the position of the object model.

5. The image processing apparatus according to claim 4, wherein, The determining component determines the position of the virtual object in the virtual space as the position of the object model that is not obscured by the virtual object when viewed from the position of the base point while the virtual object is arranged in the virtual space.

6. The image processing apparatus according to claim 1, wherein, The determining component determines the location of the virtual object in the virtual space based on input made by the user through an input device.

7. The image processing apparatus according to claim 1, wherein, The setting component sets the position of the base point in the virtual space based on input from the user via an input device.

8. The image processing apparatus according to claim 1, wherein, The texture acquisition component obtains the texture by generating a texture of the virtual object based on a virtual viewpoint image, wherein the virtual viewpoint image is obtained by virtually photographing at least a portion of the object model using a virtual camera.

9. The image processing apparatus according to claim 1, wherein, The texture acquisition component obtains the texture by generating a texture of the virtual object based on an image obtained by cutting out an image region representing at least a portion of the object from the plurality of captured images.

10. The image processing apparatus according to claim 1, wherein, The model acquisition component obtains the object model by generating the object model based on the multiple captured images.

11. An image processing method, comprising: The model acquisition step is used to obtain an object model based on multiple captured images obtained from multiple locations, the object model indicating the three-dimensional shape of an object existing in the captured area; The determination steps are used to determine the location of virtual objects in virtual space; The texture acquisition step is used to obtain the texture of the virtual object; The setup steps are used to set the position of the base point in the virtual space; as well as The image generation step is used to generate a VR image corresponding to the base point based on the position of the base point, the object model, the virtual object, and the texture of the virtual object.

12. A computer-readable storage medium storing a computer program, said computer program being implemented when executed by a processor: The model acquisition step is used to obtain an object model based on multiple captured images obtained from multiple locations, the object model indicating the three-dimensional shape of an object existing in the captured area; The determination steps are used to determine the location of virtual objects in virtual space; The texture acquisition step is used to obtain the texture of the virtual object; The setup steps are used to set the position of the base point in the virtual space; as well as The image generation step is used to generate a VR image corresponding to the base point based on the position of the base point, the object model, the virtual object, and the texture of the virtual object.

13. A computer program product comprising a computer program, said computer program being implemented when executed by a processor: The model acquisition step is used to obtain an object model based on multiple captured images obtained from multiple locations, the object model indicating the three-dimensional shape of an object existing in the captured area; The determination steps are used to determine the location of virtual objects in virtual space; The texture acquisition step is used to obtain the texture of the virtual object; The setup steps are used to set the position of the base point in the virtual space; as well as The image generation step is used to generate a VR image corresponding to the base point based on the position of the base point, the object model, the virtual object, and the texture of the virtual object.

Citation Information

Patent Citations

  • Image processing apparatus and image processing method

    JP2020068513A