Image generation apparatus and method, program, and storage medium

The image generation device addresses the issue of object occlusion by adjusting objects to maintain the desired viewpoint, ensuring complete visibility during 3D image changes.

JP2025162841APending Publication Date: 2025-10-28CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024066302
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-16
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Conventional 3D image generation systems fail to prevent objects from being obscured by obstacles while maintaining the desired user viewpoint, especially when multiple objects are involved, leading to restricted viewpoints and incomplete object visibility.

Method used

An image generation device that includes a conversion means to change the camera's viewpoint, a detection means to identify hidden areas, and an adjustment means to correct objects, ensuring the desired viewpoint is maintained by adjusting objects to prevent occlusion.

Benefits of technology

Prevents objects from being occluded in a three-dimensional image space while allowing the user to maintain their desired viewpoint, ensuring complete object visibility during viewpoint changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025162841000001_ABST
    Figure 2025162841000001_ABST
Patent Text Reader

Abstract

To provide an image generation apparatus configured to prevent an obstacle from blocking an object while maintaining a viewpoint desired by a user in a three-dimensional image space.SOLUTION: An image generation apparatus generates a second image viewed from a different camera viewpoint, from a first image having a plurality of objects including a first object and a second object different from the first object, the image generation apparatus including: a transformation unit which transforms a first image into a transformed image viewed from a different camera viewpoint; a detection unit which detects a region of the second object, in the transformed image, hidden by the first object located on the second object; and an adjustment unit which corrects one of the objects to adjust the detected hidden region of the second object, to generate a second image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image generation device that can prevent an object from becoming invisible due to being blocked by an obstacle when generating a three-dimensional image that allows the viewpoint to be moved. [Background technology]

[0002] In recent years, with the improvement in the processing power of 3D image generation devices, systems that allow the viewpoint to be moved within 3D space have been adopted in a variety of fields. At the same time, the degree of freedom of the viewpoint within 3D space has also increased, and measures are required to prevent obstacles from overlapping between the object (subject) and the viewpoint.

[0003] For example, in Patent Document 1, in a game set in a virtual three-dimensional space, character objects and various structures placed at predetermined positions are displayed as display objects, and a viewpoint position and a point of gaze that serve as the basis for the display range of the game screen are set according to the situation. This allows the player to control the character while viewing images that appear as if they were taken by a virtual camera that moves freely within the three-dimensional game field. A method has been proposed for limiting the position of the virtual camera so that obstacles such as buildings do not get between the virtual camera and the objects. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2003-334380 Summary of the Invention [Problem to be solved by the invention]

[0005] However, in the conventional technology disclosed in the above-mentioned Patent Document 1, if an obstacle overlaps between the user's desired viewpoint and the object to be focused on, the viewpoint is restricted, making it impossible to set the desired viewpoint. Also, to prevent the object to be focused on from being obscured by the obstacle, the obstacle can be displayed semi-transparently. However, if there are multiple objects to be focused on and the obstacle is one of the multiple objects to be focused on, the object cannot be deleted.

[0006] The present invention has been made in consideration of the above-mentioned problems, and its purpose is to provide an image generation device that can prevent objects from being occluded by obstacles while maintaining the viewpoint desired by the user in a three-dimensional image space. [Means for solving the problem]

[0007] The image generating device of the present invention is an image generating device that generates a second image by changing the camera's viewpoint from a first image having a plurality of objects including a first object and a second object different from the first object, and is characterized by comprising: a conversion means that converts the first image into a converted image by changing the camera's viewpoint; a detection means that detects, in the converted image, an area of ​​the second object that is hidden by the first object that is located in front of the second object; and an adjustment means that corrects one of the plurality of objects to adjust the hidden area of ​​the detected second object, and generates the second image. [Effects of the Invention]

[0008] According to the present invention, it is possible to prevent an object from being occluded by an obstacle in a three-dimensional image space while maintaining the viewpoint desired by the user. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a diagram showing the configuration of an image generating apparatus according to a first embodiment of the present invention. [Figure 2] 4 is a flowchart showing an image generation process executed by the image processing device. [Figure 3] FIG. 2 is a diagram for explaining an imaging unit. [Figure 4] FIG. 10 is a diagram for explaining a viewpoint change method. [Figure 5] FIG. 10 is a diagram for explaining calculation of a shielded area. [Figure 6A] 10 is a flowchart illustrating a shielding adjustment method. [Figure 6B] 10A and 10B are diagrams for explaining a shielding adjustment method. [Figure 6C] 10A and 10B are diagrams for explaining a shielding adjustment method. [Figure 7] 10A and 10B are diagrams for explaining another example of a shielding adjustment method. [Figure 8] 10A and 10B are diagrams for explaining another example of a shielding adjustment method. [Figure 9] FIG. 1 is a diagram for explaining the effects of the first embodiment. [Figure 10] 10 is a flowchart showing an image generation process executed by the image processing device according to the first modification of the first embodiment. [Figure 11] 10A and 10B are diagrams for explaining a shielding adjustment method in Modification 1 of the first embodiment. [Figure 12] FIG. 10 is a diagram for explaining the effect of the first modification of the first embodiment. [Figure 13] 10 is a flowchart showing an image generation process executed by an image processing device according to a second modification of the first embodiment. [Figure 14] 10 is a flowchart showing an image generation process executed by an image processing device according to a second embodiment. [Figure 15] FIG. 10 is a schematic diagram of a user interface according to the second embodiment. [Figure 16] FIG. 10 is a diagram for explaining a viewpoint change method in the second embodiment. [Figure 17] FIG. 10 is a diagram for explaining an image after a viewpoint change in the second embodiment. [Figure 18] FIG. 10 is a diagram for explaining the effect of the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0011] (First embodiment) <Configuration> 1 is a diagram showing the configuration of an image generation device according to a first embodiment of the present invention. Hereinafter, the image generation device for generating a virtual viewpoint image of this embodiment will be described with reference to FIG.

[0012] 1 generates and outputs an image (converted image) seen from a virtual camera viewpoint (hereinafter referred to as "camera viewpoint") designated by a user based on input three-dimensional image data. The image generating device 100 of this embodiment is configured to include an image processing device 200, a user interface 300, and an imaging device 400.

[0013] The user interface 300 has installed therein a dedicated application for performing various instruction operations such as an instruction to display a camera viewpoint image and an instruction to change the viewpoint on the image processing device 200.

[0014] The imaging device 400 is provided with an imaging section 401 capable of capturing three-dimensional data, and provides the captured three-dimensional data to the image processing device 200.

[0015] The image processing device 200 generates a camera viewpoint image based on the three-dimensional data sent from the image capturing device 400 and displays it on the user interface 300. In addition, the image processing device 200 receives a camera viewpoint change instruction sent from a dedicated application of the user interface 300 and automatically adjusts each object based on the relationships between multiple objects. Thereafter, the image processing device 200 generates a camera viewpoint image and provides it to the user interface 300.

[0016] The image processing device 200 is, for example, a server computer, and includes a processing unit 201, a data acquisition unit 202, an image generation unit 203, a coordinate conversion unit 204, an occlusion detection unit 205, and an occlusion adjustment unit 206. The occlusion adjustment unit 206 also includes the functions of an occlusion region calculation unit 207 and a destination determination unit 208.

[0017] The user interface 300 is, for example, a personal computer electrically connected to the image processing device 200, and has a display unit 301 and an operation unit 302. The photographing device 400 is electrically connected to the image processing device 200, and has a photographing unit 401 equipped with, for example, a camera capable of photographing three-dimensional data.

[0018] The processing unit 201 provided in the image processing device 200 is, for example, a CPU, and controls the entire image processing device 200 using pre-stored programs and data, thereby realizing each functional unit shown in Fig. 1. Note that the processing unit 201 may have one or more dedicated hardware components different from the CPU, and at least a part of the processing by the CPU may be executed by the dedicated hardware components. Examples of the dedicated hardware components include an ASIC (application-specific integrated circuit), an FPGA (field-programmable gate array), and a DSP (digital signal processor).

[0019] The data acquisition unit 202, image generation unit 203, coordinate conversion unit 204, occlusion detection unit 205, occlusion adjustment unit 206, occlusion area calculation unit 207, and movement destination determination unit 208 are stored as programs in ROM 209, and perform their functions by being executed by the CPU. RAM 210 temporarily stores data provided from the CPU and data provided from the outside via a communication I / F (not shown).

[0020] The display unit 301 provided in the user interface 300 is configured, for example, as a liquid crystal display, and displays a GUI and the like for the user to operate the image processing device. The operation unit 302 is configured, for example, as a keyboard, mouse, joystick, etc., and receives operations by the user to input various instructions to the processing unit 201. The operation unit 302 may be integrated with the display unit 301, for example, as a liquid crystal display equipped with a touch panel.

[0021] The image capturing unit 401 provided in the image capturing device 400 is configured, for example, with a stereo camera, and the stereo camera is operated using the operation unit 302 provided in the user interface 300. That is, the processing unit 201 receives instructions from the operation unit 302, and the processing unit 201 sends an image capturing instruction to the image capturing unit 401. The image capturing device 400 also has an object detection unit 402 and a face organ detection unit 403, which detect object information and face organ information from the captured data and transmit them to the processing unit 201. Note that, although the image capturing device 400 is configured to acquire three-dimensional data in this embodiment, the three-dimensional data may be prepared in advance by the user. In this case, the three-dimensional data is transmitted to the processing unit 201 by the data acquisition unit 202. The three-dimensional data may not be captured data, but may be computer graphics data.

[0022] <Overall operation> Next, the flow of processing in which the image generating device 100 of this embodiment automatically adjusts each object based on the relationships between multiple objects and then generates and outputs a camera viewpoint image will be described with reference to the flowchart in FIG. 2. The flowchart in FIG. 2 is realized by the processing unit (control unit) 201 receiving a viewpoint change processing instruction from the user interface 300 and executing a predetermined processing program. The processing is initiated, for example, when the user inputs 3D data for viewpoint change processing into the dedicated application mentioned above, or when the user takes a photograph with a camera capable of acquiring 3D data, and then selects to enable the processing. In the following description, the symbol "S" represents a step.

[0023] In S101, the processing unit 201 acquires three-dimensional data using the imaging unit 401. From the acquired three-dimensional data, the object detection unit 402 identifies objects, and the object information is associated with the three-dimensional data. Furthermore, if the object is a person, the face organ detection unit 403 detects face organs for each object, and the coordinate information is associated with the object as face organ information.

[0024] Next, in S102, the processing unit 201 displays a camera viewpoint change screen through the user interface so that the user can select a preferred camera viewpoint.

[0025] In S103, the processing unit 201 converts the coordinates of the three-dimensional data in response to the user's operation to change the camera viewpoint.

[0026] In S104, the processing unit 201 performs an occlusion determination on the three-dimensional data after the coordinate transformation to determine whether there is an occluded object, using the occlusion detection unit 205. The occlusion determination is performed based on the transformed coordinates. That is, the processing unit 201 calculates an occluded area of ​​the object from the object in the foreground, and searches for an object that exists within the occluded area.

[0027] In S105, the processing unit 201 determines whether or not there is an occluded area as a result of the occlusion detection in S104. If it is determined that there is an occluded area, the process proceeds to occlusion adjustment processing of the object in S106. On the other hand, if there is no occluded object, the process proceeds to image generation processing in S107.

[0028] In S106, the processing unit 201 adjusts the objects in the masked area on an object-by-object basis so that the objects are outside the masked area. The method of adjusting the objects will be described later, but the objects are moved, enlarged, or reduced based on a predetermined rule.

[0029] In S107 , the processing unit 201 generates an image from the camera viewpoint specified by the user, and displays it on the display unit 301 provided in the user interface 300 .

[0030] When the above-described processing is completed, the process ends. The above is an overview of the process of generating and outputting a camera viewpoint changed image in the image processing device 200.

[0031] <3D data generation> Next, the details of the acquisition of three-dimensional data in S101 mentioned above will be described. The imaging unit 401 can have any configuration as long as it functions as an acquisition means capable of acquiring three-dimensional data. More preferably, it should be configured to be capable of simultaneously acquiring distance distribution information and color distribution information from the same viewpoint. Examples of configurations that realize such functionality include a stereo camera equipped with two imaging systems, each consisting of an optical system capable of acquiring RGB images and an imaging element, and a ToF camera equipped with a ToF module capable of acquiring distance distribution information and an imaging system capable of acquiring RGB images. Below, as an example, a configuration of the imaging unit 401 that can simultaneously acquire distance distribution information and color distribution information using an imaging surface phase difference detection method will be described.

[0032] 3 shows the configuration of the imaging unit 401. The imaging unit 401 is configured with an imaging optical system 501 and an imaging element 502, and captures an image by collecting light from a subject located on an object plane with the imaging optical system 501 and receiving the light with the imaging element 502 located at a position that is approximately optically conjugate with the object plane.

[0033] The image sensor 502 is configured with, for example, tens of millions of pixels 503, each equipped with a photoelectric conversion element, arranged in a grid pattern. Each pixel is provided with a color filter that transmits specific wavelengths such as red, green, and blue, and these are arranged in, for example, a Bayer array, making it possible to acquire color distribution information.

[0034] One pixel 503 is configured with a microlens 504, a first photoelectric conversion unit 505, and a second photoelectric conversion unit 506, and acquires different optical information according to its position by the two photoelectric conversion units 505, 506 arranged in the horizontal direction (X direction). From the received light information of all pixels, a first image configured as the luminance distribution of light received by the first photoelectric conversion unit 505 and a second image configured as the luminance distribution of light received by the second photoelectric conversion unit 506 are obtained.

[0035] The incident surface of the image sensor 502 and the light receiving surfaces of the first photoelectric conversion unit 505 and the second photoelectric conversion unit 506 are in an approximately Fourier conjugate relationship via the microlens 504. Therefore, the light receiving surfaces of the photoelectric conversion units and the exit pupil of the image capturing optical system 501 are in an approximately optically conjugate relationship. Because the position distribution at the exit pupil corresponds to the position distribution at the light receiving surfaces of the photoelectric conversion units, providing two different photoelectric conversion units makes it possible to separate and receive light beams that have passed through different pupil regions of the image capturing optical system 501. The first and second images described above are luminance distribution information obtained by light beams that have passed through different pupil regions.

[0036] Light rays that form an image on the imaging surface (ideally) enter the same point on the image sensor 502 regardless of the position on the pupil of the imaging optical system 501 through which they pass. However, for defocused light rays, the position at which the light rays enter the image sensor 502 changes depending on the position on the pupil through which they pass. In other words, an image shift occurs according to the amount of defocus.

[0037] The amount of image shift can be calculated, for example, by stereo matching between the first and second images. A small patch of one image is matched along the epipolar line with the other image, and the image shift amount is determined by identifying the position with the highest correlation. The calculated image shift amount, along with the focal length and focus position obtained from the shooting information, is converted into a world coordinate system to generate 3D data.

[0038] Furthermore, the image data captured by the image capture unit 401 is grouped by object by the object detection unit 402. At this time, if the object is a person, the face organ detection unit 403 detects facial feature points and calculates face organ coordinates. This information is associated with the three-dimensional data and transmitted to the processing unit 201.

[0039] <Change of viewpoint> In S102 of Fig. 2, the user freely changes the viewpoint using the operation unit 302 provided in the user interface 300, and the viewpoint change information is sent to the processing unit 201. The viewpoint change is performed by rotation and translation. The user selects the rotation button 311 or translation button 312 and performs the operation according to the viewpoint change screen displayed on the user interface 300, as shown in Fig. 4(a).

[0040] When the rotation button 311 is selected, the user can rotate the camera viewpoint. The center of rotation is set to the center of gravity of the 3D data, and when the user specifies the direction of rotation, the rotation is performed around the set center of gravity. At this time, if the object protrudes from the screen due to the rotation, automatic scaling may be performed. Also, although the center of gravity of the 3D data is set as the center of rotation, this is not limitative. For example, the center of gravity of a certain object may be set as the center, or only the face of a certain object may be extracted and used as the center of gravity position.

[0041] Figure 4(b) shows the 3D data as viewed from the camera's viewpoint and from above. As shown in Figure 4(b), the viewpoint coordinates are defined as the horizontal direction of the screen as x, the vertical direction as y, and the depth direction as z. The object's viewpoint (orientation) change direction indicated by arrow A is the direction in which the object's orientation changes, and is the direction of the rotation change performed by the user operating along the screen as described above. The camera viewpoint rotation direction indicates that changing the object's orientation in the direction of arrow A (to the left) is synonymous with rotating the camera's viewpoint to the right by a rotation angle θ without changing the object's orientation.

[0042] Also, Figure 4(c) shows the displayed image after the user has changed the camera viewpoint, and the coordinates after the viewpoint change are defined as [x', y', z']. When rotation is performed in accordance with a user instruction, the processing unit 201 immediately performs coordinate transformation. A rotation axis vector n = [nx, ny, 0] is calculated from the direction of the viewpoint change by the user, and a rotation angle θ is determined according to the amount of rotation operated by the user. In the coordinate transformation, the coordinates after the viewpoint change, [x', y', z'], are calculated from these parameters using a rotation matrix R. In this case, the rotation matrix R is expressed as follows:

[0043]

number

[0044] Here, the rotation axis vector n is a unit vector.

[0045] When the parallel movement button 312 is pressed, the user can translate the viewpoint in six directions: up and down, left and right, and forward and backward. The coordinate conversion at this time is calculated by offsetting the coordinates by the amount of the viewpoint moved by the user.

[0046] In this embodiment, the user performs the camera viewpoint change operation using the operation unit 302 provided in the user interface 300, but it is also possible to provide multiple predetermined camera viewpoint trajectories and allow the user to select the camera viewpoint trajectory of their choice.

[0047] <Occlusion detection method> Next, we will explain the occlusion detection method. Figure 5(a) is a view of the 3D data viewed from above, Figure 5(b) is a view from camera viewpoint 1 in Figure 5(a), and Figure 5(c) is a view from camera viewpoint 2 as changed by the user in Figure 5(a). Here, the object on the right is object 1, and the object on the left is object 2. Also, with 3D data from a two-viewpoint camera, only data captured within the camera is generated, so the back side of the object does not exist.

[0048] As shown in FIG. 5(c), when the viewpoint is changed, object 2, which is the subject, may overlap with object 1 and become invisible. Therefore, in this embodiment, object occlusion detection is performed after the camera viewpoint is changed, and the occluded object is searched for. Note that in FIG. 5(c), the direction of change of the viewpoint (orientation) of the object indicated by arrow A is the direction in which the orientation of the object is changed. Changing the viewpoint (orientation) of the object in the direction of arrow A (leftward) is synonymous with rotating the camera viewpoint to the right from camera viewpoint 1 to camera viewpoint 2 without changing the orientation of the object, as shown in FIG. 5(d).

[0049] When the viewpoint is changed, the processing unit 201 first selects the object 1 in the foreground and calculates the area occluded by the object 1. As a calculation method, as shown in the hatched area in Fig. 5(d), points having the same values ​​as the coordinate group [x', y', z'] after the coordinate transformation of the object 1 are set as the occluded area. Here, the coordinates [x', y'] are the pixel position on the image, and z' is the distance from the viewpoint.

[0050] Next, the processing unit 201 checks whether the object 2 is in an occluded area. In this embodiment, the search is performed based on the facial organ information attached to the three-dimensional data. That is, if the facial organ coordinates are within the occluded area, it is determined that the object having the facial organ is occluded. Here, occlusion determination may be performed without using the facial organ information, for example, by setting a certain occlusion rate, such as the percentage of the area of ​​the object occluded by the object. Alternatively, the user may set areas that should not be occluded. In this case, the user can specify the areas that should not be occluded in advance via the user interface 300.

[0051] <Shielding adjustment method> Next, an occlusion adjustment method will be described with reference to FIGS. 6A to 6C. When occlusion is detected by the occlusion detection described above, the processing unit 201 performs occlusion adjustment by the occlusion adjustment unit 206. Specifically, the occlusion adjustment unit 206 translates or enlarges the occluded object 2, or translates or reduces the occluding object 1. In this embodiment, the translation of the occluded object 2 will be described, but the present invention is not limited to this. For example, when the object of interest is not to be moved due to selection of the object of interest, which will be described later, a process other than translation is selected.

[0052] When the viewpoint of the camera is rotated and changed, and the occlusion detection unit detects an occlusion, the occlusion adjustment unit 206 translates the occluded object. The direction of the translation is determined based on the flowchart shown in FIG. 6A.

[0053] First, in S201, the processing unit 201 receives a rotation instruction from the user and performs a rotation change of the viewpoint.

[0054] In S202, the processing unit 201 detects occlusion, and in S203, the processing unit 201 predicts the trajectory of the occluded area from the rotation axis and the rotation direction.

[0055] Figure 6B shows the trajectory of the occluded area as seen from the camera's viewpoint, and Figure 6C shows the trajectory of the occluded area as seen from above. Once the user's viewpoint change direction is determined in this way, the trajectory of the occluded area can be predicted from the change direction of the camera's viewpoint. In this case, the viewpoint that the user can change is set to θ = 80° so that facial organs are not hidden by the user's own organs (for example, so that the facial organs of object 1 are not hidden by object 1's own arms, etc.).

[0056] In S204, the processing unit 201 searches for an occlusion avoidance position to determine the destination of the occluded object 2. Based on the prediction of the occluded area and the maximum viewpoint change angle θ=80°, a search is made for a destination that can avoid the occlusion area.

[0057] In S205, the processing unit 201 determines whether or not there is an avoidance position based on the result of the search in S204. If there is an avoidance position, the process proceeds to S207, and if there is not, the process proceeds to S206.

[0058] In S207, the processing unit 201 determines the destination with the shortest (minimum) movement amount as the movement direction based on the information of the movement destination searched in S204. Also, in order to perform parallel movement without causing the user to feel uncomfortable, the movement direction is selected so that the movement of the object 2 is continuous. Once the movement direction is determined, the processing unit 201 translates the occluded object 2 in the determined movement direction so that the occluded area of ​​the object 2 is kept within a predetermined size.

[0059] In the above description, the object is moved so that the occluded area of ​​the object 2 is kept below a predetermined size. However, for example, the object may be moved so that the user can set an unoccludable area of ​​the object, and when it is detected that the unoccludable area has been occluded, the object is moved so that the unoccludable area is not hidden.

[0060] Next, a case will be described in which, as a result of trajectory prediction of the occlusion area, another object exists in the above-mentioned occlusion avoidance direction, and therefore an avoidable destination cannot be set within the screen (No in S205).

[0061] FIG. 7(a) is a diagram of a scene in which three human objects are lined up, as viewed from the camera's viewpoint, and FIG. 7(b) is a diagram viewed from above. Here, the object on the right is object 1, the object in the center is object 2, and the object on the left is object 3. FIG. 7(c) is a diagram as viewed from camera viewpoint 2 when the user rotates the viewpoint (orientation) of the object in the direction of arrow A in FIG. 7(a) and changes the viewpoint from camera viewpoint 1 to camera viewpoint 2 in FIG. 7(b). Note that the direction of change of the object's viewpoint (orientation) indicated by arrow A in FIG. 7(a) is the direction in which the user changes the orientation of the object. Changing the object's viewpoint (orientation) in the direction of arrow A (leftward) is synonymous with rotating the camera's viewpoint to the right from camera viewpoint 1 to camera viewpoint 2 without changing the orientation of the object, as shown in FIG. 7(b).

[0062] In this scene, if we consider the maximum viewpoint change angle, moving parallel as before will result in a collision with another human object, object 3 (Figure 8(b)). Furthermore, if we move in a direction to avoid the collision, the object will go off the screen.

[0063] Therefore, in this embodiment, the processing unit 201 performs overall reduction processing on the overflowing area in S207 as shown in Fig. 8(b). Therefore, it is possible to select a direction in which the movement amount is small and continuous movement is possible. However, if there is an object further above and it is difficult to avoid the occluded area, the processing unit 201 selects a movement direction in which the area that can be avoided is the largest.

[0064] Once the movement direction is determined in S206, in S207 the processing unit 201 translates the occluded object 2 in the determined movement direction so that the occluded area of ​​the object 2 is kept within a predetermined range. If the movement is subsequently restricted, the object continues to move outside the range and the entire image is displayed in a reduced size, thereby fitting the protruding object within the display range. Once the translation of the occluded object 2 is complete, the processing unit 201 generates an image of the changed viewpoint and displays it on the display unit 301 provided in the user interface 300.

[0065] Next, a process will be described for the case where the facial organ coordinates are occluded by one's own arm or the like (for example, when the facial organs of object 1 are hidden by object 1's own arm or the like).

[0066] When occlusion is detected, the processing unit 201 determines whether the object itself is occluding it, and if so, adjusts by rotating the occluded object itself. The rotation axis is parallel to the rotation axis used for changing the viewpoint and is a line passing through the center of gravity of the object itself. The rotation angle is calculated based on the occluded area calculated earlier and the amount of change in the occluded area and facial organ coordinates due to the object's own rotation, and is determined so that the facial organ coordinates are outside the occluded area.

[0067] Furthermore, as described above, 3D data generated by a two-viewpoint camera only generates data that is captured within the camera, and therefore the back side of an object does not exist. Therefore, in this embodiment, even if the back side of an object becomes visible due to a change in viewpoint by the user, the object itself may be rotated to make it invisible.

[0068] <Effects> 9(a) and 9(b) show images output as a result of image processing by the image processing device 200 when the user changes the viewpoint for the scenes shown in Fig. 6A to Fig. 6C and the scene shown in Fig. 7, respectively. In this way, by performing the processing of this embodiment, even if the user freely changes the viewpoint, it is possible to display a screen that maintains the viewpoint desired by the user without multiple objects overlapping and becoming invisible, or without losing track of an object.

[0069] <Variation 1> In the above embodiment, the operation of searching for a destination outside the occluded area for the occluded object 2 and translating only the occluded object 2 has been described. However, the effects of the present invention can also be obtained with other processing flows.

[0070] 10 will be described below as Modification 1 of the first embodiment. Also, a description of the same processes (S201, S202, S204) as those in the above-described embodiment will be omitted.

[0071] In S303, the processing unit 201 creates a group of objects for each distance z' from the viewpoint in order to more easily search for the destination. That is, in the coordinates after viewpoint transformation, objects whose distance z' from the viewpoint is within a certain range are considered to be in one group.

[0072] Next, we will explain how to create groups using the scene in Figure 7(a). First, calculate the center of gravity of each object after coordinate transformation. Next, compare the distance z' from the viewpoint of each calculated center of gravity position. Table 1 compares the z' coordinate values ​​of each object.

[0073] [Table 1]

[0074] The processing unit 201 creates object groups in ascending order of z' values. That is, first, the processing unit 201 creates group A based on object 1, which is closest to camera viewpoint 2, as shown in FIG. 11(a). A search is performed to see if there is another object that belongs to the same group and whose center of gravity position z' is within 15. In this scene, there is no object within 15, so the process moves to the next step. Next, group B is created based on objects that are closest to camera viewpoint 2 among the objects that do not belong to group A. Here, objects that belong to the same group are searched for. In this scene, object 3 is within 15 of object 2, so objects 2 and 3 are set as group B.

[0075] In S305, the processing unit 201 searches for a destination for each group. That is, since object 2 is occluded, the same movement operation is performed on objects 2 and 3 belonging to group B.

[0076] In S306, the processing unit 201 searches for a destination outside the occluded area as shown in Fig. 11(b) and determines the direction of movement that will minimize the movement distance. At this time, since movement is performed for each group, there will be no collisions with other objects at the destination. Therefore, it is possible to select a destination that will minimize the movement distance.

[0077] Once the movement direction is determined, group B is moved out of the occlusion area, preventing the object from becoming invisible due to occlusion as shown in FIG.

[0078] <Variation 2> Although the three-dimensional data used in the above-described embodiment is a still image, the present invention can achieve the same effect even when the data is a moving image. Below, as a second modification of the first embodiment, processing when the three-dimensional data is a moving image will be described with reference to the flowchart in Fig. 13. Note that the processing of S401 to S407 in Fig. 13 is basically the same as the processing of S101 to S107 in Fig. 2, so the following description will focus on the differences.

[0079] In this embodiment, a video is treated as data consisting of a series of multiple still images. When acquiring three-dimensional data in S401, the processing unit 201 breaks down the image into frames and performs processing. Specifically, when a user instructs to change the viewpoint, in S402 to S404, the same processing as in S101 to S104 is performed for each frame up to the occlusion detection processing.

[0080] When occlusion is detected by the occlusion detection process in S405, the processing unit 201 adjusts the occlusion of the object in S406. At this time, the processing proceeds while comparing the correction amount between each frame so that it is equal to or less than a predetermined value. As a result, in S407, the moving image after adjustment between frames becomes continuous, making it possible to generate a moving image that feels less strange and prevents objects from becoming invisible.

[0081] Alternatively, the correction amount may be the same for each frame. That is, the maximum correction amount is calculated for all frames, and that maximum value is then offset for all frames. This method also makes the movement between frames continuous, so it is possible to generate a video that is less awkward and prevents objects from becoming invisible.

[0082] (Second embodiment) In the first embodiment, a screen display was described in which the user can freely change the viewpoint and adjust the occlusion of objects caused by the viewpoint change. In the second embodiment, a method will be described in which the user designates an object that is particularly prioritized as an object of interest, and a screen display is created in accordance with the viewpoint change. Note that the same components as in the first embodiment are given the same reference numerals, and their description will be omitted. Also, since the processes of S501 and S503 to S508 in FIG. 14 are basically the same as S101 to S107 in FIG. 2, the following description will focus on the differences.

[0083] 14 is a flowchart illustrating the processing of this embodiment. After acquiring three-dimensional data (S501), the user can select an object of interest (S502).

[0084] The object of interest is selected via a display unit 301 and an operation unit 302 provided in the user interface 300. Figures 15(a) and 15(b) show examples of the display unit 301 and the operation unit 302 through which the user selects the object of interest.

[0085] The user selects the object of interest selection button 151 shown in Fig. 15(a), and then operates the arrow shown in Fig. 15(b) via the operation unit 302 to select an object of interest. Either one object or multiple objects may be selected.

[0086] Once the selection is complete, the user selects the Next button 152 to proceed to the next process.

[0087] In S503, as in the first embodiment, the user freely changes the viewpoint using the operation unit 302 provided in the user interface 300, and the viewpoint change information is sent to the processing unit 201. The viewpoint change is performed by rotation and translation. When an object of interest is selected as shown in FIG. 16, the center of gravity of the three-dimensional data for the object of interest is set as the rotation center, and when the user specifies the rotation direction, rotation processing is performed around the set center of gravity. At this time, if the object protrudes from the screen due to rotation, it may be automatically reduced.

[0088] In S504, the processing unit 201 converts the coordinates of the three-dimensional data in accordance with the user's operation to change the camera viewpoint, similar to S103 in the first embodiment.

[0089] Next, an occlusion detection process similar to S104 in the first embodiment is performed (S505), and then if an object is occluded in S506 as shown in FIG. 17, an occlusion adjustment process is performed (S507).

[0090] When a rotation process is performed during a change in camera viewpoint, and the occlusion detection unit 205 detects occlusion, the occlusion adjustment unit 206 reduces or enlarges objects other than the focused object. Here, unlike the first embodiment, an object of interest has been set, and therefore the object of interest is not adjusted in the occlusion adjustment process. Therefore, in this process, enlargement or reduction processing is performed on objects other than the object of interest.

[0091] When adjusting an object in the foreground, it is scaled down, and when adjusting an object in the background, it is scaled up. Furthermore, scaling up and down is performed using information that predicts the trajectory of the occluded area. In other words, the occlusion trajectory when scaling up or down is calculated, and the center where the scaling rate is smallest is determined.

[0092] 18 shows an image display generated after occlusion adjustment is performed by the processing of this embodiment. In this way, when the user selects an object of interest, the occlusion adjustment processing is performed without adjusting the position of the object of interest, thereby achieving the same effect as in the first embodiment.

[0093] In the first and second embodiments, if the rotation matrix is ​​calculated for all three-dimensional data and the screen is switched each time the user changes the viewpoint, this places a heavy load on the calculation process and can result in a huge processing time depending on the performance of the processing device. Therefore, it is also possible to extract only the outline of each object and proceed with the processing. In this case, when restoring the data to three-dimensional data, it is better to thicken the outline by a certain amount to prevent overlapping of objects.

[0094] In addition, in the first and second embodiments described above, by saving the display of the trajectory of the user's viewpoint change as each frame, it becomes possible to create a video with the same effect as in this embodiment. The user may specify the start viewpoint and end viewpoint in each frame to create a viewpoint video that connects them.

[0095] Furthermore, the image generating device 100 in the first and second embodiments described above is configured to include the photographing device 400, but it may also be configured to input three-dimensional data from outside instead of photographing with a photographing device. In this case, the photographing device 400 is not provided, and instead a three-dimensional data input unit is provided.

[0096] The disclosure of this specification includes the following image generating device, image generating method, program, and storage medium.

[0097] (Item 1) 1. An image generating device that generates a second image by changing a camera viewpoint from a first image having a plurality of objects including a first object and a second object different from the first object, a conversion means for converting the first image into a converted image obtained by changing the viewpoint of the camera; a detection means for detecting, in the converted image, an area of ​​the second object that is hidden by a first object located in front of the second object; an adjustment means for correcting any one of the plurality of objects to adjust the hidden area of ​​the detected second object and generating the second image; An image generating device comprising:

[0098] (Item 2) 2. The image generating device according to item 1, wherein the second object is a person, and the detection means detects whether or not a facial organ of the second object is occluded.

[0099] (Item 3) 3. The image generating device according to item 1 or 2, wherein the detection means detects whether an unobscurable area set by a user is hidden.

[0100] (Item 4) 4. The image generating device according to any one of items 1 to 3, wherein the adjustment means translates any one of the plurality of objects.

[0101] (Item 5) 5. The image generating device according to any one of items 1 to 4, wherein the adjustment means reduces and displays any one of the plurality of objects.

[0102] (Item 6) 6. The image generating device according to any one of items 1 to 5, wherein the adjustment means rotates any one of the plurality of objects.

[0103] (Item 7) The image generating device described in any one of items 1 to 6, characterized in that the adjustment means predicts the trajectory of the hidden area based on the trajectory of the change in the viewpoint of the camera, and determines the movement direction of one of the plurality of objects.

[0104] (Item 8) 8. The image generating device according to any one of items 1 to 7, wherein a moving image is generated by changing the viewpoint of the camera.

[0105] (Item 9) 9. The image generating device according to any one of items 1 to 8, wherein the adjustment means corrects, among the plurality of objects, objects other than a target object set by a user.

[0106] (Item 10) The image generating device described in any one of items 1 to 9, characterized in that the adjustment means divides the plurality of objects into groups based on the distance from the viewpoint of the camera, adjusts the hidden area for each divided group, and generates the second image.

[0107] (Item 11) 11. The image generating device according to any one of items 1 to 10, further comprising means for enabling the adjustment means by a user.

[0108] (Item 12) 12. The image generating device according to any one of items 1 to 11, wherein the conversion means extracts and converts only the contours of the plurality of objects.

[0109] (Item 13) 13. The image generating device according to any one of items 1 to 12, wherein the first image is a moving image, and the adjustment means determines the amount of correction so that the amount of correction for each frame of the moving image is continuous.

[0110] (Item 14) 14. The image generating device according to any one of items 1 to 13, wherein the adjustment means adjusts the hidden area so as to reduce the amount of correction of the plurality of objects, and generates the second image.

[0111] (Item 15) 1. An image generation method for generating a second image by changing a camera viewpoint from a first image having a plurality of objects including a first object and a second object different from the first object, comprising: a transformation step of transforming the first image into a transformed image in which the viewpoint of the camera is changed; a detecting step of detecting, in the converted image, an area of ​​the second object that is hidden by a first object that is positioned in front of the second object; an adjustment step of correcting any one of the plurality of objects to adjust the hidden area of ​​the detected second object and generating the second image; An image generating method comprising:

[0112] (Item 16) A program for causing a computer to function as each of the means of the image generating device according to any one of items 1 to 14.

[0113] (Item 17) A computer-readable storage medium storing a program for causing a computer to function as each of the means of the image generating device according to any one of items 1 to 14.

[0114] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0115] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0116] 201: Processing unit, 203: Image generation unit, 204: Coordinate conversion unit, 205: Occlusion detection unit, 206: Occlusion adjustment unit, 207: Occlusion area calculation unit, 208: Movement destination determination unit, 209: ROM, 210: RAM, 301: Display unit, 302: Operation unit, 401: Photography unit

Claims

1. 1. An image generating device that generates a second image by changing a camera viewpoint from a first image having a plurality of objects including a first object and a second object different from the first object, a conversion means for converting the first image into a converted image obtained by changing the viewpoint of the camera; a detection means for detecting, in the converted image, an area of ​​the second object that is hidden by a first object located in front of the second object; an adjustment means for correcting any one of the plurality of objects to adjust the hidden area of ​​the detected second object and generating the second image; An image generating device comprising:

2. 2. The image generating apparatus according to claim 1, wherein the second object is a person, and the detecting means detects whether or not a facial part of the second object is hidden.

3. 2. The image generating device according to claim 1, wherein the detecting means detects whether or not an unobstructable area set by a user is hidden.

4. 2. The image generating device according to claim 1, wherein the adjustment means translates any one of the plurality of objects.

5. 2. The image generating device according to claim 1, wherein the adjustment means reduces and displays any one of the plurality of objects.

6. 2. The image generating device according to claim 1, wherein the adjustment means rotates any one of the plurality of objects.

7. The image generating device according to claim 1 , wherein the adjustment means predicts a trajectory of the hidden area based on a trajectory of a change in the viewpoint of the camera, and determines a movement direction of one of the plurality of objects.

8. The image generating device according to claim 1, wherein a moving image is generated by changing the viewpoint of the camera.

9. 2. The image generating apparatus according to claim 1, wherein the adjustment means corrects, among the plurality of objects, objects other than a target object set by a user.

10. The image generating device according to claim 1, characterized in that the adjustment means divides the plurality of objects into groups based on their distance from the viewpoint of the camera, adjusts the hidden area for each divided group, and generates the second image.

11. 2. The image generating device of claim 1 further comprising means for user activation of said adjustment means.

12. 2. The image generating apparatus according to claim 1, wherein said conversion means extracts and converts only the contours of said plurality of objects.

13. 2. The image generating device according to claim 1, wherein the first image is a moving image, and the adjusting means determines the amount of correction so that the amount of correction for each frame of the moving image is continuous.

14. 2. The image generating device according to claim 1, wherein the adjustment means adjusts the hidden area so as to reduce the amount of correction of the plurality of objects, and generates the second image.

15. 1. An image generation method for generating a second image obtained by changing a camera viewpoint from a first image having a plurality of objects including a first object and a second object different from the first object, the method comprising: a transformation step of transforming the first image into a transformed image in which the viewpoint of the camera is changed; a detecting step of detecting, in the converted image, an area of ​​the second object that is hidden by a first object located in front of the second object; an adjustment step of correcting any one of the plurality of objects to adjust the hidden area of ​​the detected second object and generating the second image; An image generating method comprising:

16. A program for causing a computer to function as each of the means of the image generating apparatus according to any one of claims 1 to 14.

17. A computer-readable storage medium storing a program for causing a computer to function as each of the means of the image generating apparatus according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Virtual camera position control program for three- dimensional game

    JP2003334380A