Image processing device and control method for the image processing device
The image processing apparatus addresses the issue of cut-off 3D objects by determining their position and size to overlap with non-rendering areas in the virtual environment, enhancing the natural appearance and immersion of 3D compositions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2025-01-10
- Publication Date
- 2026-07-23
Smart Images

Figure 2026121008000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus and a control method for an image processing apparatus.
Background Art
[0002] Conventionally, techniques for obtaining distance distribution information using a stereo camera, a light field camera, etc. are known. By performing perspective projection conversion on the distance distribution information, a point cloud is obtained. By polygonizing the point cloud, a three-dimensional surface model having a surface is generated. By obtaining an image (color distribution information) together with the distance distribution information or the three-dimensional surface model generated from the distance distribution information, it is possible to generate a 3D (three-dimensional) object provided with texture information. Unlike a normal two-dimensional image, a 3D object has the advantage that a user can enjoy viewing from an arbitrary viewpoint.
[0003] Patent Document 1 discloses a technique for making a subject included in image data easily distinguishable by creating a refocused image with a changed focal length in the image data photographed by a light field camera.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] 3D objects created using a stereo camera or light field camera may not fit entirely within the field of view, resulting in some parts being cut off. When parts of the subject are cut off, the 3D object will appear partially missing. For example, if a 3D object is composited with a 3D background to enhance immersion, the missing portion of the subject may make it appear as if it's floating in mid-air. When a 3D object with missing parts is composited with a 3D background, users may perceive this as unnatural.
[0006] Therefore, the present invention aims to provide an image processing apparatus that reduces the unnatural appearance of an image when a 3D object with a missing part of the subject is combined with a background in a three-dimensional space. [Means for solving the problem]
[0007] The image processing apparatus according to the present invention includes an acquisition means for acquiring information about a virtual environment and information about three-dimensional objects to be placed in the virtual environment, and a determination means for determining the position and size of the three-dimensional objects when they are placed in the virtual environment, wherein the three-dimensional objects placed in the virtual environment at the position and size determined by the determination means are characterized in that a part or all of the non-rendering area of the three-dimensional objects that is not to be displayed overlaps with the non-rendering area of the virtual environment. [Effects of the Invention]
[0008] According to the present invention, it is possible to reduce the sense of incongruity in an image when a 3D object with a missing part of the subject is combined with a background in a three-dimensional space. [Brief explanation of the drawing]
[0009] [Figure 1] This diagram illustrates the configuration of an image generation device. [Figure 2] This diagram illustrates the configuration of an imaging device. [Figure 3] This is a diagram illustrating the configuration of the imaging unit. [Figure 4] This is a diagram illustrating the missing regions of a 3D object. [Figure 5] This flowchart illustrates the process for determining the placement information of 3D objects. [Figure 6] This figure shows an example of 3D object placement. [Figure 7] This diagram illustrates the area where the region of interest of a 3D object is placed. [Figure 8] This diagram illustrates the area where the entire 3D object is placed. [Figure 9] This is a diagram illustrating the area of interest in a 3D object. [Modes for carrying out the invention]
[0010] <Embodiment> Embodiments of the present invention will be described below with reference to the drawings. Note that the following embodiments do not limit the invention as defined in the claims. The various features described in the embodiments are not necessarily essential for carrying out the invention and can be combined arbitrarily. In each figure, identical components are denoted by the same reference numerals, and redundant explanations are omitted.
[0011] This embodiment is explained by an example of generating a video by arranging (compositing in 3D space) 3D objects (3-dimensional objects) generated by an imaging device in a virtual space where camera work, background, and foreground are pre-set. The examples described in this embodiment are not limiting to the present invention.
[0012] (composition) Figure 1 is a diagram illustrating the configuration of the image generation device 100. The image generation device 100 comprises an image processing device 200 and a user interface 300. Referring to Figure 1, the configuration of the image processing device 200 and the user interface 300 according to the present invention will be described.
[0013] The image generation device 100 generates and outputs an image that can be seen from a virtual camera viewpoint (hereinafter referred to as the virtual camera viewpoint) when capturing the virtual environment, based on information about the 3D object and information about the virtual environment.
[0014] The image processing device 200 generates an image as seen from a virtual camera viewpoint (hereinafter referred to as a virtual camera viewpoint image) using information about 3D objects and information about the virtual environment, and displays the generated image on the user interface 300. The user interface 300 displays the virtual camera viewpoint image to the user, which is a view of the virtual environment from a viewpoint that has been set in advance as an initial value, and accepts various operations from the user, such as instructions to change the viewpoint. The user interface 300 performs processing such as displaying the image and accepting operations using a dedicated application. The virtual environment includes a three-dimensional space, multiple components set up in the three-dimensional space, and components for extracting a predetermined 3D model or 2D image from the virtual environment, such as a virtual camera and virtual lighting. The multiple components can be foreground, background, or objects at the same distance from the virtual viewpoint when viewed from the 3D object to be synthesized. The information about the virtual environment includes data such as the position and size of these multiple components, the position and field of view size of the virtual camera, and information about the camera movement (camera work) when generating a moving image using the image extracted by the virtual camera.
[0015] The image processing device 200 is, for example, a server computer. The image processing device 200 includes a control unit 201, a data acquisition unit 202, an image generation unit 203, a region acquisition unit 204, and an object placement unit 205. The user interface 300 is, for example, a personal computer and is electrically connected to the image processing device 200. The user interface 300 includes a display unit 301 and an operation unit 302.
[0016] The control unit 201 of the image processing apparatus 200 has a memory such as a ROM (Read Only Memory), controls the entire image processing apparatus 200 using the programs and data stored in the memory, and realizes the processing of each functional unit other than the control unit 201.
[0017] Note that the control unit 201 includes one or more dedicated hardware, and at least a part of the processing by the control unit 201 may be executed by the dedicated hardware. The dedicated hardware is, for example, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), and a DSP (Digital Signal Processor), etc.
[0018] The data acquisition unit 202 acquires information on the virtual environment and information on the 3D object to be placed (synthesized in the three-dimensional space) in the virtual environment from the memory possessed by the control unit 201 or the external I / F. The data acquisition unit 202 can acquire, for example, information on the virtual environment and the 3D object selected by the user at the user interface 300 from a plurality of virtual environments and 3D objects stored in the memory. The data acquisition unit 202 can also acquire information on the 3D object generated from the image captured by the imaging device 400 and its depth information from the imaging device 400.
[0019] The image generation unit 203 places a 3D object in the virtual environment and generates a virtual camera viewpoint image of the virtual space as seen from the virtual camera viewpoint. The image generation unit 203 places the 3D object in the virtual environment based on the position and size of the 3D object in the virtual environment determined by the object placement unit 205. That is, the image generation unit 203 synthesizes the 3D object in the virtual environment, which is a three-dimensional space, based on the position and size of the 3D object in the virtual environment determined by the object placement unit 205.
[0020] The region acquisition unit 204 acquires information about non-rendered regions that are not included in the image field and therefore not rendered when the image generation unit 203 generates the virtual camera viewpoint image. The region acquisition unit 204 also acquires information about excluded regions, which are areas of the 3D object that are not displayed in the virtual camera viewpoint image for some reason, such as parts of the object that are cut off during shooting. The region acquisition unit 204 also sets the area of interest of the 3D object and sets predetermined areas that are candidates for placing the 3D object in the virtual environment. The object placement unit 205 determines the position and size of the 3D object when placing it in the virtual environment.
[0021] The functions of the data acquisition unit 202, image generation unit 203, area acquisition unit 204, and object placement unit 205 are realized by the control unit 201 executing programs corresponding to each function unit stored in ROM. The control unit 201 temporarily stores data provided from the outside via the communication interface and data used for various calculations in its own RAM (Random Access Memory), and uses the RAM as a working area to realize the processing of each function unit.
[0022] The display unit 301 of the user interface 300 is, for example, a liquid crystal display. The display unit 301 displays a graphical user interface (GUI) for the user to operate the image processing device 200. The operation unit 302 includes, for example, a keyboard, mouse, and joystick. The operation unit 302 may also be configured as an integral part of the display unit 301, or it may be, for example, a touch panel attached to the liquid crystal display. The operation unit 302 inputs various instructions based on user operations to the control unit 201 of the image processing device 200.
[0023] (Creating 3D objects) Refer to Figures 2 and 3 to illustrate an example of how to generate a 3D object. This figure illustrates the configuration of the imaging device 400. The imaging device 400 includes an imaging unit 401, an object detection unit 402, and a face detection unit 403. The imaging device 400 can communicate with the image generation device 100, the image processing device 200, and the user interface 300 via a communication interface (not shown). Figure 3 is a diagram illustrating the configuration of the imaging unit 401. Note that the imaging device 400 may be configured as an integral part of the image generation device 100 or the image processing device 200.
[0024] The imaging unit 401 can capture images of subjects in real space and generate 3D objects. The 3D objects generated by the imaging unit 401 are stored in a storage unit such as a recording medium. The imaging unit 401 can transmit information about the 3D objects to the image processing device 200 and the user interface 300.
[0025] In imaging for 3D object generation by the imaging unit 401, it is preferable to acquire distance distribution information within the field of view of the imaging unit 401 and color distribution information viewed from the same viewpoint. An imaging unit 401 capable of acquiring color distribution information is, for example, a stereo camera equipped with two imaging systems, each having an optical system and an image sensor capable of acquiring RGB images. Alternatively, an imaging unit 401 capable of acquiring color distribution information may be a ToF camera equipped with a ToF module capable of acquiring distance distribution information and an imaging system capable of acquiring RGB images. Below, an imaging unit 401 capable of acquiring distance distribution information and color distribution information using an imaging plane phase difference distance measurement method will be described.
[0026] In the example shown in Figure 3, the imaging unit 401 includes an imaging optical system 501 and an image sensor 502. The imaging unit 401 captures light from a subject located on the object surface using the imaging optical system 501 and images the image by exposing the image sensor 502, which is positioned approximately conjugate to the object surface optically.
[0027] The image sensor 502 has a structure in which pixels 503, each containing a photoelectric conversion element, are arranged in a grid of, for example, tens of millions of pixels 503. Each pixel 503 is equipped with a color filter that transmits a specific wavelength of red, green, or blue. The image sensor 502 can acquire color distribution information by arranging each pixel 503, for example, in a Bayer array.
[0028] Each pixel 503 comprises a microlens 504, a first photoelectric conversion unit 505, and a second photoelectric conversion unit 506. The first photoelectric conversion unit 505 and the second photoelectric conversion unit 506, which are arranged in the horizontal direction (X direction), acquire different light information according to their positions. From the light reception information of each pixel 503, a first image is obtained, which is composed of the brightness distribution of the light received by each first photoelectric conversion unit 505, and a second image is obtained, which is composed of the brightness distribution of the light received by each second photoelectric conversion unit 506.
[0029] The microlens 504 is designed such that the incident surface of the image sensor 502 and the light-receiving surfaces of the first photoelectric conversion unit 505 and the second photoelectric conversion unit 506 are approximately Fourier conjugate. The light-receiving surfaces of the first photoelectric conversion unit 505 and the second photoelectric conversion unit 506 and the exit pupil of the imaging optical system 501 are optically approximately conjugate. The position distribution in the exit pupil corresponds to the position distribution on the light-receiving surfaces of the first photoelectric conversion unit 505 and the second photoelectric conversion unit 506. By providing two different photoelectric conversion units, it is possible to separate and receive light beams that have passed through different pupil regions. The first image, which is composed of the luminance distribution of light received by each first photoelectric conversion unit 505, and the second image, which is composed of the luminance distribution of light received by each second photoelectric conversion unit 506, are luminance distribution information obtained by light beams that have passed through different pupil regions.
[0030] Ideally, light rays that form an image on the imaging plane enter the same point on the image sensor 502 regardless of their position over the pupil. However, the position at which defocused light rays enter the image sensor 502 changes depending on their position over the pupil. In other words, an image shift occurs depending on the amount of defocusing.
[0031] Image displacement can be calculated, for example, by stereo matching between the first and second images. Specifically, the imaging unit 401 can calculate the image displacement by matching a small patch region of one image along the direction of the epipolar lines in the other image and identifying the position with the highest correlation. The imaging unit 401 then generates a 3D object by converting the captured data to a world coordinate system using the calculated image displacement, the focal length obtained from the imaging information, and the focus position.
[0032] The object detection unit 402 groups the data captured by the imaging unit 401 by object (by type of subject). The object detection unit 402 detects the area of each grouped object as the subject area. If the detected object is a person, the face detection unit 403 detects the facial feature points and obtains the coordinates of each facial organ. Information such as the type of subject detected by the object detection unit 402 and the coordinates of each organ detected by the face detection unit 403 are linked to the 3D object generated by the imaging unit 401 and stored in a storage unit such as a recording medium.
[0033] (Missing area of a 3D object) Refer to Figures 4(A) to 4(E) to explain the missing regions of a 3D object. Figure 4(A) shows a person 600 being photographed by the imaging device 400. Figure 4(B) shows a 2D image of person 600 taken in Figure 4(A). For example, when generating a 3D object from the 2D image shown in Figure 4(B) using the method described above, the 3D object corresponding to the area below the boundary 610 (towards the feet of person 600), which is outside the field of view of the imaging device 400, will not be generated. In this way, a region in a 3D object where a missing part of the object occurs for some reason (a region where a missing part is assumed to occur and is not displayed) will be called a missing region. Figure 4(C) shows a 3D object 650 of person 600 generated from the 2D image shown in Figure 4(B). The 3D object 650 of person 600 is generated including the missing region 651.
[0034] Figure 4(D) shows the state in which the 3D object 650 is placed in the virtual environment 700. By placing the 3D object 650 in the virtual environment 700, the image processing device 200 can generate images that enhance user immersion and provide a higher quality viewing experience.
[0035] The virtual environment 700 shown in Figure 4(D) contains foreground objects 721 and background objects 722-724. The image processing device 200 can generate a virtual camera viewpoint image as seen from the virtual camera viewpoint 710.
[0036] The image processing device 200 may move the virtual camera viewpoint 710 along the virtual camera viewpoint trajectory 711. By moving the virtual camera viewpoint 710 along the virtual camera viewpoint trajectory 711, the image processing device 200 can generate moving images from the virtual camera viewpoint 710. Alternatively, the image processing device 200 may move the virtual camera viewpoint 710 within the virtual camera viewpoint region 712. By moving the virtual camera viewpoint 710 within the virtual camera viewpoint region 712, the image processing device 200 can generate images from any virtual camera viewpoint 710.
[0037] Figure 4(E) shows a virtual camera view image when a 3D object 650 is placed in the virtual environment 700 without considering the missing region 651. The virtual camera view image shown in Figure 4(E) shows the missing region 651 of the 3D object 650, which may cause discomfort to the user. Therefore, in this embodiment, the image processing device 200 determines the position and size of the 3D object to be placed in the virtual environment 700 such that the missing region 651 of the 3D object 650 is partially or completely hidden in the virtual camera view image.
[0038] (Process for determining the placement information of 3D objects) Figure 5 is a flowchart illustrating the process for determining the placement information of a 3D object. The placement information includes information on the position and size of the 3D object when it is placed in the virtual environment. The image processing device 200 determines at least one of the position and size of the 3D object to be placed in the virtual environment based on non-rendering areas that are not rendered when generating a virtual camera viewpoint image, and excluded areas of the 3D object that are not to be displayed in the virtual camera viewpoint image, within the three-dimensional space that constitutes the virtual environment. In the embodiment described below, the process of determining both the position and size of the 3D object based on information on non-rendering areas and excluded areas will be explained, but determining only one of them may reduce the area in the virtual camera viewpoint image where the excluded area is displayed.
[0039] The placement information determination process shown in Figure 5 is realized when the control unit 201, which receives an instruction from the user interface 300 to generate a virtual camera viewpoint image or change the viewpoint, executes a program corresponding to the processing of each functional unit. The placement information determination process is started, for example, when the user instructs the image processing device 200 to generate a virtual camera viewpoint image or change the viewpoint via the user interface 300.
[0040] Users can register 3D objects to be placed in a virtual environment using a dedicated application that accepts various operations from the user. From the registered 3D objects and multiple pre-registered virtual environments, users can select a desired virtual environment and the 3D objects to be placed in that virtual environment, and then instruct the generation of a virtual camera viewpoint image.
[0041] In step S101, the data acquisition unit 202 acquires information about the virtual environment and information about the 3D objects to be placed in the virtual environment. The data acquisition unit 202 displays a dialog box on the user interface 300, for example, that allows the user to select 3D objects and virtual environments. The data acquisition unit 202 only needs to acquire information about the virtual environment and 3D objects selected by the user on the user interface 300.
[0042] Furthermore, if the image processing device 200 has a function to generate 3D objects from captured images, the data acquisition unit 202 may acquire the captured image used to generate the 3D object from the user interface 300 as information about the 3D object. The captured image held by the user interface 300 is acquired, for example, from the imaging device 400. The user selects the captured image of the subject, rather than the 3D object of the subject, in the user interface 300. The data acquisition unit 202 can then generate the 3D object from the captured image selected by the user.
[0043] In the following description, 3D objects are generated by the imaging device 400. Furthermore, a 3D object is defined as a 3D representation of a region detected as a person by the object detection unit 402 of the imaging device 400. The face detection unit 403 detects facial features from the 3D object of the person, obtains the coordinates of the facial features, and records them in the storage unit of the imaging device 400.
[0044] In step S102, the region acquisition unit 204 acquires information about non-rendered areas within the virtual environment relative to the viewpoint and field of view of the virtual camera set via the user interface 300. Based on the virtual environment information acquired in step S101, the region acquisition unit 204 acquires information about non-rendered areas, which are areas where rendering is not performed, based on the positional relationship between the virtual camera viewpoint and each object in the virtual environment.
[0045] The virtual camera's viewpoint and the patterns of viewpoint movement (trajectory) are predetermined and non-rendered If the rendering region information is pre-calculated and included in the virtual environment information, the region acquisition unit 204 only needs to acquire the non-rendering region information included in the virtual environment information. If the region acquisition unit 204 acquires the virtual environment information and determines that the non-rendering region information is not included in the virtual environment information, the region acquisition unit 204 may be configured to acquire the non-rendering region information from the relationship between the virtual camera viewpoint and the foreground and background objects.
[0046] Refer to Figure 6(A) to explain the non-rendering areas. In the virtual environment 700, the non-rendering areas 821-824 are areas that, from the perspective of the virtual camera viewpoint 710, are occluded by foreground object 721 and background objects 722-724, respectively, and therefore do not need to be rendered. Areas outside the field of view (outside the angle of view) of the virtual camera viewpoint 710 are also non-rendering areas.
[0047] In step S103, the region acquisition unit 204 acquires information about areas of the 3D object that are to be excluded from display. If information about the excluded areas is attached to the 3D object as metadata, the region acquisition unit 204 may acquire the information about the excluded areas from the metadata. Alternatively, the excluded areas of the 3D object may be specified by the user via the operation unit 302 of the user interface 300. In this case, the region acquisition unit 204 acquires information about the excluded areas specified by the user. Furthermore, if the 3D object is an object that is pre-stored in a database and can be detected as a specific object, such as a 3D object of a person, the system may determine and identify where the missing part of the object is based on the information of the parts detected within the 3D object. For example, in the 3D object of a person shown in Figure 6(A), the face and torso are detected, and the lower half of the body or legs below the torso are missing, so the lower part of the torso is set as the missing area.
[0048] The excluded area is, for example, a region of a 3D object that the user does not want to display, and can be a missing region 651 that includes the boundary of the 3D object. Other excluded areas may be set to include blurred areas and occlusion areas determined from image analysis or imaging conditions.
[0049] In step S104, the object placement unit 205 determines the position and size of the 3D object within the virtual environment based on the non-rendering area and the exclusion area. For example, the object placement unit 205 determines the position and size of the 3D object such that at least a portion of the exclusion area overlaps with the non-rendering area.
[0050] The object placement unit 205 may evaluate the placement information (position and size of the 3D object) using an evaluation function that represents the degree of agreement between the non-target area of the 3D object and the non-rendering area. The degree of agreement between the non-target area and the non-rendering area is an indicator that shows how much of the non-target area overlaps with the non-rendering area. The object placement unit 205 uses various optimization methods or the Monte Carlo method to change at least one of the position and size of the 3D object from the initial position to increase the degree of agreement, thereby determining the final position and size of the 3D object. By determining the position and size of the 3D object to increase the degree of agreement between the non-target area and the non-rendering area, the object placement unit 205 can adjust the placement of the 3D object so that the non-target area is hidden by the non-rendering area.
[0051] By determining the optimal position and size of the 3D object, the object placement unit 205 can place part or all of the missing area 651, which is an area outside the scope of the 3D object 650, within the non-rendering area 821 of the virtual environment 700, as shown in Figure 6(A).
[0052] The image generation unit 203 generates a virtual camera viewpoint image in which the 3D object 650 is placed in the virtual environment 700 based on the position and size of the 3D object 650 determined by the object placement unit 205. Since the missing region 651, which is an area outside the scope of the 3D object 650, is located within the non-rendering region 821, the image generation unit 203 can generate a virtual camera viewpoint image that does not show the missing region 651 and does not appear unnatural, as shown in Figure 6(B).
[0053] As shown in Figure 5, the placement information determination process allows the image processing device 200 to determine the optimal position and size of a 3D object when placing it in the virtual environment, based on the non-rendering area and the exclusion area.
[0054] (Area within a virtual environment where 3D objects are placed) Referring to Figures 7(A), 7(B) and 8(A), 8(B), the areas in which 3D objects are placed in the virtual environment 700 will be described. The image processing device 200 can generate a suitable virtual camera viewpoint image by placing 3D objects in predetermined areas in the virtual environment 700, such as an area located in the center when viewed from the virtual camera viewpoint 710, or an area that does not overlap with other objects.
[0055] In Figures 7(A) and 8(A), the area behind the foreground object 721 and background objects 722-724 relative to the virtual camera viewpoint 710 becomes the non-rendered area 821-824. Additionally, areas outside the field of view of the virtual camera viewpoint 710 also become non-rendered areas. Based on the information of the non-rendered areas, the image processing device 200 can generate a virtual camera viewpoint image with reduced visual inconsistencies by omitting some or all of the unwanted areas, such as the missing area 651.
[0056] The image processing device 200 may set a predetermined area in the virtual environment that is suitable for the placement of the area of interest of the 3D object that is the subject (hereinafter referred to as the area of interest of the 3D object) in order to create a more natural image.
[0057] In Figure 7(A), a predetermined region 860 of the virtual environment 700 is set to the central part of the field of view as seen from the virtual camera viewpoint 710. The predetermined region 860 may be pre-set, for example, as a region for placing the area of interest of a 3D object when creating the virtual environment 700. Alternatively, the predetermined region 860 may be set based on the field of view from the virtual camera viewpoint 710 when capturing the virtual environment 700. For example, the predetermined region 860 is set to the region that includes the center of the field of view from the virtual camera viewpoint 710. Figure 7(B) shows an image as seen from the virtual camera viewpoint 710.
[0058] The object placement unit 205 of the image processing device 200 can determine the position and size of a 3D object based on a predetermined area 860 of the virtual environment 700 and the area of interest of the 3D object. The object placement unit 205 determines the position and size of the 3D object such that at least a portion of the area of interest of the 3D object overlaps with the predetermined area 860 of the virtual environment 700.
[0059] Furthermore, the area of focus of a 3D object may be the entire subject. When the entire subject is designated as the area of focus of a 3D object, the image processing device 200 sets a predetermined area in the virtual environment that is suitable for arranging the entire 3D object.
[0060] In Figure 8(A), a predetermined area 870 of the virtual environment 700 is set between background object 722 and background object 723 within the field of view as seen from the virtual camera viewpoint 710. The predetermined area 870 may be pre-set, for example, as an area for arranging the entire 3D object when creating the virtual environment 700. Alternatively, the predetermined area 870 may be set based on the field of view from the virtual camera viewpoint 710 when capturing the virtual environment 700. The predetermined area 870 may also be set based on the positional relationship with the foreground object 721 and background objects 722-724. For example, the predetermined area 870 is set to an area in the central part of the field of view from the virtual camera viewpoint 710 that does not overlap with the background objects 722 and 723. Figure 8(B) shows an image as seen from the virtual camera viewpoint 710.
[0061] Multiple predetermined areas may be set in the virtual environment 700 depending on the subject to be placed. By setting predetermined areas 860 and 870 in the virtual environment 700, the image processing device 200 can position the area of interest of the 3D object in a more suitable position and size, and generate a natural virtual camera viewpoint image.
[0062] (Area of focus of a 3D object) Figures 9(A) to 9(C) are diagrams illustrating the areas of interest of a 3D object. Figure 9(A) shows the captured image used to generate the 3D object 650. The object detection unit 402 detects the area of the person 600 as the subject area from the captured image. The face detection unit 403 detects the feature points of the person 600's face and obtains the coordinates of each organ of the face 620. Note that the processing of the object detection unit 402 and the face detection unit 403 may be performed by the image processing device 200. When the data acquisition unit 202 acquires the captured image from the imaging device 400, the area acquisition unit 204 can detect the area of the person 600 from the captured image and obtain the coordinates of each organ of the face 620.
[0063] Since the region acquisition unit 204 recognizes that the region of the person 600 is cut off by the boundary 610 of the captured image, it can acquire the missing region 651 corresponding to the boundary 610 in the 3D object 650 as an excluded region. Multiple excluded regions may be acquired.
[0064] Furthermore, the region acquisition unit 204 acquires the region in the 3D object 650 corresponding to the face 620 as the region of interest 660 of the 3D object. In addition, the region acquisition unit 204 acquires the entire 3D object 650, which is the subject region, as the region of interest 670. The region acquisition unit 204 may acquire multiple regions of interest from the 3D object depending on the purpose.
[0065] As shown in Figures 9(B) and 9(C), the region acquisition unit 204 can acquire the missing region 651 (an area not to be targeted), the area of interest 660 of the 3D object, and the area of interest 670 of the 3D object from the 3D object 650. The region acquisition unit 204 may also acquire the areas not to be targeted and the areas of interest by object detection and segmentation using various machine learning methods. In addition, the region acquisition unit 204 may acquire areas specified by the user via the user interface 300 as areas not to be targeted or areas of interest.
[0066] The image processing device 200 can position the 3D object 650 so that the missing region 651 overlaps with the non-rendering region, based on information about the missing region 651, which is an area outside the target area. The image processing device 200 can position the 3D object in the appropriate position and size by acquiring the areas of interest 660 and 670 of the 3D object 650 and positioning them so that they overlap with a predetermined area of the virtual environment 700.
[0067] (Priority of position and size when placing 3D objects) The object placement unit 205 of the image processing device 200 may determine the priority of the position and size of 3D objects before determining the placement of the 3D objects. The placement unit 205 weights the amount of change in the position and size of 3D objects based on the priority of the position and size of the 3D objects. The object placement unit 205 can adjust whether to prioritize position or size when placing 3D objects by setting the priority of the position and size of the 3D objects.
[0068] The object placement unit 205 can set the priority of the position and size of 3D objects based on the style of the background image of the virtual environment. For example, if the background image of the virtual environment acquired by the data acquisition unit 202 in step S101 is toon-style and scale is not important, the object placement unit 205 sets the priority of size higher than the priority of position. By setting the priority of size higher than the priority of position, the image generation unit 203 can generate a powerful virtual camera viewpoint image in which the subject is depicted large.
[0069] On the other hand, if the background image of the virtual environment is realistic and a sense of scale is important, the object placement unit 205 sets the priority of size lower than the priority of position. By setting the priority of size lower than the priority of position, the image generation unit 203 can generate a virtual camera viewpoint image that looks natural without changing the current dimensions of the subject as much as possible. To further emphasize the sense of scale, it is also possible to adjust only the position without changing the size from the default size set for the 3D object or the size desired by the user.
[0070] Furthermore, the object placement unit 205 may set the priority of the position and size of 3D objects based on the relative size relationship between a predetermined area of the virtual environment and the area of interest of the 3D object.
[0071] The object placement unit 205 sets the priority of the 3D object's size higher than the priority of the 3D object's position if the area of interest of the 3D object is smaller than a predetermined area of the virtual environment. If the areas of interest of the 3D object 660 and 670 are smaller than the predetermined areas 860 and 870 of the virtual environment, respectively, the priority of the 3D object's size is set higher than the priority of the 3D object's position.
[0072] The object placement unit 205 sets the priority of the 3D object's size lower than the priority of the 3D object's position if the area of interest of the 3D object is larger than a predetermined area of the virtual environment. If the areas of interest of the 3D object 660 and 670 are larger than the predetermined areas 860 and 870 of the virtual environment, respectively, the priority of the 3D object's size is set lower than the priority of the 3D object's position.
[0073] Furthermore, the object placement unit 205 may set the priority of the position and size of 3D objects based on the relationship between the distance from a predetermined area of the virtual environment to a non-rendering area and the distance from the area of interest of the 3D object to an area outside the area of interest. The distance between each area may be the shortest distance between each area, or the distance between the centroids of each area.
[0074] The object placement unit 205 sets the priority of the 3D object's size higher than the priority of the 3D object's position if the distance from the 3D object's area of interest to the area outside the target region is shorter than the distance from a predetermined area of the virtual environment to the non-rendering region. Similarly, if the distance from the 3D object's area of interest 660 to the missing area 651 of the area outside the target region is shorter than the distance from the predetermined area 860 of the virtual environment to the non-rendering region, the priority of the 3D object's size is set higher than the priority of the 3D object's position.
[0075] The object placement unit 205 sets the priority of the 3D object's size lower than the priority of the 3D object's position if the distance from the area of interest of the 3D object to the area outside the area of interest is longer than the distance from a predetermined area of the virtual environment to the non-rendering area. If the distance from the 3D object's area of interest 660 to the missing area 651 of the excluded area is longer than the distance from the predetermined area 860 of the rendering environment to the non-rendering area, the priority of the 3D object's size is set lower than the priority of the 3D object's position.
[0076] The object placement unit 205 may change both the size and position of the 3D objects to be placed in the virtual environment. The object placement unit 205 adjusts the amount of change to the size and position of the 3D objects to be placed in the virtual environment based on the priority of size and position. By appropriately adjusting the amount of change to size and position, the image generation unit 203 can generate a virtual camera viewpoint image that looks more natural.
[0077] The priority of position and size when placing 3D objects in the virtual environment may be automatically set by the object placement unit 205 at the instruction of the control unit 201. Alternatively, the priority of position and size of 3D objects can be changed by the user via the operation unit 302 of the user interface 300. By providing a GUI for changing the priority of position and size, the user can adjust the priority values automatically set by the object placement unit 205. The image generation unit 203 can generate virtual camera viewpoint images that more closely match the user's intentions and feel natural by accepting priority input according to the user's preferences.
[0078] Based on the set priority, the object placement unit 205 determines the position and size of the 3D object when placing it in the virtual environment. In step S104 of Figure 5, the object placement unit 205 changes the position of the 3D object more significantly and the size less significantly the higher the position priority. By expressing the degree of agreement between the non-rendered area of the 3D object and the non-target area using an evaluation function, the object placement unit 205 can determine the position and size of the 3D object in the virtual environment in a way that increases the degree of agreement.
[0079] The evaluation function includes, for example, terms relating to the amount of change in position and the amount of change in size. The terms relating to the amount of change in position and the amount of change in size have position priority and size priority as coefficients, respectively. When the position and size priorities are set, the object placement unit 205 adjusts the amount of change in position and size of the 3D object so that the degree of agreement with the non-rendered area of the excluded area is higher. Based on the amount of change in position and size of the 3D object obtained using the evaluation function, the object placement unit 205 can determine the position and size of the 3D object in the virtual environment.
[0080] (How to place 3D objects) This section describes a method for arranging 3D objects using an evaluation function. The object placement unit 205 determines the position and size of the 3D object in the virtual environment based on the non-rendering area and the missing area 651, which is an area outside the rendering scope. The object placement unit 205 evaluates the placement of the 3D object using an evaluation function that represents the degree of agreement between the 3D object's area outside the rendering scope and the non-rendering scope. The object placement unit 205 can hide part or all of the 3D object's area outside the rendering scope in the non-rendering scope by adjusting the position and size of the 3D object so that the area outside the rendering scope overlaps with the non-rendering scope.
[0081] Furthermore, the object placement unit 205 may evaluate the degree of agreement between the non-rendering area and the area other than the area outside the target region of the 3D object (missing area 651). The object placement unit 205 determines the position and size of the 3D object so that the area other than the area outside the target region of the 3D object is hidden as little as possible by the non-rendering area. In other words, the object placement unit 205 evaluates the degree of agreement between the non-rendering area and the area other than the area outside the target region of the 3D object The position and size of the 3D object should be determined in such a way that the degree of agreement with the region is reduced.
[0082] Furthermore, the object placement unit 205 may evaluate the degree of agreement between a predetermined area of the virtual environment and the area of interest of the 3D object.
[0083] For example, if the region acquisition unit 204 acquires a predetermined region 860 of the virtual environment and a region of interest 660 of a 3D object, a term representing the degree of agreement between the predetermined region 860 and the region of interest 660 is added to the evaluation function. The object placement unit 205 should then determine the position and size of the 3D object so that the degree of agreement between the predetermined region 860 and the region of interest 660 increases. By arranging the 3D object so that the predetermined region 860 and the region of interest 660 coincide, the image generation unit 203 can generate a virtual camera viewpoint image that zooms in on the region of interest of the 3D object.
[0084] Furthermore, when the region acquisition unit 204 acquires a predetermined region 870 of the virtual environment and a region of interest 670 of the 3D object, a term representing the degree of agreement between the predetermined region 870 and the region of interest 670 is added to the evaluation function. The object placement unit 205 should then determine the position and size of the 3D object so that the degree of agreement between the predetermined region 870 and the region of interest 670 increases. By arranging the 3D object so that the predetermined region 870 and the region of interest 670 coincide, the image generation unit 203 can generate a natural virtual camera viewpoint image without disrupting the overall arrangement of the 3D object. The method of evaluating the arrangement of the entire 3D object (region of interest 670) to determine the position and size of the 3D object is effective when face detection fails and the region of interest 660 is not acquired.
[0085] (A simple method for placing 3D objects) The object placement unit 205 is not limited to a method of placing 3D objects using an evaluation function; it is also possible to place 3D objects using a simple and fast method without using an evaluation function.
[0086] First, the object placement unit 205 determines the size of the 3D object based on the distance between a predetermined area 860 of the virtual environment 700 and a non-rendering area in the virtual camera viewpoint 710. The object placement unit 205 determines the size of the 3D object such that the distance between the area of interest 660 of the 3D object and the area outside of it (missing area 651) is greater than the distance between the predetermined area 860 of the virtual environment 700 and a non-rendering area. Next, the object placement unit 205 determines the position of the 3D object such that at least a portion of the predetermined area 860 of the virtual environment and the area of interest 660 of the 3D data overlap.
[0087] The object placement unit 205 can determine the position and size of 3D objects in the virtual environment based on non-rendering areas and excluded areas in a fast and simple manner, without using an evaluation function.
[0088] In the above embodiment, the image processing device 200 determines the position and size of a 3D object when placing it in the virtual environment, based on the non-rendering area of the virtual environment and the area outside the scope of the 3D object. The image processing device 200 can reduce the unnatural appearance of the virtual camera viewpoint image by placing the 3D object, including the area outside the scope, in the virtual environment (compositing it with the background of the 3D space) such that the area outside the scope overlaps with the non-rendering area.
[0089] A simpler method involves determining the non-rendering region based on the position of each component in the virtual environment relative to the virtual camera viewpoint 710, and then defining the non-rendering region of the 3D object within that region. The position of the 3D object may be changed and determined only so that part or all of the area overlaps. Of course, in this case as well, the size of the 3D object may be changed for other reasons (for example, so that the user can view the 3D object at a desired size).
[0090] The various controls described above may or may not be performed by a single piece of hardware (e.g., a processor or circuit). Multiple pieces of hardware (e.g., multiple processors, multiple circuits, or a combination of one or more processors and one or more circuits) may share the processing to control the entire device.
[0091] Furthermore, the above-mentioned processors are processors in a broad sense, including general-purpose processors and specialized processors. General-purpose processors include, for example, CPUs (Central Processing Units), MPUs (Micro Processing Units), and DSPs (Digital Signal Processors). Specialized processors include, for example, GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), and PLDs (Programmable Logic Devices). Programmable logic devices include, for example, FPGAs (Field Programmable Gate Arrays) and CPLDs (Complex Programmable Logic Devices).
[0092] Furthermore, the embodiments described above (including modified examples) are merely examples, and configurations obtained by appropriately modifying or changing the above-described configurations within the scope of the gist of the present invention are also included in the present invention. Configurations obtained by appropriately combining the above-described configurations are also included in the present invention.
[0093] <Other Embodiments> The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit that implements one or more functions.
[0094] This embodiment includes the following configurations, methods, programs, and media. (Composition 1) Acquisition means for acquiring information about a virtual environment and information about 3D objects placed in the virtual environment, The system includes a determination means for determining the position and size of the three-dimensional object when it is placed in the virtual environment. A 3D object placed in the virtual environment at the position and size determined by the determination means overlaps with a part or all of the non-rendering area of the 3D object that is not to be displayed in the non-rendering area of the virtual environment. An image processing apparatus characterized by the following: (Configuration 2) The determination means determines the position of the 3D object in the virtual environment based on the non-rendering area and the excluded area. The image processing apparatus according to configuration 1, characterized in that... (Composition 3) The determination means determines the size of the 3D object in the virtual environment based on the non-rendering area and the excluded area. The image processing apparatus according to configuration 1, characterized in that... (Composition 4) The determination means determines the virtual ring based on the non-rendering region and the excluded region. Determine the position and size of the 3D object at the boundary. The image processing apparatus according to configuration 1, characterized in that... (Composition 5) The determination means determines the position and size of the 3D object based on a predetermined area of the virtual environment and the area of interest of the 3D object. An image processing apparatus according to any one of configurations 1 to 4, characterized by the above. (Composition 6) The determination means determines the position and size of the 3D object such that at least a portion of the area of interest of the 3D object overlaps with a predetermined area of the virtual environment. An image processing apparatus according to any one of configurations 1 to 5, characterized by the above. (Composition 7) The image processing apparatus according to any one of configurations 1 to 6, characterized in that the determination means determines the size of the 3D object such that the distance between the area of interest of the 3D object and the area not to be targeted is greater than the distance between a predetermined area of the virtual environment and the non-rendering area, and determines the position of the 3D object such that at least a portion of the predetermined area of the virtual environment and the area of interest of the 3D object overlap. (Composition 8) The predetermined area of the virtual environment is set based on at least one of the following: the field of view from a virtual camera viewpoint when photographing the virtual environment, and the positional relationship with other objects placed in the virtual environment. An image processing apparatus according to any one of configurations 5 to 7, characterized by the above. (Composition 9) The excluded region is the region that includes the boundary of the three-dimensional object. An image processing apparatus according to any one of configurations 1 to 8, characterized by the above. (Composition 10) The determination means weights the respective amounts of change to the position and size of the 3D object based on the priority of the position and size of the 3D object. An image processing apparatus according to any one of configurations 1 to 9, characterized by the above. (Composition 11) The priority of the position and size of the three-dimensional object is set based on the style of the background image of the virtual environment. The image processing apparatus according to configuration 10, characterized in that... (Composition 12) If the area of interest of the 3D object is smaller than a predetermined area of the virtual environment, the priority of the size of the 3D object is set higher than the priority of the position of the 3D object. An image processing apparatus according to configuration 10 or 11, characterized by the above. (Composition 13) If the area of interest of the 3D object is larger than a predetermined area of the virtual environment, the priority of the size of the 3D object is set lower than the priority of the position of the 3D object. An image processing apparatus according to any one of configurations 10 to 12, characterized by the above. (Composition 14) The image processing apparatus according to configuration 10 or 11, characterized in that, if the distance from the area of interest of the 3D object to the area outside the target is shorter than the distance from a predetermined area of the virtual environment to the non-rendering area, the priority of the size of the 3D object is set higher than the priority of the position of the 3D object. (Composition 15) If the distance from the area of interest of the 3D object to the area outside the target is longer than the distance from the predetermined area of the virtual environment to the non-rendering area, the 3D object The image processing apparatus according to configuration 10, 11, or 14, characterized in that the priority of the size of the object is set lower than the priority of the position of the three-dimensional object. (Composition 16) The priority of the position and size of the aforementioned 3D object can be changed by the user. An image processing apparatus according to any one of configurations 10 to 15, characterized by the above. (Composition 17) The determination means determines the position and size of the 3D object using an evaluation function that evaluates the degree of agreement between the excluded region and the non-rendered region. An image processing apparatus according to any one of configurations 1 to 16, characterized by the above. (Composition 18) The system further includes a generation means that generates an image of the 3D object placed in the virtual environment based on the position and size of the 3D object determined by the determination means. An image processing apparatus according to any one of configurations 1 to 17, characterized by the above. (method) An acquisition step to acquire information about the virtual environment and information about the 3D objects to be placed in the virtual environment, The process includes a determination step of determining the position and size of the 3D object when placing the 3D object in the virtual environment, A 3D object placed in the virtual environment at the position and size determined in the aforementioned determination step has part or all of its non-rendered area overlapping with the non-rendered area of the virtual environment. A control method for an image processing apparatus, characterized by the features described above. (program) A program for causing a computer to function as one of the means of an image processing apparatus described in any of configurations 1 to 18. (medium) A computer-readable storage medium containing a program for causing the computer to function as one of the means of the image processing apparatus described in any of configurations 1 to 18. [Explanation of symbols]
[0095] 200: Image processing device, 201: Control unit, 202: Data acquisition unit, 204: Area acquisition unit, 205: Object placement unit
Claims
1. Acquisition means for acquiring information about a virtual environment and information about three-dimensional objects placed in the virtual environment, The system includes a determination means for determining the position and size of the three-dimensional object when it is placed in the virtual environment. A 3D object placed in the virtual environment at the position and size determined by the determination means overlaps with a part or all of the non-rendering area of the 3D object that is not to be displayed in the non-rendering area of the virtual environment. An image processing apparatus characterized by the following:
2. The determination means determines the position of the 3D object in the virtual environment based on the non-rendering region and the excluded region. The image processing apparatus according to feature 1.
3. The determination means determines the size of the 3D object in the virtual environment based on the non-rendering area and the excluded area. The image processing apparatus according to feature 1.
4. The determination means determines the position and size of the 3D object in the virtual environment based on the non-rendering area and the excluded area. The image processing apparatus according to feature 1.
5. The determination means determines the position and size of the 3D object based on a predetermined area of the virtual environment and the area of interest of the 3D object. The image processing apparatus according to feature 1.
6. The determination means determines the position and size of the 3D object such that at least a portion of the area of interest of the 3D object overlaps with a predetermined area of the virtual environment. The image processing apparatus according to feature 1.
7. The image processing apparatus according to claim 1, characterized in that the determination means determines the size of the three-dimensional object such that the distance between the area of interest of the three-dimensional object and the area not to be targeted is greater than the distance between a predetermined area of the virtual environment and the non-rendering area, and determines the position of the three-dimensional object such that at least a portion of the predetermined area of the virtual environment and the area of interest of the three-dimensional object overlap.
8. The predetermined area of the virtual environment is set based on at least one of the following: the field of view from a virtual camera viewpoint when photographing the virtual environment, and the positional relationship with other objects placed in the virtual environment. The image processing apparatus according to feature 5.
9. The excluded region is the region that includes the boundary of the three-dimensional object. The image processing apparatus according to feature 1.
10. The determination means weights the respective changes in the position and size of the three-dimensional object based on the priority of the position and size of the three-dimensional object. The image processing apparatus according to feature 1.
11. The priority of the position and size of the three-dimensional object is determined by the background image of the virtual environment. Set based on the formula The image processing apparatus according to feature 10.
12. If the area of interest of the 3D object is smaller than a predetermined area of the virtual environment, the priority of the size of the 3D object is set higher than the priority of the position of the 3D object. The image processing apparatus according to feature 10.
13. If the area of interest of the 3D object is larger than a predetermined area of the virtual environment, the priority of the size of the 3D object is set lower than the priority of the position of the 3D object. The image processing apparatus according to feature 10.
14. The image processing apparatus according to claim 10, characterized in that, if the distance from the area of interest of the three-dimensional object to the area outside the target is shorter than the distance from a predetermined area of the virtual environment to the non-rendering area, the priority of the size of the three-dimensional object is set higher than the priority of the position of the three-dimensional object.
15. The image processing apparatus according to claim 10, characterized in that, if the distance from the area of interest of the three-dimensional object to the area outside the target is longer than the distance from a predetermined area of the virtual environment to the non-rendering area, the priority of the size of the three-dimensional object is set lower than the priority of the position of the three-dimensional object.
16. The priority of the position and size of the aforementioned 3D object can be changed by the user. The image processing apparatus according to feature 10.
17. The determination means determines the position and size of the 3D object using an evaluation function that evaluates the degree of agreement between the excluded region and the non-rendered region. The image processing apparatus according to feature 1.
18. The system further includes a generation means that generates an image of the 3D object placed in the virtual environment based on the position and size of the 3D object determined by the determination means. The image processing apparatus according to feature 1.
19. An acquisition step to acquire information about the virtual environment and information about the 3D objects to be placed in the virtual environment, The system includes a determination step of determining the position and size of the three-dimensional object when placing the three-dimensional object in the virtual environment, The 3D object placed in the virtual environment at the position and size determined in the aforementioned determination step has part or all of the exclusion area of the 3D object that is not to be displayed overlapping with the non-rendering area of the virtual environment. A control method for an image processing apparatus, characterized by the features described above.
20. A program for causing a computer to function as one of the means of an image processing apparatus according to any one of claims 1 to 18.
21. A computer-readable storage medium storing a program for causing a computer to function as one of the means of an image processing apparatus according to any one of claims 1 to 18.