Image processing device, image processing method, and program

The image processing system generates high-resolution computer-generated images with varied lighting effects by using virtual light source and viewpoint information, addressing the inefficiencies of capturing multiple images under different lighting conditions.

WO2026070926A1PCT designated stage Publication Date: 2026-04-02FUJIFILM CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Generating high-resolution 3D model data requires numerous high-resolution images under various lighting conditions, which is effort-intensive and data-intensive, making it impractical to change light sources in captured images for rendering high-quality computer-generated images.

Method used

An image processing system that generates a second image by using virtual light source information and virtual viewpoint information to process a first image, allowing for the simulation of different lighting conditions without the need for additional physical photography.

Benefits of technology

Enables the creation of high-resolution computer-generated images with varied lighting effects efficiently, reducing the effort and data requirements associated with capturing multiple images under diverse lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025033727_02042026_PF_FP_ABST
    Figure JP2025033727_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an image processing device, an image processing method, and a program. This image processing device comprises a processor which: accepts three-dimensional data of an object; accepts virtual light source information relating to a virtual light source and virtual viewpoint information relating to a virtual viewpoint, with respect to the three-dimensional data; generates a CG image from the three-dimensional data on the basis of the virtual light source information and the virtual viewpoint information; and, on the basis of the virtual light source information and the virtual viewpoint information, generates a second image from at least one first image obtained by imaging the object.
Need to check novelty before this filing date? Find Prior Art

Description

Image Processing Apparatus, Image Processing Method, and Program

[0001] The present invention relates to an image processing apparatus, an image processing method, and a program.

[0002] Patent Document 1 discloses a photo search and browsing system using a three-dimensional model (3D model). This system compares virtual information regarding the position and orientation of a virtual camera corresponding to the viewpoint of a user who views the 3D model of an object on a screen with actual information regarding the position and orientation of the camera that captured a photo of the object. Based on this, the system enables the display and browsing of digital images captured from viewpoints similar to the 3D model that the user is browsing on the screen.

[0003] Japanese Unexamined Patent Application Publication No. 2006 - 309722

[0004] One embodiment of the technology according to the present disclosure provides an image processing apparatus, an image processing method, and a program that generate a CG image from the 3D data based on virtual light source information regarding a virtual light source for the 3D data and virtual viewpoint information regarding a virtual viewpoint, and generate a second image from at least one first image of an object captured based on the virtual light source information and the virtual viewpoint information.

[0005] The image processing apparatus according to the first aspect has a processor that receives 3D data of an object, receives virtual light source information regarding a virtual light source for the 3D data and virtual viewpoint information regarding a virtual viewpoint, generates a CG image from the 3D data based on the virtual light source information and the virtual viewpoint information, and generates a second image from at least one first image of the object captured based on the virtual light source information and the virtual viewpoint information.

[0006] In the image processing apparatus according to the second aspect, in the first aspect, the 3D data is stored in association with captured images respectively obtained by capturing the object from a plurality of shooting positions and a plurality of shooting directions, and when generating the CG image, the processor selects the captured images at the shooting positions and shooting directions corresponding to the virtual viewpoint information as the first images.

[0007] In the third embodiment, the image processing device, in the second embodiment, generates three-dimensional data based on captured images obtained by photographing an object from multiple shooting positions and multiple shooting directions.

[0008] The image processing apparatus according to the fourth embodiment, in any of the first to third embodiments, the three-dimensional data includes at least one of the shape information of an object, pixel value information, and material information, and the processor generates a first intermediate CG image using the shape information, virtual light source information, virtual viewpoint information, material information and pixel value information, generates a second intermediate CG image using the shape information, virtual light source information, virtual viewpoint information and pixel value information, generates a first difference image from the first intermediate CG image and the second intermediate CG image, and generates a second image from the first difference image and the first image.

[0009] The image processing device according to the fifth embodiment generates a second image from a first image based on virtual light source information relating to a virtual light source when rendering three-dimensional data, in any of the first to third embodiments.

[0010] In the sixth aspect of the image processing apparatus, in the fourth aspect, the processor performs a process to reduce the difference between the resolution of the first difference image and the resolution of the first image, obtains a second difference image, and generates a second image based on the first image and the second difference image.

[0011] In the seventh embodiment, the image processing apparatus, in the sixth embodiment, is a process that reduces the difference in resolution by bringing the image with the lower resolution among the first image and the first difference image closer to the image with the higher resolution.

[0012] In the image processing apparatus according to the eighth embodiment, in any of the first to seventh embodiments, the processor sets the color tone of the CG image based on the first image.

[0013] In the image processing apparatus according to the ninth embodiment, in any of the first to eighth embodiments, the processor sets the color tone of the second image based on the CG image.

[0014] In the image processing apparatus according to the tenth embodiment, in the fourth embodiment, the processor receives a change in the material information of the three-dimensional data and generates a second image based on the changed material information.

[0015] In the image processing apparatus according to the 11th embodiment, in any of the first to tenth embodiments, the processor sets the resolution of the second image to be generated.

[0016] In the image processing apparatus according to the 12th embodiment, in any of the first to 11 embodiments, the first image is an image captured under uniform lighting conditions during shooting, or an image captured while suppressing reflection from the object.

[0017] The image processing method according to the 13th embodiment receives 3D data of an object, receives virtual light source information and virtual viewpoint information related to a virtual light source for the 3D data, generates a CG image from the 3D data based on the virtual light source information and virtual viewpoint information, and causes the processor to perform the following processes: generate a second image from at least one first image of the object based on the virtual light source information and virtual viewpoint information.

[0018] The program according to the 14th embodiment receives 3D data of an object, receives virtual light source information and virtual viewpoint information related to a virtual light source for the 3D data, generates a CG image from the 3D data based on the virtual light source information and virtual viewpoint information, and causes the computer to perform the following processes: generate a second image from at least one first image of the object based on the virtual light source information and virtual viewpoint information.

[0019] Figure 1 is a diagram illustrating the overview of the image processing system. Figure 2 is a block diagram showing the schematic configuration of the controller. Figure 3 is a block diagram showing the processing functions realized by the processor. Figure 4 is a flowchart showing the processing in the image processing system. Figure 5 is a conceptual diagram showing the state of capturing an object with multiple cameras. Figure 6 is a functional block diagram showing the generation of 3D data. Figure 7 is a diagram showing the lighting device and virtual light source. Figure 8 is a functional block diagram showing the generation of a CG image. Figure 9 is a diagram showing the captured image and the CG image. Figure 10 is a functional block diagram explaining the generation of the second image from the first image. Figure 11 is a diagram showing an example of the display screen. Figure 12 is a diagram showing another example of the display screen. Figure 13 is a diagram showing the case where an object is captured under uniform lighting conditions. Figure 14 is a diagram showing the case where an object is captured using two polarizing filters.

[0020] Preferred embodiments of the present invention will be described below with reference to the attached drawings.

[0021] [Overview] Methods for generating 3D models are known, including photogrammetry, SfM (Structure from Motion), and MVS (Multi View Stereo). These methods involve photographing an object from multiple angles and generating a 3D model from the resulting images. The data format of the 3D model data includes, for example, data representing the shape of the object as a point cloud and its relationships, and information about pixel values ​​representing the pixel values ​​of the object's surface. Furthermore, a material map, as described later, can be added.

[0022] Material maps represent information such as light reflection intensity and surface roughness using maps. This information allows 3D models to more accurately reproduce how they appear under various lighting conditions in a virtual space, and to be freely modified.

[0023] By the way, in order to render a 3D model and obtain a high-resolution CG image, high-resolution 3D model data is required. However, generating high-resolution 3D model data requires, for example, acquiring many high-resolution still images of the object and a long 3D model generation process. Furthermore, the amount of data in the generated 3D model data becomes large, and it also takes time to render that 3D model data and generate a CG image.

[0024] On the other hand, while still images generally have higher resolution than computer-generated images, preparing a large number of images under various lighting conditions is difficult in terms of the effort required for shooting and the storage capacity (amount of data).

[0025] Therefore, the present invention provides a technology that can generate a second image when the light source is changed on a CG image by using a CG image of the object and light source information in a virtual space to generate a new second image by performing image processing on a first image of the object.

[0026] <Embodiment> Figure 1 shows an image processing system 1 including a controller 20 of an embodiment. The image processing system 1 includes a plurality of cameras 10, a lighting device 12, a controller 20, and a shooting table 40. An object 30 to be photographed is placed on the shooting table 40.

[0027] Camera 10 is a device for photographing an object 30 and acquiring an image (a still image in this example) of the object 30. As long as Camera 10 can photograph the object 30 and acquire an image, the type and structure of the lens and sensor used are not limited. Preferably, Camera 10 is equipped with an exposure adjustment unit and a white balance adjustment unit. The exposure adjustment unit and white balance adjustment unit adjust the brightness and color of the object in the image, and these adjustments may be made manually or automatically.

[0028] Exposure can be adjusted by adjusting the exposure time (shutter speed), aperture, and ISO sensitivity. When adjusting the exposure, it is preferable to increase the depth of field (use a smaller aperture) so that the entire object 30 is in focus. Also, since adjusting the ISO sensitivity changes the image quality, it is preferable for the exposure adjustment unit to adjust the exposure by adjusting the exposure time. The camera 10 may be equipped with a correction mechanism to compensate for camera shake.

[0029] The support member 44 is a member for fixing the camera 10. Multiple cameras 10 are fixed to the support member 44, and the cameras 10 are stationary. The multiple cameras 10 are arranged on the support member 44 so that each lens is pointed towards the object 30. The support member 44 consists of a mounting base 44A and an arm portion 44B. The arm portion 44B is a curved member attached to the mounting base 44A and extends in the vertical direction. Multiple cameras 10 are attached to the arm portion 44B at a predetermined height position. As a result, the multiple cameras 10 are arranged in a so-called vertical direction along the arm portion 44B. However, the arrangement of the multiple cameras 10 is not limited to the case shown in Figure 1. Also, there may be only one camera 10.

[0030] The lighting device 12 is a photographic light source that illuminates the object 30 with illumination light when photographing the object 30. The lighting device 12 is the photographic light source used when photographing the object 30. The brightness, color, position relative to the object 30, illumination range, and illumination angle of the lighting device 12 can be changed. The type of light source is not limited as long as the lighting device 12 can illuminate the object 30 with illumination light. As the lighting device 12, for example, a lighting device 12 composed of LEDs (Light-Emitting Diodes) of multiple colors (red, blue, and green, etc.) can be used. Although one lighting device 12 has been illustrated, multiple lighting devices 12 with different positions and illumination directions may be provided.

[0031] The imaging table 40 is a component for positioning the object 30. The imaging table 40 has a mounting surface 40A on its upper surface. The object 30 is positioned approximately in the center of the mounting surface 40A. The imaging table 40 is configured to rotate arbitrarily in the direction indicated by the arrow. The imaging table 40 is equipped with a motor (not shown). The motor can arbitrarily change the rotation speed, rotation angle, etc., of the imaging table 40. The imaging table 40 is a turntable. The rotation speed and rotation angle can be detected and changed during processing.

[0032] Multiple marks 40B are provided on the mounting surface 40A of the imaging table 40. The marks 40B serve as indicators of the imaging direction. The marks 40B are arranged at equal intervals (45-degree intervals) on the outer edge of the mounting surface 40A. The marks 40B may also display angles indicating the imaging direction (0 degrees, 45 degrees, 90 degrees...315 degrees). The mounting surface 40A of the imaging table 40 has a disc shape with a flat top surface. The shape of the imaging table 40 is not limited as long as it can accommodate the object 30. The size, shape, and function of the imaging table 40 are determined considering the size of the object 30, etc.

[0033] Each of the multiple cameras 10 rotates the object 30 on the shooting table 40 at predetermined angular intervals (for example, 15 degrees at a time), stops, and takes a picture with each camera while the object is stopped, thereby capturing the object 30 from multiple shooting directions. The captured images of the object 30 are stored in the camera 10 or the controller 20.

[0034] The object 30 can be any object from which you want to generate a CG image, such as a product, a work of art, or a building, and its shape, size, and color are not limited.

[0035] The controller 20 is configured to control the entire image processing system 1. The controller 20 controls the operation of the camera 10, the lighting device 12, and the shooting table 40. The controller 20 is configured to generate a 3D model (3D data) of the object 30 from multiple images of the object 30 by executing a program, and further generate a CG image from the 3D data. The controller 20 displays the CG image on a display and also accepts instructions on the viewpoint from the user. As will be described later, the controller 20 is configured to generate a second image from a first image of the object 30.

[0036] [Controller Configuration] Figure 2 is a block diagram showing the schematic hardware configuration of the controller 20. As shown in Figure 2, the controller 20 is composed of a processor 200, RAM 201 (Random Access Memory), ROM 202 (Read Only Memory), storage device 203, input / output interface 204, operation unit 205, and display 206, etc. These components are connected by a bus 207. The controller 20 can communicate with the camera 10, lighting device 12, shooting table 40, and various external devices via the input / output interface 204. The processor 200 is an example of an image processing device.

[0037] [Processor Configuration] Figure 3 is a diagram showing the functional configuration of the processor 200. As shown in Figure 3, the processor 200 comprises an image capture control unit 210, a 3D data generation unit 211, a condition reception unit 212, a CG image generation unit 213, an image processing unit 214, and an input / output control unit 215.

[0038] The shooting control unit 210 controls the operation of the camera 10, lighting device 12, and shooting table 40 according to predetermined shooting conditions. The user inputs the shooting conditions to the shooting control unit 210 via the operation unit 205. The camera 10 takes multiple first images of the object 30.

[0039] The 3D data generation unit 211 generates 3D data based on multiple first images of the object 30. The condition receiving unit 212 receives virtual light source information and virtual viewpoint information for the 3D data. The CG image generation unit 213 generates a CG image from the 3D data based on the virtual light source information and virtual viewpoint information. The image processing unit 214 generates a second image from the first images of the object 30. The input / output control unit 215 controls the input and output of various data.

[0040] The functions of the shooting control unit 210, the 3D data generation unit 211, the condition reception unit 212, the CG image generation unit 213, the image processing unit 214, and the input / output control unit 215 will be explained in detail in the processing flow described later.

[0041] The processor 200 may be composed of one or more hardware components, and the type of hardware is not limited. For example, the processor 200 may be composed of hardware such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), FPGA (Field Programmable Gate Array) or other programmable logic devices, an ASIC (Application Specific Integrated Circuit) or other dedicated circuit for executing specific processing, a GPU (Graphic Processing Unit), or an NPU (Neural Processing Unit). The processor 200 also has various parts (Units) or means (Means) that execute the various processing in this embodiment. Furthermore, the type of hardware may be a combination of different types of hardware. When multiple hardware components are configured to execute one or more processing of a processor, these multiple hardware components may be located in physically separate devices or in the same device. In addition, in any embodiment, the order of processing by the processor is not particularly limited and may be changed as appropriate. The hardware is composed of electrical circuits (circuitry) and the like, which are combinations of circuit elements such as semiconductor elements.

[0042] Furthermore, in the present embodiment, the processor 200 may be implemented by hardware, software, firmware, microcode, or a combination thereof. Software, firmware, and microcode are composed of programs. The program may be, for example, a group of program modules, and each function may be realized by a processor configured to execute each function. The program may be program code or a plurality of code segments stored in one or more non-transitory and tangible computer-readable media (such as a storage medium or other storage; it may be the ROM 202 or the storage device 203 (the same applies hereinafter)). The program may be divided and stored in a plurality of non-transitory and tangible computer-readable media existing in physically separate devices. The program code or code segment may represent any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. The program code or code segment may be connected to other code segments or hardware circuits by transmitting and receiving information, data, arguments, parameters, or the content of memory.

[0043] Note that in the present embodiment, the "non-transitory and tangible computer-readable media" does not include non-tangible storage media such as carrier signals or propagation signals themselves. The processor 200 can use the RAM 201 as a temporary storage area or working area during processing using the program.

[0044] Also, the functions of the above-described processor 200 may be realized by various types of AI (Artificial Intelligence). Such AI may be, for example, AI that performs image recognition, segmentation, feature extraction, etc., or AI that generates an image from given information. These AIs can also be realized by hardware, software, firmware, microcode, or a combination thereof as described above.

[0045] [Configuration of the operation unit and display] Returning to FIG. 2, the operation unit 205 is composed of devices such as a keyboard, mouse, buttons, switches, etc. not shown in the figure. The user can give instructions to the controller 20 through these devices. The processor 200 receives the instruction and performs processing according to the received instruction. The display 206 can be a touch panel type device. The user may be able to give instructions through the touch panel. The display 206 can be composed of a device such as a touch panel type device or a liquid crystal display device. The display 206 can display the first image, 3D data, CG image, second image, etc. Also, various information stored in the storage device 203 can be displayed on the display 206.

[0046] [Configuration of the input / output interface] The input / output interface 204 is composed of terminals, slots, etc. for connecting external devices such as displays, printers, and storage media, and communication interfaces such as Wi-Fi (registered trademark) and Bluetooth (registered trademark). The controller 20 can acquire data such as image data from external devices (server devices, storage devices, databases, imaging devices, etc.) through the input / output interface 204. The external device may be connected to the controller 20 by wire or wirelessly. Also, the external device may be connected via a network such as the Internet.

[0047] [Configuration of the storage device] The storage device 203 is composed of a storage medium (non-temporary and tangible computer-readable medium) such as a hard disk, semiconductor memory such as SSD (Solid-State Drive), various magneto-optical storage media, and their control units, and stores or saves various information.

[0048] [Processing in the system] Next, the processing in the image processing system 1 with the above-described configuration will be described. FIG. 4 is a flowchart showing the processing in the image processing system 1.

[0049] The camera captures multiple images of the object 30 (step S1). The user fixes the object 30 on the shooting table 40. Illumination light from the lighting device 12 shines on the object 30. At the start of shooting, the camera 10 and the object 30 are first set to arbitrary reference positions. Then, the shooting table 40 rotates the object 30 at regular angular intervals (for example, 15 degrees each) and stops, and while it is stopped, the camera 10 takes pictures of the object 30, and this process is repeated. Specifically, the shooting control unit 210 controls the camera 10, the lighting device 12 and the shooting table 40 according to the determined shooting conditions. The shooting control unit 210 controls the camera 10, lighting device 12, and shooting table 40 based on shooting conditions (specifically, including 1. the position and direction of shooting, 2. the degree to which the angular interval of the shooting direction is set, 3. the position and direction of illumination device, 4. the color and brightness of the illumination light, and 5. exposure and shutter speed, etc.) that take into account the shape, surface characteristics, and other features of the object 30, as well as the accuracy requirements for the 3D data.

[0050] The shooting conditions can be set by the user, and past shooting conditions can also be used. Past shooting conditions are stored, and the shooting control unit 210 may read the shooting conditions and execute processing.

[0051] Figure 5 is a conceptual diagram showing a state in which an object 30 is photographed by multiple cameras 10. As shown in Figure 5, the bottom camera 10 photographs the object 30 according to the conditions while the object 30 rotates, acquiring multiple images 61 of the object 30. In addition to the images 61, the camera 10 can acquire metadata (Exif information). Metadata includes, for example, basic information (date and time of shooting, camera model name and lens type, etc.), shooting settings (aperture value, shutter speed, ISO sensitivity, focal length and exposure compensation value, etc.), location information (GPS coordinates (latitude and longitude) and shooting location, etc.), camera orientation (image orientation (vertical and horizontal) and camera tilt, etc.). It also includes other information (image size, color space, white balance and software used, etc.). Metadata may be used in photogrammetry processing to generate 3D data.

[0052] The middle camera 10 and the top camera 10 similarly acquire multiple captured images 62 and multiple captured images 63 of the object 30. Meta information is acquired for each of the captured images 62 and 63. Captured images 61, 62 and 63 are examples of the first images of the present invention.

[0053] Although the example shows the use of three cameras 10, the number of cameras 10 is not particularly limited. Furthermore, the cameras 10 may be positioned above the top row or below the bottom row.

[0054] Multiple captured images 61, 62, and 63 are stored in the camera 10's storage device, the controller 20's storage device 203, or both.

[0055] Next, returning to Figure 4, once the acquisition of multiple images 61, 62, and 63 (step S1) is complete, the process proceeds to the generation of 3D data (step S2).

[0056] The controller 20 (processor 200) generates three-dimensional data from multiple captured images 61, 62, and 63 using photogrammetry. Note that the technique of generating three-dimensional data using photogrammetry is a known technique. The outline of the process is as follows.

[0057] Figure 6 shows a functional block diagram of the 3D data generation unit 211 of the processor 200. As shown in Figure 6, it includes an image acquisition unit 211A, a point cloud data generation unit 211B, a 3D patch model generation unit 211C, and a final data generation unit 211D.

[0058] The image acquisition unit 211A acquires, for example, multiple captured images 61, 62, and 63 from the storage device 203.

[0059] The point cloud data generation unit 211B analyzes multiple captured images 61, 62, and 63 and performs a process to generate three-dimensional point cloud data of feature points. First, the point cloud data generation unit 211B extracts feature points from each of the captured images 61, 62, and 63. Next, the point cloud data generation unit 211B matches corresponding feature points among the captured images 61, 62, and 63. Based on the matching results, the point cloud data generation unit 211B estimates the camera parameters of the camera (focal length, camera position and orientation, etc.). Then, based on the estimated camera parameters, it determines the three-dimensional positions of the feature points of the object 30. The three-dimensional coordinates of the feature points estimated in this way become point cloud data (point cloud).

[0060] The 3D patch model generation unit 211C generates a 3D patch model of the object 30 based on the 3D point cloud data of the object 30 generated by the point cloud data generation unit 211B. Specifically, it generates patches (meshes) from the generated 3D point cloud and generates a 3D patch model.

[0061] The final data generation unit 211D generates 3D data, which is the final data with textures applied, by performing texture mapping on the 3D patch model generated by the 3D patch model generation unit 211C. By mapping textures to the mesh, the final data generation unit 211D can give the 3D patch model a realistic appearance of the object 30.

[0062] The generated 3D data is stored in a storage device 203 or the like. Additionally, the 3D data is displayed on the display 206 as needed.

[0063] When creating 3D data, providing distance information between two marks 40B on the imaging table 40 can serve as a reference for the size of the 3D data. Alternatively, the positional relationship between the camera 10 and the object 30 can be determined based on multiple marks 40B on the imaging table 40, and 3D data can be generated from there.

[0064] In this example, the 3D data is generated based on the captured images 61, 62, and 63 obtained by photographing the object 30 from multiple shooting positions and multiple shooting directions.

[0065] The shooting location may be determined using GPS information in the metadata. Alternatively, image distortion may be corrected using lens information through camera calibration. Furthermore, suitable images for 3D data 70 may be selected based on ISO sensitivity and exposure information.

[0066] The three-dimensional data generated in this manner includes information about the shape of the object 30 and pixel value information related to its color. Furthermore, the three-dimensional data may also include material information such as reflective properties and surface roughness. Each piece of information and how it is assigned will be described later.

[0067] Shape information can be obtained by extracting feature points through image analysis during point cloud data generation, generating 3D point cloud data, and then generating a mesh (patch) from the point cloud data during 3D patch model generation to obtain more detailed shape information.

[0068] Pixel value information can be acquired by the camera 10 when photographing the object 30. In particular, detailed pixel value information can be obtained by photographing with the high-resolution camera 10. Furthermore, when texture mapping, the pixel value information can be applied to the surface of the 3D data to add color information.

[0069] Material information includes surface roughness information and reflection characteristics information. Such information can be estimated from images taken to generate 3D data, or it can be more accurately estimated by taking images in an environment where the shooting conditions and lighting conditions are known, thereby allowing for the estimation of the surface reflection characteristics of the object 30. It is also possible to obtain this information by manually inputting it into each pixel to generate the 3D data after acquiring the shape information and pixel value information.

[0070] Next, Figure 7 shows the relationship between the object 30 and the lighting device 12 in actual photography, and the relationship between the 3D data 70 in the virtual space and the virtual light sources 72, 73, and 74 in the 3D space.

[0071] In the actual shooting shown in Figure 7A, the positional relationship between the object 30 and the lighting device 12 is determined before the camera 10 takes a picture. While the camera 10 is shooting the object 30, the positional relationship between the object 30 and the lighting device 12 remains fixed. Therefore, it is not possible to change how the light from the lighting device 12 falls on the object in the images 61, 62, and 63 taken by the camera 10 from different angles after the initial shooting, except by changing the positional relationship between the object 30 and the lighting device 12 and reshooting. On the other hand, it is not practical to repeatedly take pictures with the camera 10 while changing the positional relationship between the object 30 and the lighting device 12, taking into account how the light falls on the object in the captured images.

[0072] In the virtual 3D space on the computer shown in Figure 7B, a virtual light source and virtual viewpoint can be selected for the 3D data 70. In Figure 7B, three virtual light sources 72, 73, and 74 are shown. By selecting a virtual light source and virtual viewpoint, the 3D data 70 can be converted and output as a 2D image (CG image) as seen from that virtual light source and virtual viewpoint. In other words, this process does not require actual photography of the object 30 by the camera 10.

[0073] A virtual light source is a light source placed in a three-dimensional space that simulates actual lighting on a computer. The settings of a virtual light source affect the appearance of 3D data, including shading, reflection, and refraction. Types of virtual light sources include point lights, parallel lights, spotlights, and area lights. A virtual light source can be further specified by parameters such as position and direction, color and intensity, and attenuation rate (attenuation of light with distance). This information constitutes the virtual light source information contained within the virtual light source.

[0074] A virtual viewpoint defines where the camera is located and which direction it is facing in three-dimensional space. Setting the virtual viewpoint is related to how three-dimensional data is displayed as a two-dimensional computer graphics image. A virtual viewpoint can be freely positioned anywhere in three-dimensional space and can face any direction. A virtual viewpoint can be defined by parameters such as position (x, y, z coordinates) and orientation (direction vector). This information constitutes the virtual viewpoint information contained within the virtual viewpoint.

[0075] Next, returning to Figure 4, after completing the generation of 3D data (step S2), the system receives virtual light source information and virtual viewpoint information (step S3). Specifically, the condition receiving unit 212 of the processor 200 receives the virtual light source information and virtual viewpoint information described above. The virtual light source information and virtual viewpoint information can be input by the user from the operation unit 205 to the condition receiving unit 212 via the input / output interface 204. Alternatively, it can be input by reading the virtual light source information and virtual viewpoint information stored in the storage device 203. Furthermore, it can be automatically input by selecting an appropriate scene (time of day, weather, etc.) in the CG image generation program.

[0076] Next, a CG image is generated (step S4). The controller 20 (processor 200) generates a CG image 80 from the virtual light source information, virtual viewpoint information, and 3D data 70. The technique for generating the CG image is a known technique. The outline of the process is as follows.

[0077] Figure 8 shows a functional block diagram of the CG image generation unit 213 of the processor 200. As shown in Figure 8, the CG image generation unit 213 includes a condition setting unit 213A, a rendering processing unit 213B, a post-processing processing unit 213C, and a final image generation unit 213D.

[0078] The condition setting unit 213A receives virtual light source information and virtual viewpoint information from the condition receiving unit 212 and sets the conditions necessary for generating CG images. Based on the virtual light source information, it sets the type of light source, the position and direction of the light source, the color and intensity of the light source, etc. Also, based on the virtual viewpoint information, it sets the viewpoint position and line of sight direction in the virtual 3D space.

[0079] Next, the rendering processing unit 213B receives the 3D data 70 and generates a 2D CG image based on the 3D data 70, virtual light source information, and virtual viewpoint information. The post-processing processing unit 213C performs color grading to adjust the image's color tone, applies a pseudo-depth of field, and various effects, and further performs processing such as noise reduction and image enhancement. The final image generation unit 213D generates the completed CG image 80 as the final image in a format appropriate for the purpose.

[0080] Figure 9 shows the captured image 81, taken with the light source used during shooting, and CG images 82, 83, and 84 with different virtual light sources. Captured image 81 is an image with the illumination from the lighting device 12 at the time of shooting as the light source. As previously described, the light source conditions cannot be changed in captured image 81. CG images 82, 83, and 84 are CG images generated using virtual light sources 72, 73, and 74, respectively. The CG image generation unit 213 can generate CG images 82, 83, and 84 with different light sources according to the respective virtual light source information and virtual viewpoint information. In this example, CG image 82 corresponds to a virtual light source 72 that emits light from above, CG image 83 corresponds to a virtual light source 73 that emits light from slightly above and to the side, and CG image 84 corresponds to a virtual light source 74 that emits light from slightly below and to the side. In captured image 81 and CG images 82, 83, and 84, the background that is not subject to processing is shown with diagonal lines.

[0081] Here, while both the captured image 81 and the CG images 82, 83, and 84 are two-dimensional images, there are the following differences. Captured image 81 generally has a higher resolution than CG images, but the light source cannot be freely changed. This is based on the premise that the light source position and light irradiation direction are fixed for an image taken from a single shooting position and direction (see 7A in Figure 7). On the other hand, with CG images 82, 83, and 84, it is possible to freely obtain CG images with a changed light source by changing the settings of the virtual light source (see 7B in Figure 7). However, generally, obtaining high-resolution 3D model data of an object is not easy in terms of the effort required for shooting, generation time, and data capacity, and as a result, generating high-resolution CG images is also not easy.

[0082] Therefore, the inventors arrived at the present invention, which generates a new second image from a first image based on a computer graphics image. This invention makes it possible to generate a high-resolution second image with a different light source from a high-resolution first image.

[0083] Returning to Figure 4, after the generation of the CG image (step S4) is completed, the second image 91 is generated from the first image 90 (step S5). Specifically, the controller 20 (processor 200) generates the second image 91 from the first image 90.

[0084] Figure 10 shows a functional block diagram of the image processing unit 214 of the processor 200 that generates the second image 91 from the first image 90. As shown in Figure 10, the image processing unit 214 includes a first intermediate CG image generation unit 214A, a second intermediate CG image generation unit 214B, a first difference image generation unit 214C, and a second image generation unit 214D. The image processing unit 214 generates a first intermediate CG image 86 and a second intermediate CG image 87 through processing in each unit, generates a first difference image 88 from the first intermediate CG image 86 and the second intermediate CG image 87, and generates the second image 91 from the first image 90 and the first difference image 88.

[0085] The first intermediate CG image generation unit 214A generates a first intermediate CG image 86 using virtual light source information, virtual viewpoint information, material information, and pixel value information. The virtual light source information and virtual viewpoint information are obtained from the condition receiving unit 212. This information includes at least one of shape information, material information, and pixel value information, and this information is included in the 3D data 70. The first intermediate CG image generation unit 214A can generate the first intermediate CG image 86 by the same processing as the CG image generation unit 213.

[0086] The second intermediate CG image generation unit 214B generates the second intermediate CG image 87 using virtual light source information, virtual viewpoint information, shape information, and pixel value information. The second intermediate CG image generation unit 214B generates the second intermediate CG image 87 by the same process as the first intermediate CG image generation unit 214A, but unlike the first intermediate CG image 86, the second intermediate CG image 87 does not include material information. As a result, the second intermediate CG image 87 lacks detailed surface characteristics such as reflectivity and transparency because it does not contain material information, and the image does not express material characteristics (such as metallic or plastic feel).

[0087] The first difference image generation unit 214C generates a first difference image 88 from the first intermediate CG image 86 and the second intermediate CG image 87. For example, the difference can be obtained by comparing the first intermediate CG image 86 and the second intermediate CG image 87 on a pixel-by-pixel basis and subtracting RGB values, brightness values, etc.

[0088] The first intermediate CG image 86 includes virtual light source information, virtual viewpoint information, shape information, material information, and pixel value information, while the second intermediate CG image 87 includes virtual light source information, virtual viewpoint information, shape information, and pixel value information. The first intermediate CG image 86 and the second intermediate CG image 87 share virtual light source information, virtual viewpoint information, shape information, and pixel value information, differing only in the presence or absence of material information. Therefore, the first difference image 88 expresses the texture of the object 30, such as gloss, transparency, and reflectivity, under the virtual light source conditions when the first and second intermediate CG images were generated. In this example, in the first difference image 88, areas with high reflectivity in the material information are represented as whitish, and areas with low reflectivity are represented as blackish.

[0089] The second image generation unit 214D generates the second image 91 from the first difference image 88 and the first image 90. Here, the first image 90 is an image (actual photograph) of the object 30.

[0090] The second image 91 can be generated, for example, by combining the first difference image 88 and the first image 90. When combining them, for example, one of them can be weighted, and the sum of the pixel values ​​can be taken to generate the second image 91. By weighting, the influence of the first difference image 88 or the first image 90 can be emphasized. The example given is the combination of the first difference image 88 and the first image 90, but the second image 91 may be generated by other methods. For example, since the first image 90 includes the influence of the light source at the time the first image was taken, the second image 91 may be generated by adding the first image, from which the influence of the light source at the time of shooting has been suppressed by processing, and the first difference image 88. Furthermore, the weighting described above may be applied when adding them together. As for the processing to reduce the influence of the light source at the time of shooting from the first image 90, for example, an AI-based process can be used.

[0091] Since the second image 91 is created based on the first image 90, the second image 91 and the first image 90 have the same resolution, and the second image 91 reproduces an image as if it were taken with a light source in a virtual space.

[0092] As previously described, the first image 90 is an image of the object 30. This image is selected from among multiple images of the object 30 taken from multiple shooting positions and directions, which were used to generate the 3D data 70. For example, it is an image taken from a shooting position and direction that matches or is similar to the virtual viewpoint information used when generating the CG images 82, 83, and 84. Specifically, the selected image may match or be closest to the shooting position and direction saved as metadata at the time of shooting, or it may match or be closest to the shooting direction and shooting position of each image estimated as camera parameters when generating the 3D data. Alternatively, the selected image may be the one most similar to the generated CG image. In this example, when any of the CG images 82, 83, and 84 is selected, the first intermediate CG image 86 and the second intermediate CG image 87 are generated to match the virtual light source information and virtual viewpoint information used when generating that CG image.

[0093] The first image 90 may be taken separately from the generation of the 3D data 70. The separately taken image is saved in association with the 3D data 70. These images are taken of the object 30 from various positions and directions, and the image with a shooting position and direction that matches the virtual viewpoint information used during CG image generation is selected as the first image 90. The selection method is the same as described above (same as when the captured image is for 3D data generation).

[0094] Next, returning to Figure 4, the process ends when the second image 91 is generated from the first image 90.

[0095] <Example of Application of Embodiment> Next, an example of using the second image 91 generated from the first image 90 of this example on the display 206 will be described. Figure 11 is a diagram showing an example of the screen of the display 206 when this example is applied. As shown in XIA of Figure 11, a CG image 82 is displayed on the upper side of the display 206. The CG image 82 is the image when the virtual light source 72 is selected. On the other hand, no image is displayed on the lower side of the display 206. In this state, the user can select the viewpoint position and line of sight direction (virtual viewpoint information) that they want to see in detail in the CG image 82 by mouse operation, touch operation on the touchscreen, etc.

[0096] When a viewpoint is selected, as shown in XIB of Figure 11, the processor 200 displays the image taken at the virtual viewpoint information and corresponding shooting position and direction as the first image 90. Furthermore, the processor 200 displays the second image 91 generated from the first image 90. As previously described, the second image 91 is generated from the first difference image 88 and the first image 90. When the processing of the image processing unit 214 is completed, the first image 90 and the second image 91 can be displayed on the display 206 simultaneously. Alternatively, the first image 90 can be displayed on the display 206 first, and then the second image 91 can be displayed on the display 206. Alternatively, only the CG image 82 may be displayed, and the first image 90 may not be displayed at the user's instruction, and the second image 91 may be displayed alongside the CG image 82. When compared to CG image 82, the second image 91 is just as high-resolution as the first image 90, and appears as if it were taken under the same lighting conditions as the virtual light source in CG image 82. Therefore, the user can observe the details of the object 30 without any sense of incongruity.

[0097] Next, we will describe the state when the material information of the 3D data 70 is changed. As shown in XIIA of Figure 12, the display 206 shows the CG image 82, the first image 90, and the second image 91. This shows the same state as XIB in Figure 11.

[0098] The user can change the material information at any time. When the material information is changed, the second image 91 is also changed accordingly. Figure 12, XIIB shows the case where the user has modified the material information to increase the reflection. When the material information is modified, the first intermediate CG image generation unit 214, when generating the first intermediate CG image 86, modifies it, for example, in a direction that increases the reflection. The image processing unit 214 generates a first difference image 88 from the first intermediate CG image 86 and the second intermediate CG image 87, which were generated based on the modified material information. This first difference image 88 is a difference image in which the reflection component is larger than before the material information was changed. By combining this first difference image 88 with the first image 90, a second image 91 with greater reflection is generated compared to when the material information is not changed. Conversely, if the user modifies the material information to reduce the reflection, a second image 91 with less reflection is generated compared to when the material information is not changed.

[0099] <Suppression of Reflected Light> Next, a preferred embodiment will be described. As previously mentioned, in this example, the second image 91 is generated by combining the first difference image 88 and the first image 90. In this case, it is preferable that the first image 90, which is also a captured image of the object 30, is an image in which the effect of reflection is suppressed. When generating the second image 91, a reduction process to reduce the effect of the light source has been described, but this time, a method for obtaining the first image 90 with little effect of reflection will be described.

[0100] Firstly, as shown in Figure 13, reflections are suppressed by shooting under uniform lighting conditions. For example, a shooting environment is created in which multiple lighting devices 12 evenly arranged around the object 30 emit light of the same intensity onto the object 30. This makes it possible to obtain a captured image in which the object 30 is not illuminated by light of uneven intensity from uneven directions. Note that the camera 10 is omitted in Figure 13 for the sake of simplicity.

[0101] Secondly, as shown in Figure 14, reflections are suppressed using two polarizing filters. This method is as follows: The first polarizing filter 12A is placed between the lighting device 12 and the object 30. The second polarizing filter 10A is placed between the camera 10 and the object 30. The first polarizing filter 12A polarizes the illumination light from the lighting device 12. The reflected light is adjusted by rotating the second polarizing filter 10A. While looking at the viewfinder or LCD screen of the camera 10, the rotation angle of the polarizing filter 10A is adjusted to the position where reflected light is suppressed. The rotation angle of the polarizing filter 10A is usually the position where the polarization axes of polarizing filter 10A and polarizing filter 12A are orthogonal, but for example, the optimal angle of polarizing filter 10A may be estimated by using image recognition processing software. Alternatively, multiple images may be taken while rotating the polarizing filter 10A, and then the image in which reflected light is most suppressed may be selected as the first image 90 from the captured images.

[0102] <Resolution Settings> The processor 200 is configured to allow setting and changing the resolution of various images.

[0103] Firstly, the user can set the resolution of the generated second image 91. The user can input resolution information from the operation unit 205 to the image processing unit 214 via the input / output interface 204. The image processing unit 214 generates the second image 91 with the set resolution. Normally, the resolution of the second image 91 is about the same as that of the first image 90, but by setting the resolution of the second image 91 to be lower than that of the first image 90, the file size is reduced. As a result, processing can be sped up.

[0104] Secondly, the processor 200 (image processing unit 214) compares the resolution of the first difference image 88 with the resolution of the first image 90, performs a process to reduce the difference, and generates a second difference image (not shown). Subsequently, the second image 91 is generated based on the first image 90 and the second difference image. By using the second difference image with improved resolution instead of the first difference image 88, the accuracy of the second image 91 can be improved. Furthermore, in this method, it is preferable that the process to reduce the difference in resolution is an upscaling process that brings the image with the lower resolution of the first image 90 and the first difference image 88 closer to the image with the higher resolution. This further improves the accuracy of the generated second image 91.

[0105] <Color Adjustment> The processor 200 is configured to allow color adjustment. Two methods are possible for adjusting the color. Here, color can include "hue, saturation, and brightness," which are expressed objectively, and "color tone and nuance," which are expressed subjectively.

[0106] Firstly, the processor 200 (CG image generation unit 213) sets the color tones of CG images 80, 82, 83, and 84 based on the first image 90. This improves the accuracy of the generated CG images 80, 82, 83, and 84. For example, color features can be extracted from the first image using a known method, and CG images can be generated to have the same color tone.

[0107] Secondly, the processor 200 (image processing unit 214) adjusts the color of the second image 91 based on the CG images 80, 82, 83, and 84. As a result, the color of the second image 91 becomes closer to the actual captured image, and the accuracy of the second image 91 is improved. For example, using known methods, image processing set by the user for the CG image can also be reflected in the second image.

[0108] <Generation of the second image> In the embodiment, the case in which the second image 91 is generated from the first image 90 and the first difference image 88 obtained from the first intermediate CG image 86 and the second intermediate CG image 87 was described. Next, another method for generating the second image will be described.

[0109] The image processing unit 214 of the processor 200 can generate a second image 91 from a first image 90 based on virtual light source information related to the virtual light source when the CG image generation unit 213 performs rendering. The virtual light source information includes the position of the light source, the type of light source, the intensity of the light source, and the color of the light source. The second image 91 is generated by applying this information to the first image 90.

[0110] <Image 1> The above example describes an embodiment in which 3D data is generated from a group of 2D images, and the light source conditions for displaying the 3D data are reflected in the original group of 2D images. In this case, the group of 2D images serves both as "photogrammetry data" and "display and editing data."

[0111] However, two-dimensional images (where the position and orientation of the object and the imaging device are known) taken separately from those used for three-dimensional data generation may also be displayed and edited. For example, a group of two-dimensional images taken under imaging conditions suitable for shape reproduction (three-dimensional model generation) is used for three-dimensional data generation. Two-dimensional images (groups) for display can be taken separately under imaging conditions suitable for creating images with changed light sources.

[0112] Furthermore, the 3D data itself can be created by scanning and measuring the object using known methods such as LiDAR (Light Detection and Ranging) or Structured Light, without using photogrammetry, or by using 3D CAD data (design drawings). Separately, a set of 2D images taken of the object with the position and orientation of the object and the imaging device known at the time of shooting may be used as images for display and editing. This is particularly effective for objects where it is not possible or necessary to take a full 360-degree image around the object (e.g., extremely large objects, or when the 3D shape of the back side is not needed).

[0113] 1 Image Processing System 10 Camera 10A Polarizing Filter 12 Lighting Device 12A Polarizing Filter 20 Controller 30 Object 40 Shooting Platform 40A Mounting Surface 40B Mark 44 Support Member 44A Mounting Platform 44B Arm 61 Captured Image 62 Captured Image 63 Captured Image 70 3D Data 72 Virtual Light Source 73 Virtual Light Source 74 Virtual Light Source 80 CG Image 81 Captured Image 82 CG Image 83 CG Image 84 CG Image 86 First Intermediate CG Image 87 Second Intermediate CG Image 88 First Difference Image 90 First Image 91 Second Image 200 Processor 201 RAM 202 ROM 203 Storage Device 204 Input / Output Interface 205 Operation Unit 206 Display 207 Bus 210 Shooting Control Unit 211 3D Data Generation Unit 211A Image Acquisition Unit 211B Point Cloud Data Generation Unit 211C 3D Patch Model Generation Unit 211D Final Data Generation Unit 212 Condition Acceptance Unit 213 CG Image Generation Unit 213A Condition Setting Unit 213B Rendering Processing Unit 213C Post-Processing Processing Unit 213D Final Image Generation Unit 214 Image Processing Unit 214A First Intermediate CG Image Generation Unit 214B Second Intermediate CG Image Generation Unit 214C First Difference Image Generation Unit 214D Second Image Generation Unit 215 Input / Output Control Unit

Claims

1. An image processing device having a processor, the processor receiving three-dimensional data of an object, receiving virtual light source information and virtual viewpoint information relating to a virtual light source for the three-dimensional data, generating a CG image from the three-dimensional data based on the virtual light source information and the virtual viewpoint information, and generating a second image from at least one first image of the object based on the virtual light source information and the virtual viewpoint information.

2. The three-dimensional data is stored in association with captured images obtained by photographing the object from multiple shooting positions and multiple shooting directions, and when the processor generates the CG image, it selects the captured image from the shooting position and shooting direction corresponding to the virtual viewpoint information as the first image, as described in claim 1.

3. The image processing apparatus according to claim 2, wherein the three-dimensional data is generated based on images obtained by photographing the object from multiple shooting positions and multiple shooting directions.

4. The image processing apparatus according to any one of claims 1 to 3, wherein the three-dimensional data includes at least one of the shape information, pixel value information, and material information of the object, the processor generates a first intermediate CG image using the shape information, the virtual light source information, the virtual viewpoint information, the material information and the pixel value information, generates a second intermediate CG image using the shape information, the virtual light source information, the virtual viewpoint information and the pixel value information, generates a first difference image from the first intermediate CG image and the second intermediate CG image, and generates the second image from the first difference image and the first image.

5. An image processing apparatus according to any one of claims 1 to 3, which generates a second image from a first image based on the virtual light source information relating to the virtual light source when rendering the three-dimensional data.

6. The image processing apparatus according to claim 4, wherein the processor performs a process to reduce the difference between the resolution of the first difference image and the resolution of the first image to obtain a second difference image, and generates the second image based on the first image and the second difference image.

7. The image processing apparatus according to claim 6, wherein the process for reducing the difference in resolution is a process for bringing the image with the lower resolution among the first image and the first difference image closer to the image with the higher resolution.

8. The image processing apparatus according to claim 1, wherein the processor sets the color of the CG image based on the first image.

9. The image processing apparatus according to claim 1, wherein the processor sets the color of the second image based on the CG image.

10. The image processing apparatus according to claim 4, wherein the processor receives a change in the material information of the three-dimensional data and generates the second image based on the changed material information.

11. The image processing apparatus according to claim 1, wherein the processor sets the resolution of the second image to be generated.

12. The image processing apparatus according to claim 1, wherein the first image is an image taken under uniform lighting conditions during shooting, or an image taken while suppressing reflection from the object.

13. An image processing method that causes a processor to perform the following processes: receive three-dimensional data of an object; receive virtual light source information and virtual viewpoint information relating to a virtual light source for the three-dimensional data; generate a CG image from the three-dimensional data based on the virtual light source information and the virtual viewpoint information; and generate a second image from at least one first image of the object based on the virtual light source information and the virtual viewpoint information.

14. A program that causes a computer to perform the following processes: receive three-dimensional data of an object; receive virtual light source information and virtual viewpoint information related to a virtual light source for the three-dimensional data; generate a CG image from the three-dimensional data based on the virtual light source information and the virtual viewpoint information; and generate a second image from at least one first image of the object based on the virtual light source information and the virtual viewpoint information.

15. A non-temporary and computer-readable recording medium on which the program described in claim 14 is recorded.

Citation Information

Patent Citations

  • Photograph search / browsing system and program, using three-dimensional model, and three-dimensional model display / operation system and program, using photograph

    JP2006309722A

  • System and method for performing 3D imaging of objects

    JP2021026759A

  • Image processing apparatus, method for controlling image processing apparatus, and program

    JP2022093262A