Image processing apparatus and method, image pickup apparatus, and storage medium

The image processing device enhances virtual subject superimposition by incorporating real-world weather and terrain data, ensuring natural-looking superimposed images.

JP2026004030APending Publication Date: 2026-01-14CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024102222
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2026-01-14

AI Technical Summary

Technical Problem

Existing image processing techniques fail to consider real-world weather and terrain information when superimposing virtual subjects, resulting in unnatural images.

Method used

An image processing device that acquires scene information using LiDAR and environmental estimation, generates virtual subjects based on weather and spatial data, and superimposes them to reflect real-world conditions.

Benefits of technology

Generates superimposed images that are more consistent with the captured scene, considering environmental and spatial information for realistic composition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026004030000001_ABST
    Figure 2026004030000001_ABST
Patent Text Reader

Abstract

To superimpose a virtual subject more consistent with the situation of a photographed video image.SOLUTION: The image processing apparatus includes an acquisition unit configured to acquire scene information of a scene captured by an imaging unit, a generation unit configured to generate a virtual subject, a processing unit configured to process the virtual subject based on the scene information, and a superimposition unit configured to superimpose the virtual subject processed by the processing unit on image data of the scene obtained from the imaging unit.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device and method, an imaging device, a program, and a storage medium, and more particularly to a technique for superimposing a virtual subject on a captured image. [Background technology]

[0002] One known shooting technique is to superimpose a virtual subject onto an image captured by a camera and use the superimposed image to consider shooting conditions even when the subject is not present. However, simply superimposing a virtual subject onto a captured image does not reflect the real lighting conditions, resulting in an unnatural superimposed image.

[0003] In response to this, Patent Document 1 discloses a technique for reflecting real light source information on a virtual subject in order to fill the gap between the lighting conditions of the captured video and the virtual subject that is superimposed. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-163610 Summary of the Invention [Problem to be solved by the invention]

[0005] However, while the technology described in Patent Document 1 can bridge the gap between the lighting conditions of a captured image and a virtual subject, it does not take into consideration reflecting real-world weather information or terrain information. As a result, unnatural superimposed images are generated in some cases, and it is not possible to generate live view images that allow for appropriate consideration of shooting conditions.

[0006] The present invention has been made in consideration of the above-mentioned problems, and has an object to make it possible to superimpose a virtual subject that is more consistent with the situation of the captured video. [Means for solving the problem]

[0007] In order to achieve the above object, the image processing device of the present invention has an acquisition means for acquiring scene information of a scene being captured by an imaging means, a generation means for generating a virtual subject which is a virtual subject, a processing means for processing the virtual subject based on the scene information, and a superimposition means for superimposing the virtual subject processed by the processing means on image data of the scene obtained from the imaging means. [Effects of the Invention]

[0008] According to the present invention, it is possible to superimpose a virtual subject that is more consistent with the situation of the captured video. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a block diagram showing the functional configuration of an imaging apparatus according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a conceptual diagram showing a method for acquiring spatial information in the first embodiment. [Figure 3] 3A to 3C are conceptual diagrams illustrating a method for generating a learning model of environmental information and an estimation method using the learning model in the first embodiment. [Figure 4] 5 is a flowchart showing a live view display process in the first embodiment. [Figure 5] 5A to 5C are diagrams showing a specific example of processing when a virtual subject is superimposed on an image of wind blowing in the first embodiment. [Figure 6] 5A to 5C are diagrams showing a specific example of processing when a virtual subject is superimposed on a video of rain in the first embodiment. [Figure 7] 5A to 5C are diagrams showing a specific example of processing when a virtual subject is superimposed on an image of falling snow in the first embodiment. [Figure 8] 5A to 5C are diagrams showing a specific example of processing when a virtual subject is superimposed on an image of a slope in the first embodiment. [Figure 9]5A to 5C are diagrams showing a specific example of processing when a virtual subject is superimposed on an image of a road with trees in the first embodiment. [Figure 10] 10 is a flowchart showing a live view display process according to the second embodiment. [Figure 11] 10A and 10B are diagrams showing a specific example of processing when a virtual subject superimposed on an image including a puddle has an effect in the second embodiment. [Figure 12] 10A and 10B are diagrams showing a specific example of processing in the case where a virtual subject superimposed on an image including accumulated snow has an effect in the second embodiment. [Figure 13] 10A and 10B are diagrams showing a specific example of processing when a virtual subject superimposed on an image including rain has an effect in the second embodiment. [Figure 14] 10A and 10B are conceptual diagrams of a method for generating a learning model of spatial information and an estimation method using the learning model in a modified example. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0011] First Embodiment First, a first embodiment of the present invention will be described. FIG. 1 is a block diagram showing an example of the functional configuration of an imaging device 100. As shown in FIG. 1, the imaging device 100 includes a CPU 101, a storage unit 102, an imaging unit 103, an image processing unit 104, a display unit 105, a space recognition unit 106, an environment information estimation unit 107, a virtual object generation unit 108, a virtual object processing unit 109, a superimposition unit 110, a communication unit 111, an operation unit 113, and a system bus 112. In the following embodiment, a digital camera will be described as an example of the imaging device 100, but the present invention is applicable to any electronic device that can be equipped with an imaging function. Examples of such electronic devices include video cameras, computer devices (personal computers, tablet computers, media players, PDAs, etc.), mobile phones, smartphones, game consoles, robots, drones, and drive recorders. These are merely examples, and the present invention is also applicable to other electronic devices.

[0012] The CPU 101 controls the entire imaging apparatus 100, and executes programs stored in a ROM (not shown) to realize each process of a flowchart described later.

[0013] The storage unit 102 is composed of a DRAM, a memory card, or the like, and records images generated by the image processing unit 104, as well as 3D objects and the actions of the 3D objects received via the communication unit 111 in response to instructions from a user using the imaging device 100. A 3D object is a three-dimensional model defined in a file format such as OBJ or FBX. The recorded 3D object is used as a virtual object representing a virtual subject in the virtual subject generation unit 108, which will be described later.

[0014] The imaging unit 103 is composed of a lens unit, an imaging element, an A / D conversion circuit, etc., and performs a series of processes to capture images and output image signals. The imaging unit 103 allows the setting of shooting conditions such as aperture value, ISO sensitivity, exposure time, zoom magnification, and selection of illumination position. The image processing unit 104 performs correction processing, encoding processing, etc. on the image signal obtained by capturing an image in the imaging unit 103. The image processing unit 104 also generates a recording image or a live view video from the image signal obtained from the imaging unit 103. The display unit 105 is configured by a liquid crystal display, an organic EL display, or the like, and displays the image generated by the image processing unit 104 or the superimposed image generated by the superimposing unit 110, which will be described later.

[0015] The communication unit 111 is an interface that connects the imaging device 100 to other devices by wire or wirelessly, and transmits and receives 3D object data, image data, etc., and can also be connected to a network such as a wireless LAN or the Internet.

[0016] The space recognition unit 106 measures the space using, for example, LiDAR (Laser Imaging Detection And Ranging) technology, and acquires spatial information about objects that make up the scene, such as the topography and the positions of obstacles in the area (scene) captured by the imaging device 100. As an example, when acquiring spatial information using LiDAR, a laser beam is irradiated onto a space such as that shown in FIG. 2(a) as shown in FIG. 2(b) and the reflected light is detected, thereby acquiring spatial information about the topography and obstacles in the space, as shown in FIG. 2(c). The spatial information includes topographical information such as the unevenness and inclination of the ground, and obstacle information such as the position and size of obstacles.

[0017] The environmental information estimation unit 107 receives as input the video data captured by the imaging unit 103, and uses a learning model stored in a ROM (not shown) to estimate environmental information related to the environment of the scene being captured, including weather information such as wind, rain, or snow contained in the video data. When wind, rain, or snow is estimated, the position, direction, and amount are also estimated, and the weather information includes these estimated pieces of information.

[0018] The virtual object generation unit 108 generates a virtual object in response to a user's specification of the virtual object, its position, size, and orientation. The virtual object can be specified, for example, by voice input, selection using a GUI display on the display unit 105, specification via the operation unit 113, or image input via a network. The virtual object is generated by selecting a 3D object stored in the storage unit 102 in response to the specification from the operation unit 113. Alternatively, the virtual object may be generated by generating a 3D object using a machine learning model that corresponds to the generation of a 3D object using the specification of the virtual object as input, or by acquiring the 3D object from a network via the communication unit 111.

[0019] The virtual object processing unit 109 processes the virtual object generated by the virtual object generation unit 108 in accordance with the environmental information estimated by the environmental information estimation unit 107 and / or the spatial information acquired by the spatial recognition unit 106. The virtual object is processed in accordance with the environmental information and / or spatial information (scene information) by inputting the virtual object and the environmental information and / or spatial information (scene information) into a learning model.

[0020] For example, if the environmental information indicates rain, the learning model predicts where the virtual subject will get wet based on the intensity and direction of the rain, and applies a water-soaked effect to the virtual subject's clothes and belongings. If the environmental information indicates wind, the learning model predicts where the virtual subject will be blown by the wind and the effect on the virtual subject's posture based on the wind intensity and direction, and applies effects to the virtual subject's hair and clothes blowing in the wind, and the virtual subject being blown by the wind. If the environmental information indicates snow, the learning model predicts where snow will accumulate on the virtual subject based on the amount and direction of snowfall, and applies effects to snow accumulating on the top of the virtual subject.

[0021] Furthermore, if the spatial information indicates unevenness or inclination of the ground, the learning model estimates the angle at which the virtual subject will tilt from the unevenness of the ground and the magnitude of the inclination, and the virtual subject is processed to tilt.If the spatial information indicates an obstacle, the learning model estimates the area in which the virtual subject's movement is restricted by the obstacle from the position and size of the obstacle, and the virtual subject is processed to move in a way that does not come into contact with the obstacle.

[0022] The superimposing unit 110 generates a superimposed image by superimposing the image data captured by the imaging unit 103 and the virtual object processed by the virtual object processing unit 109. The operation unit 113 is used by the user to input various instructions, and is made up of various operation members such as buttons, switches, a touch panel, a voice input unit, a gaze detection unit, etc. The instructions input via the operation unit 113 are input to the CPU 101, and the CPU 101 performs processing based on the input instructions. The above-mentioned components are connected to a system bus 112, and can transmit and receive necessary data to and from each other via the system bus 112.

[0023] Next, the learning model and the method for estimating environmental information will be described with reference to FIG.

[0024] As shown in FIG. 3(a), this learning model is generated by supervised learning using training data consisting of a set of video data and ground truth data, which is environmental information such as wind, rain, or snow contained in the corresponding video data. Specifically, learning is performed using a linear regression algorithm with the training data. Note that the items used as ground truth data are not limited to these items. Furthermore, other algorithms such as K-nearest neighbors and neural networks may be used in addition to linear regression.

[0025] As shown in Fig. 3(b), the estimation of environmental information is realized by inputting video data into a learning model that has been trained using environmental information as ground truth data, and having the learning model estimate the environmental information contained in the video data. Note that the input video data is video data captured by the imaging device 100. In this embodiment, the environmental information obtained as an estimation result is wind, rain, snow, or the like contained in the video data.

[0026] Fig. 4 is a flowchart showing the live view display process of the image capture device 100 according to the first embodiment. The flowchart in Fig. 4 starts when the image capture device 100 is powered on. In S401, the image processing unit 104 generates a live view image from the video data captured by the imaging unit 103, and the process proceeds to S402. In S402, the display unit 105 displays the live view image generated in S401, and the process proceeds to S403.

[0027] In S403, the CPU 101 determines whether or not a virtual object generation instruction has been issued from the user via the operation unit 113, and if a virtual object generation instruction has been issued, the process proceeds to S404; if no virtual object generation instruction has been issued, the process returns to S402 and the display of the live view video continues.

[0028] In S404, the virtual subject designated by the user via operation unit 113 is determined as the virtual subject to be generated, and the process proceeds to S405. The virtual subject designated here is, for example, a person or a car, and it is also possible to include characteristics or actions in the virtual subject, such as a person with long hair or a moving car, as necessary.

[0029] In S405, the space recognition unit 106 acquires space information of the area captured by the image capturing device 100, and the process proceeds to S406. In S406, the environmental information estimation unit 107 estimates the environmental information included in the video data acquired in S401, and the process proceeds to S407. In S407, the virtual subject generation unit 108 generates the virtual subject specified in S404, and the process proceeds to S408. The processing in S405, S406, and S407 may be performed in any order, and may also be performed in parallel.

[0030] In S408, the virtual subject processing unit 109 processes the virtual subject generated in S407 by reflecting the spatial information acquired in S405 and the environmental information estimated in S406, and the process proceeds to S409. In S409, the superimposing unit 110 generates a superimposed image by superimposing the virtual subject processed in S408 on the live view image generated in S401, and the process proceeds to S410. In S410, the display unit 105 displays the superimposed image generated in S409, and the process ends.

[0031] A specific example of the process shown in FIG. 4 will be described below with reference to FIGS. If the live view image displayed in S402 is an image in which wind is blowing as shown in Fig. 5(a), the environmental information estimated in S406 is wind. If the virtual subject generated in S407 is a person with long hair as shown in Fig. 5(b), the virtual subject processed in S408 will have the person's hair blowing as shown in Fig. 5(c), and the superimposed image displayed in S410 will be as shown in Fig. 5(d).

[0032] Furthermore, if the live view image displayed in S402 is an image of rain as shown in Fig. 6(a), the environmental information estimated in S406 is rain. If the virtual subject generated in S407 is a person as shown in Fig. 6(b), the virtual subject processed in S408 will have wet clothing as shown in Fig. 6(c), and the superimposed image displayed in S410 will be as shown in Fig. 6(d).

[0033] Furthermore, if the live view image displayed in S402 is an image of falling snow as shown in Fig. 7(a), the environmental information estimated in S406 is snow. If the virtual subject generated in S407 is a car as shown in Fig. 7(b), the virtual subject processed in S408 will be a car with snow piled up on it as shown in Fig. 7(c), and the superimposed image displayed in S410 will be as shown in Fig. 7(d).

[0034] Furthermore, if the live view image displayed in S402 is an image of a slope as shown in Fig. 8(a), the spatial information acquired in S405 is the unevenness and inclination of the ground. If the virtual subject generated in S407 is a car as shown in Fig. 8(b), the virtual subject processed in S408 will be a car that is tilted as shown in Fig. 8(c), and the superimposed image displayed in S410 will be as shown in Fig. 8(d).

[0035] Furthermore, if the live view image displayed in S402 is an image of a road with trees as shown in Fig. 9(a), the spatial information acquired in S405 is the position and size of obstacles. If the virtual subject generated in S407 is a walking person as shown in Fig. 9(b), the virtual subject processed in S408 will have walked while avoiding the trees as shown in Fig. 9(c), and the superimposed image displayed in S410 will be as shown in Fig. 9(d).

[0036] 5 to 8 show cases where either environmental information or spatial information is reflected in the virtual subject, but when both are reflected, the virtual subject can be processed as follows. For example, when snow is obtained as environmental information and tilt is obtained as spatial information, the virtual subject is processed to appear as if snow has piled up on the car shown in FIG. 7(d) instead of the car shown in FIG. 8(d). In this way, when multiple pieces of environmental information and spatial information are obtained, the virtual subject is processed in S408 to correspond to each piece of information.

[0037] As described above, according to the first embodiment, a virtual subject that reflects environmental information or spatial information of the real space is generated, and a superimposed image that is consistent with the real space is generated and displayed as a live view image, making it possible to consider the angle of view and composition under conditions that are close to reality.

[0038] In the first embodiment, the spatial information and the environmental information are described as being acquired, but it is also possible to acquire either one of them and process the virtual subject based on the acquired information. Even in this case, it is possible to generate a more natural superimposed image compared to the conventional method.

[0039] <Second embodiment> The second embodiment of the present invention will be described below. Note that the imaging device in the second embodiment can have the same configuration as the imaging device 100 described in the first embodiment with reference to FIG. 1, so a description thereof will be omitted here.

[0040] However, in the second embodiment, the superimposing unit 110 not only generates a superimposed image by superimposing the video data captured by the imaging unit 103 and the virtual subject processing unit 109, but also estimates the influence of the virtual subject on the video data using a learning model and performs processing to reflect this in the video data before superimposing the video data and the virtual subject. The influence of the virtual subject on the video data is estimated by inputting the video data and the virtual subject into the learning model.

[0041] The effect of a virtual subject on video data is a phenomenon that occurs due to the action of a virtual subject in real space, such as a virtual subject being reflected in the video data when there is a reflective object such as glass, a mirror, or the surface of water, or the shape of snow or the surface of water being distorted by the presence of a virtual subject.

[0042] Fig. 10 is a flowchart showing the live view display process of the imaging device 100 according to the second embodiment. In Fig. 10, the same processes as those shown in Fig. 4 are denoted by the same reference numerals, and descriptions thereof will be omitted where appropriate.

[0043] In S408, the virtual subject is processed based on the spatial information acquired in S405 and the environmental information estimated in S406. In the next step S1009, the superimposition unit 110 estimates the effect of the virtual subject processed in S408 on the live view video, and the process proceeds to S1010. In S1010, the superimposing unit 110 reflects the effect on the live-view video estimated in S1009 in the live-view video, and the process proceeds to S1011.

[0044] In S1011, the superimposing unit 110 generates a superimposed image by superimposing the live view image reflecting the influence of the virtual object in S1010 and the virtual object processed in S408, and the process proceeds to S1012. In S1012, the display unit 105 displays the superimposed video generated in S1011, and the process ends.

[0045] A specific example of the process shown in FIG. 10 will be described below with reference to FIGS. If the live view video displayed in S402 includes a puddle as shown in FIG. 11(a) and the virtual subject generated in S407 is a person as shown in FIG. 11(b), in S1009, a reflection image reflected in the puddle is estimated using a learning model based on the distance between the puddle and the virtual subject and the position of the light source, and in S1010, this is reflected in the video data so that the reflection image of the virtual subject is drawn in the puddle. Then, in S1011, the virtual subject processed in S408 is superimposed on the video data with the reflective layer drawn. The superimposed video obtained in this way and displayed in S1012 is an image in which the reflection image of the person generated in the puddle is reflected, as shown in FIG. 11(c).

[0046] Furthermore, if the live view video displayed in S402 includes accumulated snow as shown in Fig. 12(a) and the virtual subject generated in S407 is a moving car as shown in Fig. 12(b), in S1009, the snow that will be crushed by the virtual subject is predicted using a learning model, and in S1010, this is reflected in the video data so that some of the accumulated snow in the video data is crushed. Then, in S1011, the virtual subject processed in S408 is superimposed on the video data in which some of the snow has been crushed. The superimposed video obtained in this way and displayed in S1012 is an image in which the snow has been crushed in areas where the car has passed, as shown in Fig. 12(c).

[0047] Furthermore, if the live view video displayed in S402 includes rain as shown in Fig. 13(a) and the virtual subject generated in S407 is a person holding an umbrella as shown in Fig. 13(b), in S1009, the learning model estimates the area where the virtual subject blocks the rain based on the direction of the rain, and in S1010, this is reflected in the video data so that rain does not fall behind the virtual subject. Then, in S1011, the virtual subject processed in S408 is superimposed on the video data in which rain has been prevented from falling behind the virtual subject. The superimposed video obtained in this way and displayed in S1012 is an image in which rain does not fall under the person's umbrella, as shown in Fig. 13(c).

[0048] As described above, according to the second embodiment, when generating an image in which a real space and a virtual subject are superimposed, the influence of the virtual subject on the real space is estimated and reflected, thereby generating a superimposed image in which the real space and the virtual subject are consistent. This makes it possible to consider the angle of view and composition under conditions close to reality.

[0049] <Modification> In the first and second embodiments described above, the space recognition unit 106 has been described as acquiring space information such as the topography of the space and the positions of obstacles by using, for example, LiDAR.

[0050] In contrast to this, in the modified example, the space recognition unit 106, like the environment information estimation unit 107, uses a learning model to acquire space information in the video data.

[0051] As shown in FIG. 14(a), the learning model used to acquire spatial information is generated by supervised learning using training data consisting of a set of video data and ground truth data, which is spatial information such as terrain and obstacles in the corresponding video data. Specifically, learning is performed using a linear regression algorithm with the training data. Note that the items used as ground truth data are not limited to these items. Furthermore, other algorithms such as K-nearest neighbors and neural networks may be used instead of linear regression.

[0052] 14(b), the estimation of environmental information is realized by inputting video data into a learning model that has been trained using spatial information as ground truth data, and having the learning model estimate the spatial information contained in the video data. Note that the input video data is video data captured by the imaging device 100. In this embodiment, the spatial information obtained as an estimation result is the topography, obstacles, etc. in the video data.

[0053] In this modification, in S405 of Fig. 4 or 10, the spatial recognition unit 106 acquires spatial information obtained using a learning model, and thereafter performs the same processing as that described in the first and second embodiments. As a result, this modification also achieves the same effects as the first and second embodiments. Furthermore, since a light-emitting element for irradiating laser light and a detection element for detecting reflected light are not required, the configuration of the imaging device 100 can be simplified.

[0054] <Other embodiments> The present invention may be applied to a system made up of a plurality of devices, or to an apparatus made up of a single device.

[0055] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0056] <Summary> The disclosure of this embodiment includes the following configuration.

[0057] (Item 1) an acquisition means for acquiring scene information of a scene being photographed by the imaging means; A generating means for generating a virtual subject that is a virtual subject; a processing means for processing the virtual subject based on the scene information; a superimposing means for superimposing the virtual subject processed by the processing means on the image data of the scene obtained from the imaging means; 1. An image processing device comprising: (Item 2) The image processing device described in item 1, characterized in that the acquisition means estimates the scene information of the image data of the scene obtained from the imaging means using a learning model trained using the image data and the scene information. (Item 3) the acquiring means estimates, as the scene information, at least one of environmental information relating to an environment of the scene and spatial information relating to objects constituting the scene; The processing means processes the virtual subject based on at least one of the environmental information and the spatial information. 3. The image processing device according to item 2, (Item 4) the scene information includes at least one of environmental information relating to an environment of the scene and spatial information relating to objects constituting the scene; The acquisition means an estimation means for estimating the environmental information of the image data of the scene obtained from the imaging means using a learning model trained using the image data and the environmental information; a detection means for measuring the space of the scene and acquiring spatial information; and The processing means processes the virtual subject based on at least one of the environmental information and the spatial information. 2. The image processing device according to item 1, (Item 5) 5. The image processing device according to item 4, wherein the detection means measures the space by irradiating the space of the scene with laser light and detecting reflected light. (Item 6) 6. The image processing device according to any one of items 3 to 5, wherein the environmental information includes weather information. (Item 7) The weather information includes information about the position, direction, and amount of at least one of wind, rain, and snow, 7. The image processing device according to item 6, wherein the processing means reflects the influence of at least one of the wind, rain, and snow on the virtual subject. (Item 8) 8. The image processing device according to any one of items 3 to 7, wherein the spatial information includes at least one of topographical information and obstacle information. (Item 9) the topographical information includes information regarding at least one of the unevenness of the ground and the inclination of the ground, 9. The image processing device according to item 8, wherein the processing means reflects the influence of at least one of the unevenness of the ground and the inclination of the ground on the virtual subject. (Item 10) the obstacle information includes information regarding at least one of the position and size of the obstacle; 10. The image processing device according to item 8 or 9, wherein the processing means reflects the influence of the obstacle on the virtual subject. (Item 11) further comprising an estimation unit that estimates an influence of the virtual subject on the scene when the virtual subject is superimposed by the superimposing unit, 11. The image processing device according to any one of items 1 to 10, wherein the superimposing means superimposes the virtual subject processed by the processing means on the image data processed based on the influence estimated by the estimation means. (Item 12) 12. The image processing device according to any one of items 1 to 11, further comprising a display unit that displays image data obtained by superimposing by the superimposing unit. (Item 13) An image processing device according to any one of items 1 to 12, the imaging means; An imaging device comprising: (Item 14) an acquisition step of acquiring scene information of a scene being photographed by an imaging means; a generation step of generating a virtual subject that is a virtual subject; a processing step of processing the virtual subject based on the scene information; a superimposing step of superimposing the virtual subject processed in the processing step on the image data of the scene obtained from the imaging means; An image processing method comprising: (Item 15) 13. A program for causing a computer to function as each means of the image processing device according to any one of items 1 to 12. (Item 16) Item 16. A computer-readable storage medium storing the program described in item 15.

[0058] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0059] 100...imaging device, 101...CPU, 102...storage unit, 103...imaging unit, 104...image processing unit, 105...display unit, 106...space recognition unit, 107...environment information estimation unit, 108...virtual object generation unit, 109...virtual object processing unit, 110...superimposition unit, 111...communication unit, 112...system bus, 113...operation unit

Claims

1. an acquisition means for acquiring scene information of a scene being photographed by the imaging means; A generating means for generating a virtual subject that is a virtual subject; a processing means for processing the virtual subject based on the scene information; a superimposing means for superimposing the virtual subject processed by the processing means on the image data of the scene obtained from the imaging means; 1. An image processing device comprising:

2. The image processing device according to claim 1 , wherein the acquisition means estimates the scene information of the image data of the scene obtained from the imaging means using a learning model trained using the image data and the scene information.

3. the acquiring means estimates, as the scene information, at least one of environmental information relating to an environment of the scene and spatial information relating to objects constituting the scene; The processing means processes the virtual subject based on at least one of the environmental information and the spatial information.

3. The image processing device according to claim 2.

4. the scene information includes at least one of environmental information relating to an environment of the scene and spatial information relating to objects constituting the scene; The acquisition means an estimation means for estimating the environmental information of the image data of the scene obtained from the imaging means using a learning model trained using the image data and the environmental information; a detection means for measuring the space of the scene and acquiring spatial information; and The processing means processes the virtual subject based on at least one of the environmental information and the spatial information.

2. The image processing device according to claim 1, wherein:

5. 5. The image processing apparatus according to claim 4, wherein the detecting means measures the space by irradiating the space of the scene with laser light and detecting reflected light.

6. 5. The image processing device according to claim 3, wherein the environmental information includes weather information.

7. The weather information includes information about the position, direction, and amount of at least one of wind, rain, and snow, 7. The image processing device according to claim 6, wherein the processing means reflects the influence of at least one of the wind, rain, and snow on the virtual subject.

8. 5. The image processing device according to claim 3, wherein the spatial information includes at least one of topographical information and obstacle information.

9. the topographical information includes information regarding at least one of the unevenness of the ground and the inclination of the ground, 9. The image processing device according to claim 8, wherein the processing means reflects the influence of at least one of the unevenness of the ground and the inclination of the ground on the virtual subject.

10. the obstacle information includes information regarding at least one of the position and size of the obstacle; 9. The image processing device according to claim 8, wherein the processing means reflects the influence of the obstacle on the virtual subject.

11. further comprising an estimation unit that estimates an influence of the virtual subject on the scene when the virtual subject is superimposed by the superimposing unit, 2. The image processing apparatus according to claim 1, wherein the superimposing means superimposes the virtual subject processed by the processing means on the image data processed based on the influence estimated by the estimating means.

12. 2. The image processing apparatus according to claim 1, further comprising display means for displaying the image data obtained by the superimposing means.

13. The image processing device according to claim 1 ; the imaging means; An imaging device comprising:

14. an acquisition step of acquiring scene information of a scene being photographed by an imaging means; a generation step of generating a virtual subject that is a virtual subject; a processing step of processing the virtual subject based on the scene information; a superimposing step of superimposing the virtual subject processed in the processing step on the image data of the scene obtained from the imaging means; An image processing method comprising:

15. A program for causing a computer to function as each of the means of the image processing apparatus according to claim 1.

16. A computer-readable storage medium storing the program according to claim 15.

Citation Information

Patent Citations

  • Image processing apparatus and image processing method

    JP2009163610A