Information processing device, imaging apparatus, information processing method, and program

The information processing device estimates and generates virtual subject movement using a learning model to facilitate accurate photography of moving subjects by determining imaging parameters, addressing the challenge of setting parameters for non-present subjects.

JP2025187455APending Publication Date: 2025-12-25CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024096267
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2025-12-25

AI Technical Summary

Technical Problem

Existing photography techniques struggle to set appropriate shooting parameters for a moving subject when it is not present, as the virtual subject is fixed, making it difficult to accurately focus and adjust other photography settings.

Method used

An information processing device estimates motion information of a virtual subject using a learning model based on captured images, generates a virtual subject image with associated movement information, and determines imaging parameters such as focus position, F-number, and shutter speed to facilitate accurate photography of moving subjects.

Benefits of technology

Enables the setting of appropriate shooting parameters for moving subjects even when they are not present, improving the accuracy and ease of photography by allowing users to visualize and adjust settings based on virtual subject movement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025187455000001_ABST
    Figure 2025187455000001_ABST
Patent Text Reader

Abstract

To solve the problem in which it is difficult to appropriately set shooting parameters for a moving subject when a subject does not exist.SOLUTION: An information processing device includes estimation means for estimating motion information indicating the movement of a virtual subject virtualized on the basis of a captured image, and generation means for generating a virtual subject on the basis of characteristic information indicating the characteristics of the virtual subject, associating the image of the virtual subject with the motion information of the virtual subject, and generating a virtual image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an imaging device, an information processing method, and a program. [Background technology]

[0002] One known photography technique is to set the focus position (in focus) in advance to a location where the subject is predicted to arrive before the subject arrives within the angle of view, and then photograph the subject when it arrives at that location (hereinafter referred to as pre-focus photography).In pre-focus photography, the focus position and other photography parameters are set when the subject is not present, making it difficult to set them appropriately.

[0003] Patent Document 1 discloses a technology in which a virtual subject is generated at a placement position set by a user, distance information to the virtual subject is acquired, and then focusing is performed based on the acquired distance information. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-251801 Summary of the Invention [Problem to be solved by the invention]

[0005] However, in the above-mentioned Patent Document 1, the virtual subject is fixed, so it is difficult to set appropriate shooting parameters for a moving subject in a state where the subject does not exist.

[0006] Therefore, the present invention provides a technique that can appropriately set shooting parameters for a moving subject even in a situation where the subject is not present. [Means for solving the problem]

[0007] In order to solve this problem, for example, an information processing device of the present invention has the following configuration: an estimation means for estimating motion information indicating a motion of a virtual subject obtained by virtualizing a subject based on a captured image; a generation means for generating the virtual subject based on feature information indicating features of the virtual subject, and for generating a virtual image by associating an image of the virtual subject with movement information of the virtual subject; Equipped with. [Effects of the Invention]

[0008] According to the present invention, it is possible to set appropriate shooting parameters for a moving subject even in a situation where the subject is not present. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 2 is a block diagram illustrating the overall configuration of a control system of the imaging apparatus according to the embodiment. [Figure 2] 5A to 5C are diagrams illustrating a process of estimating motion information in a learning model unit of the estimation unit. [Figure 3] 10A and 10B are conceptual diagrams illustrating a method for generating an image by superimposing a virtual subject image generated by a generating unit on a captured image. [Figure 4] FIG. 4 is a flowchart showing processing for determining shooting parameters according to the first embodiment. [Figure 5] 10A and 10B are conceptual diagrams illustrating the display of a moving image in which a virtual subject image is superimposed on a captured image according to the embodiment. [Figure 6] 5A and 5B are conceptual diagrams illustrating a method for determining imaging parameters according to an embodiment. [Figure 7] FIG. 10 is a flowchart showing processing for determining shooting parameters according to the second embodiment. [Figure 8] FIG. 11 is an image diagram of a moving image in which a plurality of virtual subject images are superimposed according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0011] (First embodiment) The embodiments relate to an imaging device and an information processing device mounted on the imaging device. In particular, the embodiments relate to an imaging device and an information processing device that determine shooting parameters by using a learning model to generate an image in which an image of a virtual subject having motion information is superimposed on a captured image. The virtual subject is a virtual object created by virtualizing a subject. The term "image" may include still images, videos, images, and data thereof.

[0012] Fig. 1 is a block diagram illustrating the overall configuration of a control system of an image capturing apparatus according to an embodiment. Fig. 1(a) is a block diagram illustrating the hardware configuration of the image capturing apparatus according to an embodiment. Hereinafter, an image capturing apparatus 100 according to a first embodiment will be described with reference to Fig. 1(a).

[0013] The imaging device 100 has a lens unit 1001, a lens control unit 1002, a CPU 1003, a CPU bus 1004, a RAM bus 1005, a ROM 1006, a display control unit 1007, an evaluation unit 1008, a recording control unit 1009, a development unit 1010, an imaging unit 1011, a RAM control unit 1012, a RAM 1013, an operation unit 1014, a communication unit 1015, a display unit 1016, and a recording unit 1017.

[0014] The lens unit 1001 is a group of lenses including a zoom lens and a focus lens, and guides light from a subject to the imaging unit 1011.

[0015] The lens control unit 1002 has a lens control processing function for controlling the focal length and aperture state of the lens unit 1001 based on information about the subject obtained by the imaging unit 1011, which will be described later.

[0016] The CPU 1003 is an abbreviation for Central Processing Unit and is a type of processor also known as a central processing unit. The CPU 1003 performs overall control of the image capturing apparatus 100. The CPU 1003, for example, reads out computer programs stored in the ROM 1006 and the recording unit 1017 and loads the programs into the RAM 1013, thereby implementing various functions and controlling the image capturing apparatus 100. Instead of or in addition to the CPU 1003, the image capturing apparatus 100 may include other processors such as an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), and a QPU (Quantum Processing Unit).

[0017] A CPU bus 1004 performs communication between the CPU 1003 and each functional block.

[0018] The RAM bus 1005 performs communication between a RAM control unit 1012 (to be described later) and each functional block. The RAM bus 1005 also has a function of arbitrating access from each functional block to a RAM 1013 (to be described later).

[0019] ROM 1006 is an abbreviation for Read Only Memory, and is a read-only non-volatile storage medium. The ROM 1006 stores programs such as firmware, learning models, and the like.

[0020] The display control unit 1007 performs predetermined display processing on image data and video data developed by a development unit 1010 (described later). The display control unit 1007 outputs the data after display processing to a display unit 1016 (described later) and executes display control processing to display the image.

[0021] The evaluation unit 1008 has a function of performing evaluation value calculation processing to calculate evaluation values ​​for the focus state, exposure state, and the like based on image data and video data generated by an image capturing unit 1011 (described later).

[0022] The recording control unit 1009 controls the recording unit 1017 (described later) in accordance with instructions from the CPU 1003 via the CPU bus 1004 and RAM bus 1005 .

[0023] The development unit 1010 performs Bayer processing on the image data and video data generated by the imaging unit 1011 (described later) to convert them into signals consisting of luminance signals and color difference signals, and has development processing functions such as removing noise contained in each signal, correcting optical distortion, and optimizing the image.

[0024] The imaging unit 1011 converts optical signals captured by the lens unit 1001 into electrical signals using an imaging sensor. The imaging sensor may be a CMOS (Complementary Metal-Oxide-Semiconductor) sensor, a CCD (Charge Coupled Device) sensor, or the like. The imaging unit 1011 performs correction processing on the obtained image data and video data to correct lens aberrations, interpolation processing to interpolate defective pixels of the imaging sensor, and the like. The imaging unit 1011 outputs the image data that has undergone correction processing, etc., via buses 1004 and 1005.

[0025] A RAM control unit 1012 controls access to a RAM 1013 (described later) based on a RAM access request from each functional block.

[0026] The RAM 1013 is an abbreviation for Random Access Memory, and is a volatile storage medium that allows high-speed reading and writing of information. The RAM 1013 is used as a working area when the CPU 1003 processes information.

[0027] The operation unit 1014 receives various settings, operations, and instructions from the user for the imaging device 100. The operation unit 1014 has a touch panel, physical buttons, etc. The operation unit 1014 may also have an external keyboard, etc.

[0028] The communication unit 1015 is a communication interface that connects the imaging device 100 to other devices via a wired or wireless connection and transmits and receives image files, video files, etc. The communication unit 1015 can also be connected to a network such as the Internet via a wireless LAN (Local Area Network) or a wired LAN.

[0029] The display unit 1016 displays an image based on image data or video data. The display unit 1016 is, for example, an image display device such as a liquid crystal display.

[0030] The recording unit 1017 is a non-volatile storage medium from / to which information can be read and written. The recording unit 1017 may be a hard disk drive (HDD) or a solid state drive (SSD). The recording unit 1017 stores image data and video data obtained by the imaging unit 1011. The recording unit 1017 may also store programs and learning models.

[0031] 1(b) is a block diagram illustrating the functional configuration of an information processing device installed in the imaging device. The information processing device 1020 is, for example, a computer. The information processing device 1020 includes, as hardware, the above-mentioned CPU 1003, ROM 1006, recording control unit 1009, RAM control unit 1012, RAM 1013, buses 1004 and 1005, etc. The information processing device 1020 has the functions of an input unit 1021, an estimation unit 1018, a generation unit 1019, and a determination unit 1022.

[0032] Some or all of the functions of the input unit 1021, the estimation unit 1018, the generation unit 1019, and the determination unit 1022 may be functions that are realized by one or more processors such as the CPU 1003 reading out a program stored in the ROM 1006 or the recording unit 1017 and loading it into the RAM 1013. Some or all of the functions of the input unit 1021, the estimation unit 1018, the generation unit 1019, and the determination unit 1022 may be realized by one or more circuits such as an ASIC (Application Specific Integrated Circuit) and a PLD (Programmable Logic Device) including an FPGA (Field Programmable Gate Array).

[0033] The input unit 1021 accepts input from the user. The input unit 1021 may accept input from the user via the operation unit 1014, such as a touch panel or a keyboard. The input unit 1021 accepts, for example, feature information of a virtual subject from the user. The input unit 1021 accepts an instruction to import a captured image. The input unit 1021 passes the accepted feature information of the virtual subject and the import instruction to the estimation unit 1018.

[0034] The estimation unit 1018 estimates motion information indicating the motion of the virtual subject based on the captured image. For example, the estimation unit 1018 uses feature information of the virtual subject and the captured image as input data, and estimates the motion of the virtual subject using a learning model unit having one or more learning models stored in the ROM 1006 or the recording unit 1017, to generate motion information of the virtual subject. Note that the captured image acquired by the estimation unit 1018 may be an image captured by the imaging unit 1011. For example, the captured image may be a real-time image (video) captured by the imaging unit 1011.

[0035] 2 is a diagram illustrating the flow of processing for estimating motion information in the learning model unit of the estimation unit. A method for estimating motion information of a virtual subject and a learning model stored in ROM 1006 or recording unit 1017 will be described with reference to FIG.

[0036] The learning model unit 21 of the estimation unit 1018 estimates the motion information VM of the virtual subject using the captured image IM and the feature information SF of the virtual subject as input data. The feature information SF of the virtual subject includes at least the type of the subject and may also include other information about the subject, such as size information of the subject. The input format of the feature information SF of the virtual subject may be text, an image such as a still image, or a video, and the input format is not particularly important. For example, for a captured image IM of a railroad track that does not include a train, which is the main subject, the user inputs the text "train" as the type, which is the feature information SF of the virtual subject. The input captured image IM is assumed to be a captured image obtained by the imaging unit 1011. The motion information VM of the virtual subject obtained as an estimation result includes information such as the speed and direction of movement of the virtual subject.

[0037] FIG. 2(a) is a diagram illustrating the learning model unit of the first embodiment. The learning model unit 21 of the first embodiment will be described in more detail using FIG. 2(a). The learning model unit 21 of the first embodiment can be used in the first pattern of FIG. 2(a). The learning model unit 21 has a two-stage structure. Specifically, the learning model unit 21 has a first learning model 22 and a second learning model 23 that is subsequent to the first learning model 22.

[0038] In the learning model unit 21 of the first pattern shown in Figure 2(a), the first learning model 22 uses a captured image IM as input data to estimate shooting position information SP, which is information about the location of the imaging device. The first learning model 22 is a learning model that learns from training data a group of image data and a group of video data, including shooting position information, taken at various locations on Earth and obtained from artificial satellites, surveillance cameras, and personally owned imaging devices. In this way, the first learning model 22 estimates the shooting position information SP from the captured image IM.

[0039] The second learning model 23 estimates the movement information VM of the virtual subject using the shooting position information SP and feature information SF of the virtual subject output by the first learning model 22 as input. The second learning model 23 is a learning model trained using image data and video data, including location and time information, obtained from satellites, surveillance cameras, and personally owned imaging devices as training data to estimate what type of subject may appear at what time and at what speed at any location on Earth. Specific learning algorithms may include one or more algorithms, including nearest neighbor algorithms, naive Bayes algorithms, decision trees, and support vector machines. Another example of an algorithm is deep learning, which uses a neural network to generate features and connection weighting coefficients for learning. Available algorithms from the above may be used as appropriate and applied to this embodiment.

[0040] In this way, the estimation unit 1018 estimates the movement information VM of the virtual subject from the captured image IM and the feature information SF of the virtual subject using the learning model unit 21, which is composed of two stages: the first learning model 22 and the second learning model 23.

[0041] Furthermore, if the imaging device 100 has a position information acquisition unit capable of acquiring the position of the imaging device 100, such as a Global Navigation Satellite System (GNSS), the learning model unit 21 may acquire the shooting position information SP without using the first learning model 22. Note that a series of processes for estimating the motion information VM of the virtual subject and the feature information VF of the virtual subject using the learning model unit 21 shown in Fig. 2 may be referred to as generation AI (Artificial Intelligence) in this embodiment.

[0042] The estimation unit 1018 passes the motion information and imaging position information of the generated virtual subject to the generation unit 1019 together with the acquired feature information and captured image of the virtual subject.

[0043] The generation unit 1019 generates a virtual subject by virtualizing a subject based on the feature information of the virtual subject and the movement information of the virtual subject acquired from the estimation unit 1018. Furthermore, the generation unit 1019 generates a virtual subject image, which is an image of the virtual subject having movement information. The generation unit 1019 generates an image in which the generated virtual subject image and the movement information of the virtual subject are associated with each other, as a virtual image. For example, the generation unit 1019 generates a virtual image (e.g., a video) in which the virtual subject image and the movement information of the virtual subject are superimposed on a captured image. The generation unit 1019 displays the generated virtual image on the display unit 1016 via the display control unit 1007. Note that a virtual image may also be referred to as an "image."

[0044] 3 is an image diagram illustrating a method for generating a virtual image by superimposing a virtual subject image on a captured image by the generation unit 1019. The method for generating a virtual subject image having movement information by the generation unit 1019 will be described with reference to FIG.

[0045] 3A is a diagram illustrating the acquisition of distance information within a captured image. As shown in FIG. 3A, the generation unit 1019 acquires distance information to an object appearing in the captured image based on an imaging position (0,0,0) that is the position of the imaging device 100. The generation unit 1019 may acquire the distance information using a distance measurement sensor such as a TOF (Time-Of-Flight) sensor mounted on the imaging device 100 (not shown). The generation unit 1019 may also acquire the distance information by comparing the imaging position information of the imaging device 100 acquired by the above-described position information acquisition method with map data obtained using the communication unit 1015.

[0046] As a result, as shown in FIG. 3(a), the generation unit 1019 acquires position information (Xa, Ya, Za) of the left end of the track and position information (Xb, Yb, Zb) of the right end of the track in the captured image. Note that the position information Xa, Ya and the position information Xb, Yb indicate position information on a plane (for example, a horizontal plane). The position information Za and the position information Zb indicate position information in the height direction. By acquiring position information for all pixels in the captured image, the generation unit 1019 can acquire position information for all of the subjects included in the captured image.

[0047] Next, as shown in FIG. 3(b), the generation unit 1019 displays a virtual image generated by superimposing a virtual subject image on the captured image. The generation unit 1019 generates a virtual subject image by calculating position information (X1, Y1, Z1) of the front of the train and position information (X2, Y2, Z2) of the rear of the train as position information for placing the train, which is the virtual subject image, from the track position information obtained in FIG. 3(a). Note that the position information X1, Y1 and the position information X2, Y2 indicate position information on a plane (e.g., a horizontal plane), and the position information Z1 and the position information Z2 indicate position information in the height direction. The generation unit 1019 associates the virtual subject image with the motion information of the virtual subject image estimated using the learning model described above, thereby displaying the virtual subject image on a video having time information. For example, the generation unit 1019 may generate a video including a virtual subject image moving on the captured image by calculating the position information of the front end of the vehicle (X1, Y1, Z1) and the position information of the rear end of the vehicle (X2, Y2, Z2) for each frame of the video based on the movement information.

[0048] The generation unit 1019 may superimpose an image related to the movement information on the captured image. Specifically, as shown in Fig. 3(b), the generation unit 1019 may display the direction in which the virtual subject image moves with a hollow arrow superimposed on the captured image, and may display the speed of the virtual subject image with a speech bubble containing text (here, 80 km / h) superimposed on the captured image. The generation unit 1019 displays the speed of the virtual subject image as 80 km / h, but may correct the displayed speed based on information such as the type of train, operation information, and the degree of curvature of the tracks.

[0049] The determination unit 1022 determines the imaging parameters. For example, the determination unit 1022 may determine the imaging parameters based on movement information of the virtual subject. The imaging parameters are parameters related to imaging of the subject, including the focus position (in focus), F-number, aperture value, shutter speed, etc. The determination unit 1022 stores the determined imaging parameters in either the recording unit 1017 or the RAM 1013. Here, the determination unit 1022 may determine the imaging parameters by accepting one or more positions (arrangement positions) of the virtual subject from the user. For example, the determination unit 1022 may accept one or more arrangement positions of the virtual subject from the user and determine multiple arrangement positions based on the accepted arrangement positions. Furthermore, when the determination unit 1022 determines multiple arrangement positions, the determination unit 1022 may store the imaging parameters associated with each of the arrangement positions in the recording unit 1017, etc.

[0050] Fig. 4 is a flowchart showing the processing for determining shooting parameters according to the first embodiment. Referring to Fig. 4, the processing performed by the information processing device 1020 of the imaging device 100 from when the user inputs feature information of the virtual subject to when the information processing device 1020 generates a virtual image on which a virtual subject having movement information is superimposed, and when the imaging parameters are determined, will be described. Each step of the flowchart in Fig. 4 is executed by the CPU 1003 reading a program for processing to determine shooting parameters when the power switch (not shown) of the imaging device 100 is on.

[0051] In step S401, the input unit 1021 accepts feature information of the virtual subject input by the user. For example, the user inputs feature information of the virtual subject via the operation unit 1014. The feature information of the virtual subject is, for example, the type of subject, "train."

[0052] In step S402, the input unit 1021 instructs the estimation unit 1018 to capture a captured image.

[0053] In step S403, upon receiving an instruction to import from the input unit 1021, the estimation unit 1018 acquires a captured image and also acquires feature information of the virtual subject input by the user in step S401 from the input unit 1021. Here, the estimation unit 1018 may acquire a captured image in real time while the imaging unit 1011 is capturing images. The estimation unit 1018 may acquire the captured image not from the imaging unit 1011 but from the recording unit 1017.

[0054] In step S404, the estimation unit 1018 estimates the movement information of the virtual subject. Specifically, the estimation unit 1018 acquires a learning model stored in the ROM 1006 or the recording unit 1017. The estimation unit 1018 estimates the movement information of the virtual subject using the learning model that receives as input the feature information of the virtual subject and the captured image.

[0055] In step S405, the generation unit 1019 generates a virtual subject image having motion information. For example, the generation unit 1019 generates the virtual subject image based on feature information of the virtual subject. The generation unit 1019 generates the virtual subject image having motion information by associating the motion information of the virtual subject estimated by the estimation unit 1018 with the virtual subject image.

[0056] In step S406, the generation unit 1019 generates a moving image as a virtual image by superimposing the virtual subject image having the movement information generated in step S405 on the captured image acquired in step S402, and displays the generated moving image on the display unit 1016.

[0057] Fig. 5 is an image diagram illustrating the display of a video in which a virtual subject image generated by the generation unit 1019 is superimposed on a captured image. Fig. 5 is an image diagram of a video when "train," which is feature information of the virtual subject, is input as a prompt input to the learning model unit, which is the generation AI. Note that the processing of the generation AI refers to a series of processes that utilize the learning model shown in Fig. 2 described above to estimate the motion information VM of the virtual subject and the feature information VF of the virtual subject.

[0058] The method for generating the virtual subject image is as explained using Fig. 3. Images 500, 501, 502, 503, and 504 shown in Fig. 5 are images taken during a video in which the virtual subject image superimposed on the captured image moves, and represent frames of the video. As shown in Fig. 5, the generation unit 1019 generates a video including an image of a train, which is the virtual subject image moving on the captured image, as virtual images 500, 501, 502, 503, and 504, and displays them on the display unit 1016.

[0059] In step S407, the determination unit 1022 acquires the user's selection of the placement position of the virtual subject. For example, the user uses the operation unit 1014 to select an image that represents the desired placement position of the virtual subject from the video in which the virtual subject image having movement information is moving, which was displayed in step S406. As a selection method, the user may stop the video when the virtual subject image reaches the desired position during playback and select the placement position of the virtual subject. Alternatively, the user may select the placement position of the virtual subject in a video that is being rewound, fast-forwarded, or frame-by-frame played.

[0060] In step S408, the determination unit 1022 determines shooting parameters. For example, the determination unit 1022 may determine the shooting parameters based on information about the shooting parameters (which may be the shooting parameters themselves, for example) input by a user viewing a video in which the virtual subject image having the motion information generated in step S407 is moving. The user may input shooting parameters for shooting a subject in accordance with the virtual subject image including the motion information. Furthermore, the determination unit 1022 may accept some of the shooting parameters (for example, shutter speed) from the user and determine the remaining shooting parameters (for example, aperture value) based on the motion information and the input shooting parameters, etc.

[0061] Fig. 6 is an image diagram illustrating a method for determining shooting parameters by the determination unit 1022 based on user input. Fig. 6(a) is a plan view of a shooting area illustrating the arrangement of virtual subjects. Fig. 6(b) is a diagram of multiple frame images on which virtual subject images at each arrangement position are superimposed.

[0062] As shown in FIG. 6(a), it is assumed that the user refers to the image on which the virtual subject image is superimposed and selects four placement positions, and the determination unit 1022 accepts the selection. Here, it is assumed that four locations indicated by (X0, Y0, Z0), (X1, Y1, Z1), (X2, Y2, Z2), and (X3, Y3, Z3) are selected as placement positions. Note that in step S407, the determination unit 1022 may accept only the main placement position (X0, Y0, Z0) from the user, and automatically set the remaining past placement positions based on preset time intervals, position intervals, frame intervals, etc. The determination unit 1022 may also set placement positions that are further in the future than the main placement position.

[0063] The "main" image on the right side of FIG. 6(b) is an image in which a virtual subject image is superimposed on a captured image at the placement position selected by the user as the main in step S407. The determination unit 1022 can display an image in which a focused virtual subject image is superimposed on the main image at the frame immediately before (frame -1 in the figure), the frame immediately before (frame -2 in the figure), or the like, depending on the setting status of the shooting parameters. The virtual subject image and the image to be displayed may be generated by the generation unit 1019. This allows the user to easily set optimal shooting parameters for desired shooting by referring to the image in which the virtual subject image is superimposed at each placement position. The determination unit 1022 determines the shooting parameters based on the information on the placement positions and the user's settings. Here, the determination unit 1022 may determine shooting parameters in association with each of a plurality of placement positions. In this case, the determination unit 1022 may store a table in the recording unit 1017 that associates the placement positions with the shooting parameters.

[0064] As described above, according to the first embodiment, a virtual image such as a video is generated by associating a virtual subject with the movement of the virtual subject, so that a user can easily set appropriate shooting parameters for a moving subject based on the virtual image, even in a situation where the subject is not present.

[0065] For example, by using a learning model such as a generation AI to estimate the type of subject, movement information, and placement position from the captured image and characteristic information of the virtual subject, and then setting shooting parameters for the moving virtual subject image generated from the estimation results, it is possible to accurately reflect the user's intentions and reduce shooting errors, resulting in a suitable captured image, without requiring special know-how, even when using shooting methods such as fixed focus shooting.

[0066] In the first embodiment, the estimation unit 1018 estimates the motion information of the virtual subject based on the feature information of the virtual subject, and therefore the estimation accuracy of the motion information can be improved.

[0067] In the first embodiment, the input unit 1021 receives feature information of the virtual subject from the user, and therefore it is possible to improve the accuracy of identifying the type of virtual subject, etc.

[0068] In the first embodiment, the estimation unit 1018 estimates the direction and speed of movement of the virtual subject, so the generation unit 1019 can make the movement of the virtual subject to be superimposed closer to that of the actual subject.

[0069] In the first embodiment, the generation unit 1019 displays the direction and speed of movement of the virtual subject together with the image of the virtual subject, allowing the user to imagine the movement of the subject with greater accuracy.

[0070] In the first embodiment, the determination unit 1022 determines the shooting parameters based on the user's input received while a virtual subject image having movement information is displayed. This allows the user to easily and appropriately determine the shooting parameters by imagining a moving subject even in a situation where the subject does not exist.

[0071] In the first embodiment, the determination unit 1022 acquires the placement position of the virtual subject from the user and determines the shooting parameters based on the information on the placement position, so that it is possible to determine appropriate shooting parameters according to the placement position.

[0072] In the first embodiment, the determining unit 1022 accepts the placement position of the virtual subject while the virtual subject image is displayed, allowing the user to select a more appropriate placement position.

[0073] In the first embodiment, the determination unit 1022 determines the shooting parameters for each of the plurality of arrangement positions, and stores the arrangement positions and the shooting parameters in association with each other in the recording unit 1017, etc. This allows the user to more appropriately photograph the subject at the plurality of arrangement positions.

[0074] (Second embodiment) An imaging device according to the second embodiment will be described below. In the first embodiment, an imaging device in which a user inputs feature information of a virtual subject was described. The imaging device according to the second embodiment does not require feature information of a virtual subject input by a user, but automatically estimates feature information of candidate virtual subjects, generates virtual subjects, and displays them superimposed on a captured image.

[0075] The configuration of the second embodiment is basically the same as that of the imaging device 100 described in the first embodiment, so a description of the same parts will be omitted and only the functions of the different parts will be described in detail.

[0076] FIG. 2(b) is a diagram illustrating the learning model unit of the estimation unit of the second embodiment. First, the function of the learning model unit of the estimation unit will be described with reference to FIG. 2(b). As shown in FIG. 2(b), the second pattern learning model unit 21 of the second embodiment may take the captured image IM as input and generate motion information VM of the virtual subject and feature information VF of the virtual subject when feature information SF of the virtual subject is not input. Specifically, the first learning model 22 takes the captured image IM as input and generates and outputs shooting position information SP. The second learning model 23 takes the shooting position information SP and the captured image IM as input and generates and outputs motion information VM of the virtual subject and feature information VF of the virtual subject. The feature information VF of the virtual subject includes at least the type of virtual subject and may also include other information about the virtual subject, such as size information of the virtual subject. Furthermore, the second learning model 23 may estimate the feature information VF of the virtual subject and then estimate the motion information VM of the virtual subject.

[0077] 7 is a flowchart showing the process of determining the shooting parameters in the second embodiment. The process of determining the shooting parameters in the second embodiment will be described with reference to FIG.

[0078] In step S701, the input unit 1021 acquires an instruction to import a captured image that is input by the user operating the operation unit 1014. The input unit 1021 transfers the instruction to the estimation unit 1018.

[0079] In step S702, in response to the instruction in step S701, the estimation unit 1018 acquires a captured image from the imaging unit 1011. For example, the estimation unit 1018 may acquire a captured image in real time while the imaging unit 1011 is capturing images.

[0080] In step S703, the estimation unit 1018 estimates feature information and movement information of a candidate virtual subject. In this embodiment, the estimation unit 1018 may estimate feature information and movement information of multiple virtual subjects. For example, the estimation unit 1018 acquires a learning model stored in the ROM 1006 or the recording unit 1017. The estimation unit 1018 uses the captured image acquired in step S702 as input and estimates feature information and movement information of the virtual subject by using the learning model. Assume that the captured image includes railroad tracks and roads. In this case, the estimation unit 1018 may estimate the "type" of the feature information of the virtual subject to be "train" based on the railroad tracks. Furthermore, the estimation unit 1018 may estimate the "type" of the feature information of the virtual subject to be "automobile" or "person" based on the roads. The estimation unit 1018 may estimate the movement information of each virtual subject based on the "type" of the estimated feature information. In this case, since the moving direction, position, and speed of the virtual subjects "train," "car," and "person" are different, the estimation unit 1018 estimates the motion information for each virtual subject taking these into consideration. Therefore, when the estimation unit 1018 estimates the motion information for each virtual subject, it generates the motion information in association with the virtual subject.

[0081] In step S704, the generation unit 1019 generates a virtual subject image having motion information based on the estimation results of the feature information and motion information of the virtual subject obtained in step S703. If the estimation unit 1018 has estimated multiple virtual subjects, the generation unit 1019 generates multiple virtual subject images based on the respective estimation results.

[0082] In step S705, the generation unit 1019 generates a moving image as a virtual image by superimposing the virtual subject image having the movement information generated in step S704 on the captured image acquired in step S702, and causes the display unit 1016 to display the virtual image.

[0083] FIG. 8 is a conceptual diagram of a video in which a virtual subject image generated by the generation unit 1019 is superimposed. The image 800 displays a train as a candidate for the virtual subject image. As shown in FIG. 8, the image 800 also displays a UI (User Interface) screen including buttons for "Type," "Position," "Size," "Orientation," and "Movement" that allow the user to select and instruct changes to information about the virtual subject image, and a "Confirm" button that instructs the user to confirm the virtual subject image. This enables the determination unit 1022 to accept selections related to the virtual subject through UI operations from the user. The user can select and press or touch one of the "Type," "Position," "Size," or "Orientation" buttons to enable or disable the corresponding information change. This allows the user to specify which parameters to change.

[0084] In step S706, the determination unit 1022 determines whether the virtual subject image on the displayed moving image is the virtual subject image desired by the user. For example, if the user selects the "Type" button, the determination unit 1022 determines that it is not the desired virtual subject image, and proceeds to step S707.

[0085] In step S707, the generation unit 1019 displays a moving image in which another virtual subject image is superimposed on the captured image based on the user's selection. For example, while the generation unit 1019 is displaying a moving image in which a virtual subject image is superimposed on the captured image in step S705, the generation unit 1019 refers to another estimation result obtained by the estimation unit 1018 in step S703 and displays a new moving image in which another virtual subject image is superimposed on the captured image on the display unit 1016.

[0086] Specifically, suppose that in step S705, the user presses the "type" button to issue an instruction to change the virtual subject image while the generation unit 1019 is displaying a moving image in which a virtual subject image of a "train" is superimposed, as in image 800 shown in Fig. 8. In this case, the generation unit 1019 may display a new moving image in which a "person," which is a different type of virtual subject, is superimposed on the captured image as a virtual subject image, as in image 801 in Fig. 8. Note that the generation unit 1019 may correct the speed, which is movement information of the virtual subject image, based on the appearance, age, etc. of the person.

[0087] In step S706, if the user selects the "OK" button, the determination unit 1022 determines that the virtual subject image superimposed on the video is the virtual subject image desired by the user, determines the type of feature information, and proceeds to step S708.

[0088] The determination unit 1022 executes the processes of steps S708 and S709 based on the type of feature information of the determined virtual subject. The processes of steps S708 to S709 are similar to the processes of steps S407 to S408 in the first embodiment, and therefore will not be described again.

[0089] As described above, according to the second embodiment, the imaging device 100 estimates feature information of a virtual subject and automatically generates and displays an image of the virtual subject for the user, thereby saving the user the trouble of inputting the feature information of the virtual subject.

[0090] In the second embodiment, the generation unit 1019 displays an image including buttons for selecting feature information of a virtual subject, so that the user can more easily input feature information of a virtual subject.

[0091] In all of the above-described embodiments, it is assumed that the shooting conditions are set in accordance with the virtual subject image displayed on the display unit 1016 and subsequent shooting instructions are given manually by the user, but shooting may also be performed automatically by the imaging device 100. In this case, shooting may be performed automatically when the imaging device 100 determines that an actual subject similar to the virtual subject image has reached the position where the virtual subject image is displayed.

[0092] In all of the above-described embodiments, the virtual subject image is displayed superimposed on the captured image acquired when the virtual subject image is generated. However, it is also possible that the user may move the imaging device 100 after the virtual subject image is displayed superimposed, changing the imaging angle of view. In this case, the image on which the virtual subject image is displayed superimposed may be the current captured image, rather than the captured image acquired when the virtual subject image is generated. In this case, the position and size of the virtual subject image may be adjusted so that they change in real time to match the current captured image, and then displayed superimposed.

[0093] In all of the above-described embodiments, the motion information of the virtual subject image and the feature information of the virtual subject are estimated using a learning model stored in ROM 1006, but the estimation may also be performed by using the communication unit 1015 to communicate with an external server of the imaging device 100 and using a learning model stored in the external server.

[0094] In the above-described embodiment, the motion information of the virtual subject is estimated using a learning model, but the estimation of the motion information is not limited to this. For example, the estimation unit may estimate the motion information of the virtual subject based on motion information input by a user. Specifically, while a still image in which the image of the virtual subject is superimposed on the captured image is displayed, the user may input the motion information of the virtual subject by sliding a finger or the like on a touch panel, and the estimation unit may estimate the motion information of the virtual subject based on the input.

[0095] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0096] The disclosure of this specification includes the following information processing device, imaging device, information processing method, and program. (Item 1) an estimation means for estimating motion information indicating a motion of a virtual subject obtained by virtualizing a subject based on a captured image; a generation means for generating the virtual subject based on feature information indicating features of the virtual subject, and for generating a virtual image by associating an image of the virtual subject with movement information of the virtual subject; An information processing device comprising: (Item 2) The estimation means estimates the movement information based on feature information of the virtual subject. 2. The information processing device according to item 1, (Item 3) an input means for receiving characteristic information of the virtual subject input by a user; 3. The information processing device according to item 1 or 2, comprising: (Item 4) The feature information of the virtual subject includes the type of the virtual subject. 4. The information processing device according to any one of items 1 to 3, (Item 5) The estimation means estimates the movement information including at least a direction in which the virtual subject moves and a speed at which the virtual subject moves. 5. The information processing device according to any one of items 1 to 4, wherein: (Item 6) The generating means generates and displays an image including an image of the virtual subject and an image showing movement information associated with the virtual subject. 6. The information processing device according to any one of items 1 to 5, (Item 7) The generating means generates a screen for receiving a selection regarding the virtual subject from a user. 7. The information processing device according to any one of items 1 to 6, (Item 8) The generating means displays an image of the virtual subject based on the selection received from the user. 8. The information processing device according to item 7, (Item 9) A determination means for determining photographing parameters relating to photographing the subject. 9. The information processing device according to any one of items 1 to 8, comprising: (Item 10) The determining means determines the shooting parameters based on the movement information of the virtual subject. 10. The information processing device according to item 9, (Item 11) The determining means determines the photographing parameters based on information input by a user. 11. The information processing device according to item 9 or 10, (Item 12) The determining means obtains the placement position of the virtual subject from a user, and determines the shooting parameters based on information about the placement position. 12. The information processing device according to any one of items 9 to 11, (Item 13) the generating means generates a virtual image by superimposing an image of the virtual subject on the captured image, The determining means acquires information about the placement position from the user while the virtual image is being displayed. Item 13. The information processing device according to item 12. (Item 14) The determining means determines a plurality of placement positions and determines imaging parameters associated with each of the plurality of placement positions. 14. The information processing device according to item 12 or 13, (Item 15) The estimation means estimates movement information of the virtual subject using a learning model. 15. The information processing device according to any one of items 1 to 14, (Item 16) the estimation means estimates feature information of the virtual subject using a learning model; The generating means generates the virtual subject based on the estimated feature information of the virtual subject. 16. The information processing device according to any one of items 1 to 15, (Item 17) the estimation means acquires real-time captured images and estimates movement information of the virtual subject; The generating means displays an image in which an image of the virtual subject is superimposed on the captured image. 17. The information processing device according to any one of items 1 to 16, (Item 18) the information processing device according to item 1; an imaging means for capturing an image of the subject and generating an image; An imaging device comprising: (Item 19) an estimation step of estimating motion information indicating a motion of a virtual subject obtained by virtualizing a subject based on a captured image; a generation step of generating the virtual subject based on feature information indicating features of the virtual subject, and associating an image of the virtual subject with movement information of the virtual subject to generate a virtual image; An information processing method comprising: (Item 20) A program for causing a computer to function as each means of the information processing device according to any one of items 1 to 17.

[0097] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0098] 100: Imaging device; 1020: Information processing device; 1021: Input unit; 1018: Estimation unit; 1019: Generation unit; 1022: Determination unit; 21: Learning model unit; 22: First learning model; 23: Second learning model.

Claims

1. an estimation means for estimating motion information indicating a motion of a virtual subject obtained by virtualizing a subject based on a captured image; a generation means for generating the virtual subject based on feature information indicating features of the virtual subject, and for generating a virtual image by associating an image of the virtual subject with movement information of the virtual subject; An information processing device comprising:

2. The estimation means estimates the movement information based on feature information of the virtual subject.

2. The information processing apparatus according to claim 1, wherein:

3. an input means for receiving characteristic information of the virtual subject input by a user; 2. The information processing apparatus according to claim 1, further comprising:

4. The feature information of the virtual subject includes the type of the virtual subject.

2. The information processing apparatus according to claim 1, wherein:

5. The estimation means estimates the movement information including at least a direction in which the virtual subject moves and a speed at which the virtual subject moves.

2. The information processing apparatus according to claim 1, wherein:

6. The generating means generates and displays an image including an image of the virtual subject and an image showing movement information associated with the virtual subject.

2. The information processing apparatus according to claim 1, wherein:

7. The generating means generates a screen for receiving a selection regarding the virtual subject from a user.

2. The information processing apparatus according to claim 1, wherein:

8. The generating means displays an image of the virtual subject based on the selection received from the user.

8. The information processing apparatus according to claim 7,

9. A determination means for determining photographing parameters relating to photographing the subject.

2. The information processing apparatus according to claim 1, further comprising:

10. The determining means determines the shooting parameters based on the movement information of the virtual subject.

10. The information processing apparatus according to claim 9,

11. The determining means determines the photographing parameters based on information input by a user.

10. The information processing apparatus according to claim 9,

12. The determining means obtains the placement position of the virtual subject from a user, and determines the shooting parameters based on information about the placement position.

10. The information processing apparatus according to claim 9,

13. the generating means generates a virtual image by superimposing an image of the virtual subject on the captured image, The determining means acquires information about the placement position from the user while the virtual image is being displayed.

13. The information processing apparatus according to claim 12.

14. The determining means determines a plurality of placement positions and determines imaging parameters associated with each of the plurality of placement positions.

13. The information processing apparatus according to claim 12.

15. The estimation means estimates movement information of the virtual subject using a learning model.

2. The information processing apparatus according to claim 1, wherein:

16. the estimation means estimates feature information of the virtual subject using a learning model; The generating means generates the virtual subject based on the estimated feature information of the virtual subject.

2. The information processing apparatus according to claim 1, wherein:

17. the estimation means acquires real-time captured images and estimates movement information of the virtual subject; The generating means displays an image in which an image of the virtual subject is superimposed on the captured image.

2. The information processing apparatus according to claim 1, wherein:

18. The information processing device according to claim 1 ; an imaging means for capturing an image of the subject and generating an image; An imaging device comprising:

19. an estimation step of estimating motion information indicating a motion of a virtual subject obtained by virtualizing a subject based on a captured image; a generation step of generating the virtual subject based on feature information indicating features of the virtual subject, and associating an image of the virtual subject with movement information of the virtual subject to generate a virtual image; An information processing method comprising:

20. A program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Imaging apparatus and virtual subject distance decision method

    JP2013251801A