Image processing apparatus, method for controlling image processing apparatus, and program

The image processing device generates virtual subjects to replicate natural movements of hidden subjects, addressing the unnatural reproduction issue in existing technologies by superimposing them on obstacles, ensuring continuous tracking.

JP2025186665APending Publication Date: 2025-12-24CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024094888
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-12
Publication Date
2025-12-24

AI Technical Summary

Technical Problem

Existing image processing technologies fail to reproduce fine movement patterns of subjects hidden by obstacles, resulting in unnatural images.

Method used

An image processing device that acquires video data, detects subjects, generates virtual subjects conforming to their movement patterns when obscured, and superimposes these on obstacles in the video data.

Benefits of technology

Enables the display of images with natural subject movements even when the subject is hidden, allowing continuous tracking during shooting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025186665000001_ABST
    Figure 2025186665000001_ABST
Patent Text Reader

Abstract

To enable display of an image including a subject reproduced with natural movement when a subject to be captured becomes hidden.SOLUTION: When performing live-view display of image data generated by an imaging unit, in a case where a subject selected by a user becomes hidden by an obstacle in the image data, movement of the subject hidden by the obstacle is predicted on the basis of a movement pattern of the subject in a frame displayed in the live-view display, a virtual subject image following the movement pattern is generated, and the generated virtual subject image is superimposed and displayed at a position of the obstacle in the image data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing apparatus, a control method for an image processing apparatus, and a program. [Background technology]

[0002] Conventionally, a technique has been proposed for generating a pseudo image by estimating changes in a subject when it becomes impossible to capture the subject. Patent Document 1 discloses a technique for calculating the relative speed of a moving subject during a period when an imaging device is unable to capture the subject, estimating changes in the subject, and generating and displaying a pseudo image. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-108182 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the technology described in Patent Document 1 cannot reproduce even the fine movement patterns of the subject, resulting in an unnatural image.

[0005] In view of the above-mentioned problems, the present invention has an object to make it possible to display an image including a subject whose natural movements are reproduced when the subject to be photographed is hidden. [Means for solving the problem]

[0006] The image processing device of the present invention is characterized by having an acquisition means for acquiring video data, a subject detection means for detecting a subject from the video data acquired by the acquisition means, a generation means for, when the subject is hidden by an obstacle in the video data, generating an image of a virtual subject that conforms to the movement pattern of the subject hidden by the obstacle based on the movement pattern of the subject, and a display control means for displaying the image of the virtual subject generated by the generation means on a display unit, superimposed on the position of the obstacle in the video data. [Effects of the Invention]

[0007] According to the present invention, when a subject to be photographed is hidden, it is possible to display an image including the subject with its natural movements reproduced. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a block diagram illustrating an example of the internal configuration of an imaging device according to a first embodiment. [Figure 2] 10 is a flowchart showing an example of a processing procedure for superimposing and displaying a virtual subject on a live view display in the first embodiment. [Figure 3] 5A to 5C are diagrams for explaining the difference in live view images when a virtual subject is superimposed in the first embodiment. [Figure 4] FIG. 10 is a block diagram illustrating an example of the internal configuration of an imaging device according to a second embodiment. [Figure 5] 10 is a flowchart showing an example of a processing procedure for superimposing and displaying a virtual subject on a live view display in the second embodiment. [Figure 6] 10A and 10B are diagrams for explaining the difference in live view images when a virtual subject is superimposed in the second embodiment. [Figure 7] 10A and 10B are diagrams for explaining the difference in live view images when a virtual subject is superimposed in the second embodiment. [Figure 8] FIG. 10 is a block diagram illustrating an example of the internal configuration of an imaging device according to a third embodiment. [Figure 9] 11 is a flowchart showing an example of a processing procedure for superimposing and displaying a virtual subject on a live view display in the third embodiment. [Figure 10] 13A to 13C are diagrams for explaining the difference in live view images when a virtual subject is superimposed in the third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the present invention is not limited to the embodiments described below, and various forms within the scope of the gist of the present invention are also included in the present invention. Furthermore, each embodiment described below merely represents one embodiment of the present invention, and each embodiment can be combined as appropriate.

[0010] (First embodiment) 1 is a block diagram showing an example of the internal configuration of an image capture device 100 according to this embodiment. The detailed configuration of the image capture device 100 will be described below as an example of an image processing device. 1, a lens 1001 is a lens group including a zoom lens and a focus lens. A lens control unit 1002 controls the focal length and aperture state of the lens 1001.

[0011] The control unit 1003 includes a portion of a nonvolatile memory (not shown) that stores a program, and a CPU (not shown) functions as the control unit 1003 to control the entire imaging device 100. Note that a GPU (Graphics Processing Unit) may be used instead of the CPU. The control unit bus 1004 is a bus for communication between the control unit 1003 and each functional block. The RAM control unit 1007 controls access to the RAM 1006 based on a RAM access request from each functional block. The RAM bus 1005 is a bus for communication between the RAM control unit 1007 and each functional block. The RAM bus 1005 also has a function of arbitrating access to the RAM 1006 from each functional block.

[0012] The imaging unit 1008 converts optical signals captured by the lens 1001 into electrical signals using an imaging sensor (not shown), and performs processing to correct lens aberrations on the obtained image data and processing to interpolate defective pixels of the imaging sensor. The development unit 1009 performs de-Bayer processing on the image data generated by the imaging unit 1008 to convert it into signals consisting of luminance signals and color difference signals, and performs development processing such as removing noise contained in each signal, correcting optical distortion, and optimizing the image.

[0013] The display control unit 1012 outputs the image data developed by the development unit 1009 to the display unit 1015 as a live view image. The display unit 1015 is configured, for example, with a liquid crystal display. The display unit 1015 also displays information related to the shooting mode, a live view image, a confirmation image after shooting, a display related to subject selection, and a virtual subject, which will be described later. In this embodiment, the image data (frames) displayed as a live view are temporarily stored in the RAM 1006 in order to generate a virtual subject, as will be described later. The number of frames stored in the RAM 1006 is the number of frames required to generate a virtual subject.

[0014] The moving image encoding unit 1013 compresses and encodes the image data developed by the developing unit 1009 using a predetermined moving image compression encoding method such as MPEG4 Video, and converts it into a moving image file with a compressed amount of information. The recording control unit 1014 records the image data developed by the developing unit 1009 in the recording unit 1016. The recording unit 1016 is a recording medium such as a non-volatile memory card or a hard disk. Furthermore, the recording control unit 1014 reads image data from the recording unit 1016 as necessary, and the moving image encoding unit 1013 decodes the read image data. The display control unit 1012 then displays the decoded image data on the display unit 1015 as a moving image.

[0015] The microphone 1017 receives audio input and converts it into an audio signal. The microphone control unit 1018 is connected to the microphone 1017 and controls the microphone 1017, starts and stops sound collection, and acquires collected audio data. Control of the microphone 1017 includes, for example, gain adjustment and status acquisition. The audio encoding / decoding unit 1019 encodes or decodes the audio signal input from the microphone 1017 using a predetermined encoding method such as MPEG4 Audio AAC. The speaker 1020 outputs the audio signal decoded by the audio encoding / decoding unit 1019 as audio.

[0016] The communication unit 1022 is a communication interface that connects the imaging device 100 to other devices via a wired or wireless connection and transmits and receives image data, audio data, and the like, and can also be connected to networks such as a wireless LAN or the Internet. The communication unit 1022 can transmit image data acquired by the imaging device 100 and image data recorded in the recording unit 1016 to the outside, and can receive image data and various information from external devices. The operation unit 1023 accepts various operations from the user to configure various settings of the imaging device 100. Note that if the display unit 1015 is equipped with a touch panel, it is configured as part of the operation unit 1023.

[0017] When the subject to be photographed is included in the image displayed on the display unit 1015, the subject selection unit 1050 selects the subject from the image. When the display unit 1015 is a liquid crystal display equipped with a touch panel, the user can select the subject by directly touching the liquid crystal display. Alternatively, the subject may be selected via another operation unit 1023.

[0018] The subject detection unit 1051 detects the type of subject. Specifically, it detects the subject based on dictionary data for each subject generated by machine learning. Each dictionary data is, for example, data in which the characteristics and type of the corresponding subject are registered, and in subject detection, the subject is detected by sequentially switching between the dictionary data for each subject. The dictionary data storage unit 1055 is configured as part of a non-volatile memory and stores the dictionary data for each subject.

[0019] The control unit 1003 determines which dictionary data from among a plurality of dictionary data to use for subject detection using the subject detection unit 1051 based on preset subject priorities and settings of the imaging device 100. Examples of dictionary data for subject detection include dictionary data for detecting people as subjects, dictionary data for detecting animals, and dictionary data for detecting vehicles. Furthermore, dictionary data for detecting the entire person and dictionary data for detecting the person's face may be stored separately in the dictionary data storage unit 1055. Furthermore, the dictionary data for subject detection may be stored in the recording unit 1016.

[0020] Specific examples of machine learning algorithms include nearest neighbor algorithms, naive Bayes algorithms, decision trees, and support vector machines. Deep learning, which uses a neural network to generate its own features and connection weighting coefficients for learning, is also an example. Any of the above algorithms can be used as appropriate and applied to this embodiment. The subject detection system may also include an error detection unit and an update unit. The error detection unit obtains the error between the training data and output data output from the output layer of the neural network based on input data input to the input layer. The error detection unit may use a loss function to calculate the error between the output data from the neural network and the training data. The update unit updates the connection weighting coefficients between the nodes of the neural network based on the error obtained by the error detection unit to reduce the error. The update unit updates the connection weighting coefficients using, for example, backpropagation. The backpropagation is a technique for adjusting the connection weighting coefficients between the nodes of each neural network to reduce the error.

[0021] The subject detection unit 1051 also detects obstacles, detecting any object that partially obscures the subject as an obstacle. Examples of obstacles include benches, pillars, utility poles, houses, and other large and small buildings, as well as natural objects such as rocks and grass, and people crossing in front of the camera or in a crowd. Furthermore, the subject detection unit 1051 recognizes the type and position of the subject selected by the subject selection unit 1050 using one or more frames temporarily stored in RAM 1006, and generates subject information including the subject's movement pattern. Here, the subject's movement pattern includes information on speed, acceleration, angular velocity, and angular acceleration.

[0022] The virtual subject generation unit 1052 predicts the movement of a subject hidden by an obstacle based on subject information related to the movement pattern of the subject selected by the subject selection unit 1050, and generates an image of the virtual subject. In this embodiment, the image of the virtual subject is generated using a machine learning model that inputs the subject and the subject's movement pattern. The image of the virtual subject generated here may be a two-dimensional image or a three-dimensional image containing information that allows for three-dimensional display. The movement pattern of the subject also includes partial movement of the subject. If the subject is a person, examples of the image include a person walking, a person pedaling a bicycle, a person dancing, etc. If the subject is an animal, examples of the image include an animal walking, an animal running, or a bird flying. If the subject is a vehicle, examples of the image include a car's tires rotating or a locomotive running while emitting smoke. In this way, the virtual subject generation unit 1052 generates an image of the virtual subject based on the subject information of the subject detected by the subject detection unit 1051.

[0023] Furthermore, the virtual subject generation unit 1052 may recognize the subject, input one or more past time-series images (frames) of the subject, predict the subject's movement, and generate an image of the virtual subject using a machine learning model. In this case, the image of the virtual subject may be generated using multiple time-series images and taking into account subject information such as speed, acceleration, angular velocity, and angular acceleration as the subject's movement pattern. Furthermore, by recognizing the position of the subject within the screen, an image of the virtual subject may be generated that follows the subject's movement pattern within the screen. Furthermore, the image of the virtual subject may be generated by recognizing the subject's movement itself. For example, by recognizing whether an animal is walking or running, the time from when the subject hides behind an obstacle until it emerges can be estimated and an image of the virtual subject can be generated.

[0024] Alternatively, the subject detection unit 1051 may generate text data indicating the characteristics of the subject, and the virtual subject generation unit 1052 may input the text data into a machine learning model to generate an image of the virtual subject. In this case, similarly, text data indicating the characteristics of the subject's movement pattern may be generated and input into the machine learning model to generate an image of the virtual subject. The display control unit 1012 displays an image in which the generated image of the virtual subject is superimposed on an obstacle on the display unit 1015. Here, the subject selection unit 1050, the subject detection unit 1051, and the virtual subject generation unit 1052 can be realized by dedicated circuits. Alternatively, the subject selection unit 1050, the subject detection unit 1051, and the virtual subject generation unit 1052 may be realized by a CPU (not shown) executing a program.

[0025] 2 is a flowchart showing an example of a processing procedure for superimposing a virtual subject on a live view display in this embodiment. The processing shown in FIG. 2 is realized by the control unit 1003 loading a program recorded in the nonvolatile memory or the recording unit 1016 into the RAM 1006 and executing it. The processing starts when the user operates the operation unit 1023 to instruct the start of live view display.

[0026] In step S201, the control unit 1003 starts live view display through a series of image capturing processes. Specifically, under the control of the control unit 1003, the imaging unit 1008 captures an image of a subject to generate an image signal, and the development unit 1009 performs a predetermined development process on the generated image signal. The display control unit 1012 then outputs the developed image signal to the display unit 1015 as video data.

[0027] Next, in step S202, the control unit 1003 controls the subject selection unit 1050 to select a subject through a user operation. In this process, if the display unit 1015 is equipped with a touch panel, the subject selection unit 1050 selects a subject when the user touches the screen of the display unit 1015 in the live view video displayed on the display unit 1015. More specifically, the subject detection unit 1051 detects a subject within a predetermined range from the position touched by the user, and the subject selection unit 1050 recognizes the detected subject as the subject selected by the user. Note that if the display unit 1015 is not equipped with a touch panel, the subject detection unit 1051 may be configured to detect subjects in the entire live view video and allow the user to select from multiple detected subjects. Then, once a subject is selected, the control unit 1003 instructs the subject detection unit 1051 to track the selected subject.

[0028] Next, in step S203, the control unit 1003 determines whether the subject selected in step S202 is hidden by an obstacle in the frame to be processed in the live view video displayed on the display unit 1015. In this process, the control unit 1003 determines whether the subject selected in step S202 is hidden by an obstacle through tracking by the subject detection unit 1051 in the live view video. Here, the criterion for determining that the subject is hidden by an obstacle does not necessarily require that the subject is completely hidden. For example, the subject may be determined to be hidden by an obstacle when only a portion of the subject is hidden and its movement cannot be confirmed, or the subject may be determined to be hidden by an obstacle when only a small portion of the subject is hidden. In this embodiment, the subject is determined to be hidden by an obstacle when only a small portion of the subject is hidden. Furthermore, the determination criterion may be changed depending on the type of subject. If the control unit 1003 determines that the subject is hidden by an obstacle, the process proceeds to step S204; otherwise, the process proceeds to step S206.

[0029] In step S204, the control unit 1003 instructs the virtual subject generation unit 1052 to generate an image of a virtual subject of a subject hidden by an obstacle. In this process, the virtual subject generation unit 1052 generates an image of the virtual subject using a machine learning model from the subject information generated by tracking by the subject detection unit 1051 and / or the time-series frames stored in the RAM 1006 after the live view display has ended.

[0030] Next, in step S205, the control unit 1003 instructs the display control unit 1012 to superimpose the image of the generated virtual subject on the obstacle in the live view image and output the superimposed live view image to the display unit 1015.

[0031] FIG. 3(a) is a diagram showing an example of a live view image when a virtual subject is superimposed by the processing of this embodiment, and FIG. 3(b) is a diagram showing an example of an image actually obtained by the imaging unit 1008. As shown in FIG. 3(b), an animal 301, which is actually the subject, runs from the right side of the screen and is hidden by an obstacle 302. In this case, the user may lose sight of the subject and be unable to track the subject, which may interfere with shooting. Therefore, in this embodiment, as shown in FIG. 3(a), a virtual subject is generated and displayed superimposed on the obstacle 302. This allows the user to track the subject during shooting without losing sight of the subject.

[0032] Next, in step S206, the control unit 1003 determines whether to end the live view display. In this process, if the frame following the frame of the live view video for which it was determined in step S203 whether the subject was hidden by an obstacle is input, the control unit 1003 determines to continue the live view display. On the other hand, if the frame following the frame of the live view video for which it was determined in step S203 whether the subject was hidden by an obstacle is not input and an instruction to end the live view display is received via the operation unit 1023, the control unit 1003 determines to end the live view display. If the control unit 1003 determines to end the live view display, it ends the process; otherwise, the process returns to step S203.

[0033] As described above, according to this embodiment, when a subject is hidden by an obstacle in a live view image, a virtual subject is displayed superimposed on the obstacle, with movement conforming to the movement pattern of the subject. This makes it possible to realize an image similar to a natural live view display, and to track the subject during shooting without losing sight of it.

[0034] Furthermore, although the present embodiment has been described with respect to a case where the subject is moving, the present invention can also be applied to a case where the user moves the imaging device 100, changing the imaging range and, as a result, obscuring the subject behind an obstacle. For example, an animal may be stationary, but the user may move while holding the imaging device 100 and the animal may become obscured by an obstacle. In this case, even if the animal is obscured by an obstacle, the movement of the subject can be confirmed by superimposing a virtual subject, and the user can also consider factors such as the angle of view, composition, and angle. Furthermore, if half of the subject's body is obscured, the user can move to a location where the subject's entire body can be captured.

[0035] Furthermore, in this embodiment, an example has been described in which a virtual subject is generated for a subject selected by the user. However, if there is only one moving subject in the live view video, the imaging device may select that subject and generate a virtual subject. For example, when capturing a natural landscape and displaying the live view, if a bird is detected, the subject detection unit 1051 detects the bird and the subject selection unit 1050 selects the bird. Then, if the bird is hidden by a tree, for example, a virtual subject may be generated and superimposed on the live view video. Furthermore, although this embodiment focuses on live view video, the processing of this embodiment can also be applied to video data recorded in the recording unit 1016.

[0036] (Second embodiment) In this embodiment, an example will be described in which, when a subject is hidden by an obstacle, the obstacle is displayed semi-transparently. Fig. 4 is a block diagram showing an example of the internal configuration of the imaging device 100 in this embodiment. Hereinafter, the same components as in the first embodiment are assigned the same reference numerals as in Fig. 1, and their description will be omitted. Only the differences from the first embodiment will be described below.

[0037] The background complementing unit 1053 virtually complements the part of the background hidden by the obstacle in order to display the obstacle as if it were translucent. Specifically, the background complementing unit 1053 uses the background surrounding the obstacle to generate the background hidden by the obstacle as a virtual background on the obstacle. At this time, an image is generated in which the obstacle appears translucent. The background complementing unit 1053 may also be realized by a dedicated circuit, or may be realized by a CPU (not shown) executing a program.

[0038] Fig. 5 is a flowchart showing an example of a processing procedure for superimposing a virtual subject on a live view display in this embodiment. The processing shown in Fig. 5 is realized by the control unit 1003 loading a program recorded in the non-volatile memory or recording unit 1016 into the RAM 1006 and executing it. The processing starts when the user operates the operation unit 1023 to instruct the start of live view display. Note that steps S201 to S206 are the same as those in Fig. 2, and therefore a description thereof will be omitted.

[0039] In step S501, the control unit 1003 instructs the background complementing unit 1053 to complement the background in which the subject is hidden by an obstacle. In this process, the background complementing unit 1053 predicts the background hidden by the obstacle from the background around the obstacle using a machine learning model. The background complementing unit 1053 then generates an image using the predicted background as a virtual background, and the display control unit 1012 superimposes the virtual background on the obstacle in the live view video and outputs the image to the display unit 1015. By superimposing the virtual background in this way and making the obstacle transparent, the presence of the obstacle can be recognized.

[0040] FIG. 6(a) is a diagram showing an example of a live view image when a virtual subject is superimposed by the processing of this embodiment, and FIG. 6(b) is a diagram showing an example of an image actually obtained by the imaging unit 1008. As shown in FIG. 6(b), it is assumed that an animal 601, which is actually the subject, runs from the right side of the screen and is hidden by an obstacle 602. In this embodiment, as shown in FIG. 6(a), not only is a virtual subject generated and displayed superimposed on the obstacle 602, but a virtual background is also generated on the obstacle 602, and the obstacle 602 is displayed as if it were transparent. This allows the user to track the subject during shooting without losing sight of the subject and without it appearing unnatural against the background.

[0041] In this embodiment, when an image of a virtual background is generated and superimposed, obstacles are made semi-transparent, but in reality, not only the background but also the subject is hidden by the obstacle, so the virtual subject may also be made to appear in the same manner as the virtual background. In this case, in step S204 of Fig. 5, when generating an image of a virtual subject, the virtual subject generation unit 1052 generates an image in which the virtual subject can be seen through the obstacle.

[0042] FIG. 7(a) is a diagram showing an example of a live view image in which a virtual subject and virtual background are superimposed by the processing of this embodiment so that the virtual subject and virtual background can be seen through the obstacle, and FIG. 7(b) is a diagram showing an example of an image actually obtained by the imaging unit 1008. As shown in FIG. 7(b), an animal 701, which is actually the subject, runs from the right side of the screen and is hidden by an obstacle 702. Therefore, when a virtual subject is generated and superimposed on the obstacle 702 as shown in FIG. 7(a), the virtual subject is displayed so that the virtual subject can be seen through the obstacle 702. This allows the user to track the subject more naturally during shooting without being aware of the presence of the obstacle.

[0043] As described above, according to this embodiment, obstacles are displayed as if they are transparent, so that the subject can be tracked during shooting without losing sight of the subject and without any sense of discomfort.

[0044] (Third embodiment) In this embodiment, in addition to the contents of the second embodiment, an example will be described in which the color of the background other than the subject specified by the user is changed when displaying a live view. Fig. 8 is a block diagram showing an example of the internal configuration of the imaging device 100 in this embodiment. Hereinafter, the same components as those in the second embodiment are assigned the same reference numerals as in Fig. 4, and their description will be omitted. Only the differences from the second embodiment will be described below.

[0045] The color change unit 1054 changes the color of the area other than the subject selected by the subject selection unit 1050. When changing the color of the area other than the subject, at least one of the hue, brightness, and saturation is changed. For example, the subject can be displayed in its original color while the color of the background other than the subject is changed to a lighter color, thereby making the subject stand out more. Another method is to display the subject in color while changing the color of the background other than the subject to monochrome when a live view display is being performed using color video. The color change unit 1054 may also be implemented by a dedicated circuit, or may be implemented by a CPU (not shown) executing a program.

[0046] Fig. 9 is a flowchart showing an example of a processing procedure for superimposing a virtual subject on a live view display in this embodiment. The processing shown in Fig. 9 is realized by the control unit 1003 loading a program recorded in the non-volatile memory or the recording unit 1016 into the RAM 1006 and executing it. The processing starts when the user operates the operation unit 1023 to instruct the start of live view display. Note that steps S201 to S206 and step S501 are the same as those in Fig. 5, and therefore description thereof will be omitted.

[0047] In step S901, the control unit 1003 instructs the color modification unit 1054 to modify the color of the background other than the subject selected in step S202. Specifically, the color modification unit 1054 modifies the color of the background other than the selected subject for the image signal developed by the development unit 1009. Thereafter, the display control unit 1012 outputs the image signal in which the color of the background other than the subject has been modified to the display unit 1015 as video data.

[0048] FIG. 10(a) is a diagram showing an example of a live view image in which the color of the background other than the subject is changed and a virtual subject is superimposed by the processing of this embodiment, and FIG. 10(b) is a diagram showing an example of an image actually obtained by the imaging unit 1008. As shown in FIG. 10(b), it is assumed that an animal 1060, which is actually the subject, runs from the right side of the screen and is hidden by an obstacle 1070. Therefore, as shown in FIG. 10(a), the color of the background other than the animal 1060 (including the obstacle 1070) is changed, and further, a virtual subject and virtual background are superimposed as in FIG. 7(a). By changing the color of the background other than the subject in this way, the subject can be made to stand out more, and it becomes possible to track the subject during shooting more naturally without losing sight of the subject.

[0049] As described above, according to this embodiment, the color of the background other than the subject in the live view video is changed, making the subject more noticeable. Furthermore, since the virtual subject is displayed in a superimposed manner, it is possible to track the subject during shooting without losing sight of it. Note that, although obstacles are displayed as if they are transparent, it is not necessary to complement the background hidden by the obstacle. It is also possible to change the color of the background other than the subject and display only the virtual subject in a superimposed manner, as in the first embodiment.

[0050] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0051] The disclosure of this embodiment includes the following configuration, method, and program.

[0052] (Configuration 1) an acquisition means for acquiring video data; a subject detection means for detecting a subject from the video data acquired by the acquisition means; a generating means for generating, when the subject is hidden by an obstacle in the video data, an image of a virtual subject that conforms to the movement pattern of the subject based on the movement pattern of the subject; a display control means for displaying the image of the virtual subject generated by the generation means on a display unit in such a manner that the image is superimposed on the position of the obstacle in the video data; 1. An image processing device comprising:

[0053] (Configuration 2) the subject detection means detects a pattern of movement of the subject from one or more frames of the video data and generates subject information including the pattern of movement of the subject; 2. The image processing device according to configuration 1, wherein the generating means generates an image of the virtual subject based on the subject information. (Configuration 3) 3. The image processing device according to claim 1, wherein the generating means generates an image of the virtual subject from one or more frames of the video data using a machine learning model. (Configuration 4) 4. The image processing device according to any one of configurations 1 to 3, wherein the subject detection means detects the subject based on a position designated by a user operation.

[0054] (Configuration 5) a complementing means for complementing a background hidden by the obstacle by making the obstacle in the video data semi-transparent; The image processing device according to any one of configurations 1 to 4, wherein the display control means causes the display unit to display an image of the virtual subject superimposed on the position of the translucent obstacle whose background has been complemented by the complement means. (Configuration 6) 6. The image processing device according to configuration 5, wherein the generating means generates an image in which the virtual subject can be seen through the obstacle. (Configuration 7) The image processing device according to any one of configurations 1 to 6, further comprising a change means for changing at least one of hue, brightness, and saturation in an area other than the object detected by the object detection means in the video data.

[0055] (method) an acquisition step of acquiring video data; a subject detection step of detecting a subject from the video data acquired in the acquisition step; a generation step of generating, when the subject is hidden by an obstacle in the video data, an image of a virtual subject that follows the movement pattern of the subject based on the movement pattern of the subject; a display control step of superimposing the image of the virtual subject generated in the generation step on a position of the obstacle in the video data and displaying the image on a display unit; 1. A method for controlling an image processing apparatus, comprising:

[0056] (program) A program for causing a computer to function as each means of the image processing device according to any one of configurations 1 to 7. [Explanation of symbols]

[0057] 1003 control unit, 1012 display control unit, 1050 object selection unit, 1051 object detection unit, 1052 virtual object generation unit

Claims

1. an acquisition means for acquiring video data; a subject detection means for detecting a subject from the video data acquired by the acquisition means; a generating means for generating, when the subject is hidden by an obstacle in the video data, an image of a virtual subject that conforms to the movement pattern of the subject based on the movement pattern of the subject; a display control means for displaying the image of the virtual subject generated by the generation means on a display unit, the image being superimposed on the position of the obstacle in the video data; 1. An image processing device comprising:

2. the subject detection means detects a pattern of movement of the subject from one or more frames of the video data and generates subject information including the pattern of movement of the subject; The image processing device according to claim 1 , wherein the generating means generates an image of the virtual subject based on the subject information.

3. The image processing device according to claim 1 , wherein the generating means generates an image of the virtual subject from one or more frames of the video data using a machine learning model.

4. 2. The image processing apparatus according to claim 1, wherein the subject detection means detects the subject based on a position designated by a user's operation.

5. a complementing means for complementing a background hidden by the obstacle by making the obstacle in the video data semi-transparent, 2. The image processing device according to claim 1, wherein the display control means causes the display unit to display the image of the virtual subject superimposed on the position of the semi-transparent obstacle whose background has been complemented by the complementing means.

6. 6. The image processing device according to claim 5, wherein the generating means generates an image in which the virtual subject can be seen through the obstacle.

7. 2. The image processing device according to claim 1, further comprising a change unit that changes at least one of hue, brightness, and saturation in an area other than the object detected by the object detection unit in the video data.

8. an acquisition step of acquiring video data; a subject detection step of detecting a subject from the video data acquired in the acquisition step; a generation step of generating, when the subject is hidden by an obstacle in the video data, an image of a virtual subject that follows the movement pattern of the subject based on the movement pattern of the subject; a display control step of superimposing the image of the virtual subject generated in the generation step on a position of the obstacle in the video data and displaying the image on a display unit; 1. A method for controlling an image processing apparatus, comprising:

9. A program for causing a computer to function as each of the means of the image processing apparatus according to claim 1.

Citation Information

Patent Citations

  • Vehicle driving support apparatus

    JP2010108182A