Imaging device, control method, program, and imaging system
The system addresses the lack of posture guidance in photography by estimating skeletal landmarks and generating pose instruction images, ensuring accurate and effective shooting.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-11-25
- Publication Date
- 2026-06-04
AI Technical Summary
Existing photography techniques fail to provide clear instructions on how to adjust posture when the preset posture of the subject does not match the actual posture, leading to inappropriate shooting.
A system comprising a photographing device and a display device connected via a network, which estimates skeletal landmarks, acquires pose change instructions, generates pose instruction images, and transmits them to the display device for the subject to follow, ensuring appropriate shooting.
Enables photographers to take appropriate pictures by providing clear guidance on posture adjustments to the subject, enhancing the quality of captured images.
Smart Images

Figure 2026091528000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for assisting photography.
Background Art
[0002] There is a technique that can make it easier for a photographer to take an image with a desired composition by estimating the posture or skeleton of a person or the like to be photographed (hereinafter referred to as "subject") before shooting. Patent Document 1 discloses a technique of audibly emitting a message prompting a person of the subject to change their posture when the preset posture of the person does not match the posture of the person of the subject.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the technique disclosed in Patent Document 1, when the preset posture of the person does not match the posture of the person of the subject, a message prompting the person of the subject to change their posture is issued, but no indication of how to change the posture is given. Therefore, the technique disclosed in Patent Document 1 has a problem that the person of the subject cannot recognize how to change their own posture, and thus there is a problem that appropriate shooting cannot be performed.
Means for Solving the Problems
[0005] The photographing device according to this disclosure includes: a skeleton estimation means for estimating skeletal landmarks in the image of a person to be photographed included in the captured image; an instruction acquisition means for acquiring a pose change instruction based on the user's specification of the position of the skeletal landmark after movement; an image generation means for generating a pose instruction image corresponding to the change instruction; and a transmission means for transmitting the signal of the pose instruction image to a display device positioned at a location where the person to be photographed can be seen. [Effects of the Invention]
[0006] According to this disclosure, it is possible to provide support to photographers in order to take pictures appropriately. [Brief explanation of the drawing]
[0007] [Figure 1] This figure shows an example of the configuration of the imaging system according to Embodiment 1. [Figure 2] This is a block diagram showing an example of the hardware configuration of the imaging device and display device according to Embodiment 1. [Figure 3] This is a block diagram showing an example of the logical configuration of the imaging device and display device according to Embodiment 1. [Figure 4] This flowchart shows an example of the processing flow of the imaging device according to Embodiment 1. [Figure 5] This figure shows an example of a live view image according to Embodiment 1. [Figure 6] This figure illustrates an example of the configuration and learning method of the skeleton estimation DL model according to Embodiment 1. [Figure 7] This figure shows an example of a skeletal landmark image according to Embodiment 1. [Figure 8] This figure shows an example of a pause instruction input screen according to Embodiment 1. [Figure 9] This figure shows an example of a display image according to Embodiment 1. [Figure 10] This is a block diagram showing an example of the logical configuration of the imaging device and display device according to Embodiment 2. [Figure 11]It is a flowchart showing an example of the processing flow of the imaging device according to Embodiment 2. [Figure 12] It is a diagram showing an example of the pose instruction input screen according to Embodiment 2. [Figure 13] It is a diagram showing an example of the skeletal movable range data according to Embodiment 2. [Figure 14] It is a diagram showing an example of the pose instruction image corresponding to the alternative pose according to Embodiment 2. [Figure 15] It is a diagram showing an example of the display image according to Embodiment 2. [Figure 16] It is a diagram showing an example of the configuration of the imaging system according to Embodiment 3. [Figure 17] It is a diagram showing the hardware configuration of the imaging device and the display device according to Embodiment 3. [Figure 18] It is a block diagram showing an example of the logical configuration of the imaging device and the display device according to Embodiment 3. [Figure 19] It is a diagram showing an example of the live view image according to Embodiment 3. [Figure 20] It is a diagram showing an example of the skeletal landmark image according to Embodiment 3. [Figure 21] It is a diagram showing an example of the pose instruction input screen according to Embodiment 3. [Figure 22] It is a diagram showing an example of the display image according to Embodiment 3.
Embodiments for Carrying Out the Invention
[0008] Hereinafter, embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the configuration for solving the problems according to the present disclosure. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential for solving the problems according to the present disclosure, and a plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are given the same reference numerals, and redundant explanations are omitted.
[0009] [Embodiment] FIG. 1 is a diagram showing an example of the configuration of a photographing system according to Embodiment 1. The photographing system includes a photographing device 101 and a display device 102. The photographing device 101 and the display device 102 are communicably connected to each other via a network 105. Hereinafter, the case where the photographing system has one photographing device 101 will be described. The photographing device 101 is a device having a photographing function and a communication function, which is constituted by a digital still camera, a digital video camera, a smartphone, or the like. The display device 102 is a device for displaying a received image, which is constituted by a television receiver, a display for a PC (personal computer), or the like.
[0010] The photographing device 101 receives an input of a pose change instruction (hereinafter referred to as a "pose instruction") from a photographer 104 who is a user of the photographing device 101, and transmits image data regarding the pose instruction to the display device 102 via the network 105. The display device 102 displays an image regarding the pose instruction received from the photographing device 101. A person 103 as a photographing subject (hereinafter referred to as a "subject") takes a pose according to the image regarding the pose instruction displayed on the display device 102.
[0011] Figure 2 is a block diagram showing an example of the hardware configuration of the imaging device 101 and display device 102 according to Embodiment 1. The imaging device 101 has a hardware configuration of a CPU 200, ROM 201, RAM 202, communication unit 203, storage medium 204, display unit 205, input unit 208, and imaging unit 207. These hardware components of the imaging device 101 are connected to each other so as to be able to communicate via a bus 206. The CPU 200 is a control unit including at least one processor or processing circuit, and controls the entire imaging device 101. The ROM 201 is an electrically erasable and recordable memory, and stores various data and programs used in processing by the CPU 200. The program referred to here is a computer program for executing the processing of various flowcharts described later in this embodiment. The RAM 202 is a memory used as the work area of the CPU 200, and the data used in processing by the CPU 200 and programs read from the ROM 201 are stored in the RAM 202.
[0012] The communication unit 203 is an interface for communicating with external devices such as network equipment or USB devices, and performs data communication via the network 105, or data transmission and reception with external devices. The storage medium 204 is a non-volatile memory composed of semiconductor memory such as a memory card. The input unit 208 is a device that receives input from the photographer 104, composed of buttons or a touch panel. The display unit 205 is composed of a display device such as a liquid crystal monitor and displays a GUI (Graphical User Interface) related to the operating status and setting changes of the imaging device 101. The imaging unit 207 is an image sensor composed of a CCD (Charge-Coupled Device) or CMOS (Complementary Metal-Oxide-Semiconductor) element, which converts optical images into electrical signals. The CPU 200 also operates as a control unit that controls the input unit 208, the display unit 205, and the imaging unit 207.
[0013] The display device 102 has a hardware configuration consisting of a CPU 210, ROM 211, RAM 212, communication unit 213, storage medium 214, and display unit 215. These hardware components of the display device 102 are connected to each other via a bus 216 so as to be able to communicate with one another. The CPU 210 is a control unit including at least one processor or processing circuit, and controls the entire display device 102. The ROM 211 is an electrically erasable and recordable memory, and stores various data and programs used in processing by the CPU 210. Here, "program" refers to a computer program for executing the processing of various flowcharts described later in this embodiment. The RAM 212 is a memory used as the work area of the CPU 210, and the data used in processing by the CPU 210 and programs read from the ROM 211 are stored in the RAM 212.
[0014] The communication unit 213 is an interface for communicating with external devices such as network equipment or USB devices, and performs data communication via the network 105 or transmission and reception of data with external devices. The storage medium 214 is a non-volatile memory composed of semiconductor memory such as a memory card. The display unit 215 is composed of a display device such as a liquid crystal monitor and displays images acquired via the communication unit 213.
[0015] Figure 3 is a block diagram showing an example of the logical configuration of the imaging device 101 and display device 102 according to Embodiment 1. The imaging device 101 has the following logical configuration: an image acquisition unit 300, a skeletal estimation unit 302, an instruction acquisition unit 305, an image generation unit 307, a transmission unit 308, and a display control unit 309. Each logical configuration of the imaging device 101 is realized by the CPU 200 loading a program stored in the ROM 201 into the RAM 202 and executing it.
[0016] The image acquisition unit 300 acquires the image obtained by the imaging unit 207 as a live view image and stores the acquired live view image data in the storage medium 204. The skeleton estimation unit 302 acquires skeleton landmark data by estimating the skeleton landmarks in the image of the person 103, which is included in the live view image acquired by the image acquisition unit 300. Specifically, for example, first the skeleton estimation unit 302 inputs the live view image data into a skeleton estimation DL (Deep Learning) model obtained as a result of learning by machine learning, etc. Next, the skeleton estimation unit 302 acquires the skeleton landmark data output as an estimation result from the skeleton estimation DL model. The configuration and learning method of the skeleton estimation DL model will be described later. The skeleton landmark data acquired by the skeleton estimation unit 302 is converted into a skeleton landmark image by the image generation unit 307, and the skeleton landmark image is displayed on the display unit 205 via the display control unit 309. Hereinafter, the skeleton landmark image will be described as being displayed on the display unit 205 superimposed on the live view image.
[0017] The instruction acquisition unit 305 acquires pose instructions from the photographer 104 received by the input unit 208 and stores the acquired pose instruction data (hereinafter referred to as "pose instruction data") in the storage medium 204. For example, the photographer 104 inputs pose instructions while referring to the skeletal landmark image displayed on the display unit 205. The pose instructions acquired by the instruction acquisition unit 305 are converted into a pose instruction image by the image generation unit 307, and the pose instruction image is displayed on the display unit 205 via the display control unit 309. Specifically, for example, the image generation unit 307 generates a display image by superimposing the pose instruction image, the live view image, and the skeletal landmark image, and the display control unit 309 displays the display image generated by the image generation unit 307 on the display unit 205. The pose instruction image is also transmitted to the display device 102 as a display image via the transmission unit 308 and the network 105. Specifically, for example, the image generation unit 307 generates a display image that superimposes the pose instruction image and the live view image, and the transmission unit 308 transmits the display image generated by the image generation unit 307 to the display device 102 via the network 105.
[0018] The display device 102 has a logical configuration consisting of a receiving unit 310 and a display control unit 311. The logical configurations of the display device 102 are realized by the CPU 210 loading a program stored in the ROM 211 into the RAM 212 and executing it. The receiving unit 310 receives data of the display image transmitted from the imaging device 101. The display control unit 311 displays the display image received by the receiving unit 310 on the display unit 215.
[0019] Figure 4 is a flowchart showing an example of the processing flow of the imaging device 101 according to Embodiment 1. The processing shown in the flowchart in Figure 4 is realized by the CPU 200 of the imaging device 101 loading the program stored in the ROM 201 into the RAM 202 and executing it. In addition, some or all of the processing shown in the flowchart in Figure 4 may be executed by a processing circuit. Furthermore, in the following description, the symbol "S" means a step (process).
[0020] First, in S400, the image acquisition unit 300 acquires a live view image. The data of the live view image acquired by the image acquisition unit 300 is stored in the storage medium 204. The live view image acquired by the image acquisition unit 300 is then displayed on the display unit 205 via the display control unit 309. Figure 5 shows an example of a live view image 500 displayed on the display unit 205 according to Embodiment 1. The live view image 500 includes an image 501 of the subject, person 103. Next, in S401, the skeleton estimation unit 302 acquires skeleton landmark data by estimating the skeleton landmarks in the image 501 of the subject, person 103, included in the live view image 500 acquired in S400. The skeleton landmark data acquired by the skeleton estimation unit 302 is stored in the storage medium 204.
[0021] Here, with reference to Figure 6, the skeleton estimation DL model used by the skeleton estimation unit 302 for estimating skeleton landmarks will be described. Figure 6 is a diagram illustrating an example of the configuration and learning method of the skeleton estimation DL model 602 according to Embodiment 1. The skeleton estimation DL model 602 has an input layer, one or more intermediate layers, and an output layer, and each layer contains one or more nodes. Images corresponding to each frame of the live view image are input to the input layer of the skeleton estimation DL model 602 as training data. Skeletal landmark data is output from the output layer of the skeleton estimation DL model 602 as the estimation result of the skeleton landmarks corresponding to the images input to the input layer.
[0022] In training the skeletal estimation DL model 602, first, the difference between the skeletal landmark data 601, which is the ground truth data corresponding to the image input to the input layer, and the skeletal landmark data output from the output layer is calculated using a loss function or the like. Next, the weight parameters of each node in the hidden layer are updated using backpropagation or the like so that this difference becomes smaller. By repeating this process, for example, until the above difference falls within a predetermined range, the trained skeletal estimation DL model 602 is obtained.
[0023] Following S401, in S402, the image generation unit 307 generates a skeletal landmark image by imaging the skeletal landmark data acquired in S401. The skeletal landmark image generated by the image generation unit 307 is displayed on the display unit 205 via the display control unit 309. The image generation unit 307 may also generate an image by superimposing the generated skeletal landmark image onto the live view image acquired in S400. In this case, the image generated by the image generation unit 307 is displayed on the display unit 205 via the display control unit 309.
[0024] Figure 7 shows an example of a skeletal landmark image displayed on the display unit 205 according to Embodiment 1. Specifically, Figure 7(a) shows an example of a skeletal landmark image 700 generated by the image generation unit 307. Figure 7(b) shows an example of a skeletal landmark image 710 generated by the image generation unit 307 by superimposing the skeletal landmark image 700 and the live view image 500. The black circle 701 is a skeletal landmark corresponding to important parts of the skeleton, such as the acromion and vertex, in the image 501 of the subject person 103.
[0025] The skeletal landmark data stores the coordinates indicating the position of each skeletal landmark and information indicating the connection relationships between skeletal landmarks. In Figure 7, the connection relationships between linked skeletal landmarks are represented by line segments connecting the black circles 701. In other words, the skeletal landmark data stores a list of each skeletal landmark, the 3D coordinates of each skeletal landmark, and information about other skeletal landmarks connected to each skeletal landmark. By processing the skeletal landmark data, it is possible to calculate the connection relationships between skeletal landmarks and the distances between interconnected skeletal landmarks. For example, the distance between the acromion and the elbow can be identified as corresponding to the length of the upper arm, and this distance can be calculated from the 3D coordinates of the skeletal landmarks corresponding to the acromion and elbow, and information indicating the connection relationships between these skeletal landmarks.
[0026] After S402, in S403, the instruction acquisition unit 305 acquires the pose instruction from the photographer 104 received by the input unit 208 and stores the pose instruction data in the storage medium 204. Figure 8 is a diagram showing an example of the pose instruction input screen 800 displayed on the display unit 205 according to Embodiment 1. Referring to Figure 8(a), the method of inputting a pose instruction by the photographer 104 will be described. Hereinafter, the display unit 205 and the input unit 208 are composed of a liquid crystal panel and a touch sensor, or a touch panel, and the photographer 104 will be described as inputting a pose instruction by performing a touch operation on the pose instruction input screen 800 displayed on the display unit 205. First, the photographer 104 selects a skeletal landmark as the skeletal landmark to be moved by touching any skeletal landmark from among the multiple skeletal landmarks displayed in the skeletal landmark image 710 shown as an example in Figure 7(b). Hereinafter, as an example, the case in which the photographer 104 selects the skeletal landmark 801 as the skeletal landmark to be moved (hereinafter simply referred to as the "target landmark") shown in Figure 8 will be described.
[0027] When skeletal landmark 801 is selected as the target landmark, the imaging device 101 displays skeletal landmark 802 for pose instruction and GUI component 803 in the vicinity of skeletal landmark 801. GUI component 803 is a GUI component that allows the photographer 104 to input the direction of movement, represented by the X, Y, and Z axes. For example, if the photographer 104 selects the X axis on GUI component 803 by touching it, the photographer 104 can move skeletal landmark 802 for pose instruction only in the direction of the X axis. Similarly, if the photographer 104 selects the Y or Z axis on GUI component 803, the photographer 104 can move skeletal landmark 802 for pose instruction only in the direction of the selected axis.
[0028] Furthermore, movement in the X-axis and Y-axis directions is represented, for example, by the horizontal and vertical display positions of the skeletal landmark 802 on the screen. Movement in the Z-axis direction is represented, for example, by the display of the 3D coordinates 804 of the skeletal landmark 802. The representation of movement in the Z-axis direction is not limited to this and may also be represented by changing the shape or other aspects of the skeletal landmark 802. For example, movement in the Z-axis direction may be represented by the size of the white circle inside the skeletal landmark 802. Specifically, for example, the white circle may be made larger as the skeletal landmark 802 moves in the positive Z-axis direction, and smaller as the skeletal landmark 802 moves in the negative Z-axis direction.
[0029] Figures 8(b) and 8(c) show other embodiments of the pose instruction input screen 800. The pose instruction input screen 800 shown in Figure 8(b) includes a slider bar 815 for inputting the amount of movement in each axis direction as a GUI component for pose instruction. The photographer 104 specifies the position of the skeletal landmark 802 for pose instruction by changing the grip position of the slider bar 815.
[0030] The pose instruction input screen 800 shown in Figure 8(c) includes a "Move Back" button 826 and a "Move Forward" button 827 as GUI components for pose instruction. For movement in the X-axis and Y-axis directions, the photographer 104 specifies the destination by touching the desired position on the screen where they want to move the skeletal landmark 802. For movement in the Z-axis direction, the photographer 104 instructs by touching the "Move Back" button 826 or the "Move Forward" button 827. The skeletal landmark 802 in the pose instruction input screen 800 shown in Figure 8(c) is an example of when the skeletal landmark 802 is moved to the back side when the "Move Back" button is pressed. In the example shown in Figure 8(c), the movement of the skeletal landmark 802 to the back side is represented by the inner white circle becoming smaller.
[0031] Following S403, in S404, the image generation unit 307 generates a pose instruction image using the live view image, pose instruction data, and skeletal landmark data acquired in S400, S401, or S403. The pose instruction image generated by the image generation unit 307 is displayed on the display unit 205 via the display control unit 309 as a display image. The image generation unit 307 may also generate a display image by superimposing the generated pose instruction image onto the live view image acquired in S400. Alternatively, the image generation unit 307 may also generate a display image by superimposing the generated pose instruction image onto the live view image acquired in S400 and the skeletal landmark image generated in S402. The display image generated by the image generation unit 307 is displayed on the display unit 205 via the display control unit 309 for the photographer 104.
[0032] Next, in S405, the transmission unit 308 transmits the display image data generated in S404 to the display device 102 via the network 105. The display device 102 receives the display image data transmitted in S405 and displays the display image on the display unit 215 for the subject, person 103. The display image displayed on the display unit 205 and the display image transmitted to the display device 102 via the network 105 may be the same or different. That is, the display image for the photographer 104 and the display image for the subject, person 103, may be the same or different. After S405, the shooting device 101 completes the processing shown in the flowchart in Figure 4.
[0033] Figure 9 shows an example of a display image shown in the display unit 205 or display unit 215 according to Embodiment 1. The display image 900 shown in Figure 9(a) shows the position of the target landmark 901 and the position of the skeletal landmark 902 after the target landmark 901 has moved, corresponding to the pose instruction. The target landmark 901 and the moved skeletal landmark 902 are represented in a manner that allows the photographer 104 and the subject person 103 to determine the direction of movement of the target landmark 901 in the front-to-back direction. Specifically, the target landmark 901 is represented by a white circle containing a black circle, and the moved skeletal landmark 902 is represented by a black circle containing a white circle. When a skeletal landmark is represented by a white circle containing a black circle, it indicates that the position of the skeletal landmark is in front of the reference position. Conversely, when a white circle contains a black circle, it indicates that the position of the skeletal landmark is behind the reference position.
[0034] The amount of displacement of the skeletal landmark from a reference position in the front-to-back direction can be represented, for example, by the size of the circle on the striking side contained within the outer circle. For example, the inner black circle increases as the position of the skeletal landmark shifts toward the front relative to the reference position, and the inner white circle increases as the position of the skeletal landmark shifts toward the back relative to the reference position. The subject, person 103, can recognize which parts of their body should be moved, in what direction, and by how much by checking the display image shown on the display unit 215. Note that the above method of representing the amount of displacement in the front-to-back direction is merely an example and is not limited to this.
[0035] The display image 900 shown in Figure 9(a) includes information 903 indicating the number of mismatches. The number of mismatches is the number of target landmarks 901 of the person 103, which is the subject, whose position does not match the position of the skeletal landmark 902 used for pose instructions. For example, the camera 101 determines whether the distance between the skeletal landmark 902 used for pose instructions and the target landmark 901 of the person 103, which is the subject, is within a given threshold, and counts those for which the distance is not within the threshold. Here, the threshold may be a default value given in advance by the system, or it may be given by input from the photographer 104. Note that the number of mismatches is not limited to the number of target landmarks 901 whose position does not match the position of the skeletal landmark used for pose instructions. For example, if the person 103, which is the subject, changes their posture according to the pose instruction image, the camera 101 may also count the number of skeletal landmarks whose position has shifted from their original position as part of the number of mismatches.
[0036] The number of mismatches becomes 0 when the subject person 103 assumes the instructed pose. If the position of the skeletal landmark 902 for pose instruction and the position of the target landmark 901 of the subject person 103 coincide, the imaging device 101 may generate, display, and transmit a display image as shown below. Specifically, for example, in this case, the imaging device 101 expresses that they coincide by changing at least one of the shape, color, and transparency of the skeletal landmark 902 or target landmark 901 in the display image.
[0037] The display image 910 shown in Figure 9(b) is a pose instruction image superimposed on the live view image, which includes only the target landmark 901 and the pose instruction skeletal landmark 902 from the skeletal landmarks estimated in S402. The shooting device 101 may generate such a display image 910 and display and transmit it. The display image 920 shown in Figure 9(c) is a pose instruction image superimposed on the live view image, which includes not the skeletal landmark of the person 103, but the body outline 924 when the skeletal landmark is moved according to the pose instruction. The body outline 924 may be represented by a dashed line or the like to show the contour of the body. The body outline 924 may also be represented by an image of the specified pose, generated using artificial intelligence (AI) or image processing. Furthermore, whether it is the foreground or background direction may be indicated using a caption 925 or by an image.
[0038] The display image 930 shown in Figure 9(d) is a display image generated using the same representation method as the display image 900 shown in Figure 9(a), and is a display image when there are two target landmarks 901 and 931. In display image 930, the skeletal landmark 902 for pose instruction is represented by a small inner white circle, indicating that target landmark 901 should be moved towards the back. Similarly, the skeletal landmark 932 for pose instruction is represented by a large inner white circle, indicating that target landmark 931 should be moved towards the front.
[0039] With the shooting system configured as described above, the subject can recognize how to change their posture. As a result, the shooting system configured as described above can assist the photographer in taking appropriate shots.
[0040] [Embodiment 2] Embodiment 1 described a method of displaying an image based on a pose instruction (pose instruction image) as a display image, directed toward the subject, person 103. However, a pose based on a pose instruction may be a pose that a natural person cannot assume. Therefore, this embodiment describes a method of generating a pose instruction image while taking into account the range of motion of the skeleton of the natural person 13 who is the subject.
[0041] Figure 10 is a block diagram showing an example of the logical configuration of the imaging device 101 and display device 102 according to Embodiment 2. The imaging device 101 according to Embodiment 2 is the same as the imaging device 101 according to Embodiment 1, except that the instruction acquisition unit 305 has been changed to an instruction acquisition unit 1005. Furthermore, the logical configuration of the display device 102 according to Embodiment 2 is the same as the logical configuration of the display device 102 according to Embodiment 1, so its explanation is omitted.
[0042] The instruction acquisition unit 1005 acquires the pose instructions from the photographer 104 received by the input unit 208. In addition to the above processing, the instruction acquisition unit 1005 uses skeletal range of motion data pre-stored in the ROM 201 or the like to determine whether the pose instructed by the photographer 104 can be assumed within the range of motion of a natural person's skeleton. If it is determined that a natural person can assume the pose instructed by the photographer 104, the instruction acquisition unit 1005 saves the acquired pose instruction data (pose instruction data) to the storage medium 204. If it is determined that a natural person cannot assume the pose instructed by the photographer 104, the instruction acquisition unit 1005 acquires an alternative pose that is close to the pose instructed by the photographer 104 and can be assumed by a natural person. In this case, the instruction acquisition unit 1005 saves the acquired alternative data as pose instruction data to the storage medium 204.
[0043] Figure 11 is a flowchart showing an example of the processing flow of the imaging device 101 according to Embodiment 2. In the processing of the flowchart shown in Figure 11, the same processes as those shown in the flowchart in Figure 4 are denoted by the same reference numerals and their explanation is omitted. First, the imaging device 101 executes the processes from S400 to S403. After S403, in S1101, the instruction acquisition unit 1005 determines, based on the pose instruction from the photographer 104 acquired in S403, whether the pose instructed by the photographer 104 can be taken within the range of motion of a natural human skeleton.
[0044] Figure 12 shows an example of a pose instruction input screen 1200 displayed on the display unit 205 according to Embodiment 2. In the pose instruction input screen 1200, the skeletal landmark 1201 is a skeletal landmark for pose instruction set based on input from the photographer 104. Specifically, the skeletal landmark 1201 is a skeletal landmark that indicates the destination of the skeletal landmark 1202 to be moved, as instructed by input from the photographer 104. The instruction acquisition unit 1005 determines whether it is possible to move the skeletal landmark 1202 to the position of the skeletal landmark 1201 by moving only the skeletal landmark 1202.
[0045] As described above, the skeletal landmark data stores a list of each skeletal landmark, the 3D coordinates of each skeletal landmark, and information about other skeletal landmarks connected to each skeletal landmark. In order to make the above determination, the instruction acquisition unit 1005 first uses the skeletal landmark data to calculate the 3D angle and length between skeletal landmarks. Next, the instruction acquisition unit 1005 uses the inverse equation of motion to determine whether or not skeletal landmark 1202 can be moved to the position of skeletal landmark 1201 by moving only skeletal landmark 1202.
[0046] Specifically, the instruction acquisition unit 1005 first identifies the skeletal landmarks that are directly or indirectly connected to the skeletal landmark 1202 of the object to be moved. When the skeletal landmark 1202 of the object to be moved, which corresponds to the right shoulder of the person 103, is moved upward, as in the example shown in Figure 12, the skeletal landmarks that are directly or indirectly connected to the skeletal landmark 1202 are also pulled upward. Specifically, in this case, the skeletal landmark 1206 corresponding to the right elbow and the skeletal landmark corresponding to the right wrist are also pulled upward. In addition, in this case, the skeletal landmark 1207, etc., which corresponds to the opposite shoulder (left shoulder) and is directly or indirectly connected to the skeletal landmark 1202 of the object to be moved, are similarly pulled in the direction of the skeletal landmark 1201 used for pose instruction.
[0047] Therefore, the instruction acquisition unit 1005 then uses the inverse equation of motion to calculate the amount of movement of skeletal landmarks other than the target landmark when the target landmark is moved. Subsequently, based on the result of this calculation, the instruction acquisition unit 1005 determines whether it is possible to move only the target landmark within the range of motion of a natural person's skeleton without moving any other skeletal landmarks.
[0048] Figure 13 shows an example of skeletal range of motion data according to Embodiment 2. The skeletal range of motion data includes the following items: part name 1301, direction of movement 1302, range of motion 1303, basic axis 1304, and movement axis 1305. Part name 1301 stores information as an item value indicating the name of the part that can be estimated as a skeletal landmark in S401. For example, the skeletal landmark corresponding to the shoulder girdle is skeletal landmark 1202 in Figure 12. Direction of movement 1302 stores information as an item value indicating the direction in which each skeletal landmark moves. Range of motion 1303 stores information as an item value indicating the range in which an average natural person can move in the direction stored as an item value in direction of movement 1302. Basic axis 1304 stores information indicating the direction that serves as the reference direction when calculating the angle of the direction of movement stored as an item value in direction of movement 1302. The moving axis 1305 stores information indicating the other direction that forms with the reference direction stored as an item value in the base axis 1304, which is used when calculating the angle of the direction of motion stored as an item value in the direction of motion 1302.
[0049] For example, in the skeletal landmark 1202 corresponding to the shoulder girdle, when the direction of movement is vertical, the basic axis corresponds to the line segment 1204 shown by the solid line, and the movement axis corresponds to the line segment 1205 shown by the dashed line. Therefore, the upward movement angle 1203 of the skeletal landmark 1202 corresponding to the shoulder girdle can be obtained by calculating the angle formed by the line segment 1204 and the line segment 1205.
[0050] The following is an example of a case where the movement angle 1203 of the skeletal landmark 1202 is 30 degrees in response to a pose instruction from photographer 104. According to the skeletal range of motion data shown as an example in Figure 13, the average upward range of motion of the skeletal landmark 1202 corresponding to the shoulder girdle in a natural human is 0 to 20 degrees. Therefore, it is impossible to move only the skeletal landmark 1202 to the position of the skeletal landmark 1201 used for upward pose instructions. Consequently, in this case, the judgment in S1101 determines that the pose instructed by photographer 104 cannot be achieved within the range of motion of a natural human skeleton.
[0051] If it is determined in S1101 that the pose instructed by photographer 104 can be assumed within the range of motion of a natural human skeleton, the camera 101 executes processes S404 and S405 to complete the flowchart shown in Figure 11. If it is determined in S1101 that the pose instructed by photographer 104 cannot be assumed within the range of motion of a natural human skeleton, the instruction acquisition unit 1005 executes process S1102. Specifically, in this case, in S1102, the instruction acquisition unit 1005 acquires an alternative pose that is close to the pose instructed by photographer 104 and can be assumed by a natural human. As an example, the case in which the movement angle 1203 of the skeletal landmark 1202 is 30 degrees in the pose instruction from photographer 104, and the average range of motion of a natural human is 0 to 20 degrees, will be described below. In this case, the pose instruction from photographer 104 is at an angle that cannot be assumed by an average natural human, but it may be possible for a person with a flexible body. The instruction acquisition unit 1005 takes these circumstances into consideration and acquires alternative poses that a natural person could take. In this case, the instruction acquisition unit 1005 stores the acquired alternative data as pose instruction data in the storage medium 204.
[0052] Following S1102, in S1103, the image generation unit 307 generates a pose instruction image using the data of alternative poses that a natural person can take, acquired in S1102, and the live view image and skeletal landmark data acquired in S400 or S401. The pose instruction image generated in S1103 is a pose instruction image corresponding to the alternative pose. The process in S1103 is the same as the process in S404 shown in Figure 4, except that alternative data is used instead of pose instruction data, so the explanation is omitted.
[0053] Figure 14 shows an example of a pose instruction image 1400 corresponding to an alternative pose, which is displayed on the display unit 205 according to Embodiment 2. The pose instruction image 1400 includes a slider bar 1401. The "small" setting on the slider bar 1401 minimizes the movement of the target landmark. In other words, "small" means moving the target landmark while minimizing the burden on the body of the subject, person 103. The burden on the body means that the amount of movement of the target landmark is set to the maximum value (including approximate maximum value) of the range of motion stored in the skeletal range of motion data, or slightly above that value.
[0054] In contrast, the "Large" setting on the slider bar 1401 maximizes the number of skeletal landmarks that can be moved while suppressing the burden on the body. Suppressing the burden on the body means, for example, setting the upper limit of the movement of the target landmark to about half of the range of motion stored in the skeletal range of motion data. For example, if the range of motion is 0 to 20 degrees, suppressing the burden on the body means setting the movement of the target landmark to about 0 to 10 degrees. Note that the threshold for determining the upper limit of the range of motion is not limited to about half, and may be given in advance as a default value of the imaging device 101, or may be input by the photographer 104.
[0055] For example, if the grip position of the slider bar 1401 is set to "large", the instruction acquisition unit 1005 will perform the following processing, for example. First, the instruction acquisition unit 1005 calculates the positions of skeletal landmarks 1206 and 1207 after moving skeletal landmark 1202 to the position of skeletal landmark 1201 so that it is approximately half the range of motion of skeletal landmark 1202. For example, the instruction acquisition unit 1005 calculates the positions of skeletal landmarks 1206 and 1207 after moving based on the three-dimensional angle and length between the skeletal landmarks using the inverse equation of motion.
[0056] Next, if the position of at least one of the skeletal landmarks 1206 and 1207 changes in the calculation results described above, the instruction acquisition unit 1005 performs the following processing. Specifically, the instruction acquisition unit 1005 calculates the position of another skeletal landmark connected to the skeletal landmark after it has moved, using the inverse equation of motion, when the skeletal landmark whose position has changed is moved so that it is approximately half of its range of motion. Thereafter, this calculation is repeated, for example, until the positions of all skeletal landmarks are determined. In the above description, the instruction acquisition unit 1005 is described as calculating the position of the skeletal landmark after movement based on the range of motion of the skeleton, but the method of acquiring the position of the skeletal landmark after movement is not limited to this. For example, the instruction acquisition unit 1005 may use a trained model obtained as a result of learning by deep learning to acquire the positions of skeletal landmarks such that the movement of all skeletal landmarks is approximately half of their range of motion.
[0057] Figure 14(a) shows an example of a pose instruction image 1400 corresponding to an alternative pose when the slider bar 1401 is pinched in the "large" position. Specifically, Figure 14(a) shows the position of each skeletal landmark before movement and the position of each skeletal landmark after movement when the burden on the subject person 103's body is minimized. As a result of considering how to minimize the burden on the subject person 103's body, the movement of almost all skeletal landmarks is instructed. Figure 14(c) shows an example of a pose instruction image 1400 corresponding to an alternative pose when the slider bar 1401 is pinched in the "small" position. Specifically, Figure 14(a) shows the position of each skeletal landmark before movement and the position of each skeletal landmark after movement when the burden on the subject person 103's body is minimized while moving only the target skeletal landmark as much as possible. Figure 14(b) shows an example of a pose instruction image 1400 corresponding to an alternative pose when the slider bar 1401 is held between "large" and "small".
[0058] After S1103, the imaging device 101 executes the process in S405 and transmits data of a display image, including a pose instruction image corresponding to an alternative pose generated in S1103, to the display device 102 via the network 105. After S405, the imaging device 101 completes the process shown in the flowchart in Figure 11. Figure 15 is a diagram showing an example of a display image 1500 displayed in the display unit 205 or display unit 215 according to Embodiment 2. The display images 1500 shown in Figures 15(a), (b), and (c) are display images corresponding to the display images 900, 910, and 920 shown in Figures 9(a), (b), and (d), respectively, so their description is omitted.
[0059] With the shooting system configured as described above, the subject can recognize how to change their posture. Furthermore, even if the pose based on the pose instruction is a pose that an average person could not assume, the shooting system configured as described above can generate and display a pose instruction image that is tailored to the characteristics of the subject. As a result, the shooting system configured as described above can assist the photographer in taking appropriate photographs.
[0060] [Embodiment 3] Embodiment 1 described a method for generating a display image using a live view image obtained by shooting from one direction. This embodiment describes a method for generating a display image using multiple live view images obtained by shooting from multiple directions. Figure 16 is a diagram showing an example of the configuration of the shooting system according to Embodiment 3. The configuration of the shooting system according to Embodiment 3 is the same as the configuration of the shooting system according to Embodiment 1, except that a shooting device 1601 is added to shoot the subject, person 103, from, for example, above. The shooting device 1601 is a device having both shooting and communication functions, such as a digital still camera, digital video camera, or smartphone.
[0061] Figure 17 shows the hardware configuration of the imaging device 101, display device 102, and imaging device 1601 according to Embodiment 3. The hardware configuration of the imaging device 101 and display device 102 according to Embodiment 3 is the same as the hardware configuration of the imaging device 101 and display device 102 according to Embodiment 1, so a description is omitted. The imaging device 1601 has a CPU 1700, ROM 1701, RAM 1702, communication unit 1703, storage medium 1704, and imaging unit 1707. These hardware components of the imaging device 1601 are connected to each other so as to be able to communicate via a bus 1706.
[0062] The CPU 1700 is a control unit consisting of at least one processor or circuit, and controls the entire imaging device 1601. The ROM 1701 is an electrically erasable and recordable memory, and stores various data and programs used in processing by the CPU 1700. The program referred to here is a computer program for executing various flowcharts described later in this embodiment. The RAM 1702 is a memory used as the work area of the CPU 1700, and stores data used in processing by the CPU 1700 and programs read from the ROM 1701.
[0063] The communication unit 1703 is an interface for communicating with external devices such as network equipment or USB devices, and performs data communication via the network 105, or transmission and reception of data with external devices. The storage medium 1704 is a non-volatile recording medium composed of semiconductor memory such as a memory card. The imaging unit 1707 is an image sensor composed of a CCD or CMOS element that converts an optical image into an electrical signal. The CPU 1700 also operates as a control unit that controls the imaging unit 1707.
[0064] Figure 18 is a block diagram showing an example of the logical configuration of the imaging device 101, display device 102, and imaging device 1601 according to Embodiment 3. The imaging device 1601 has an image acquisition unit 1801 and a transmission unit 1802 as its logical configuration. The image acquisition unit 1801 acquires images obtained by imaging by the imaging unit 1707 as live view images and stores the acquired live view image data in the storage medium 1704. The transmission unit 1802 transmits the live view image data acquired by the image acquisition unit 1801 to the imaging device 101 via the network 105. The imaging device 101 according to Embodiment 3 is a modified version of the imaging device 101 according to Embodiment 1, in which the image acquisition unit 300 and image generation unit 307 are replaced with an image acquisition unit 1800 and an image generation unit 1807. The image generation unit 1807 acquires live view images transmitted from the imaging device 1601 in addition to live view images obtained by imaging by the imaging unit 207. Details of the processing performed by the image generation unit 1807 will be described later. Furthermore, the logical configuration of the display device 102 according to Embodiment 3 is the same as that of the display device 102 according to Embodiment 1, so its explanation will be omitted.
[0065] Referring to Figure 4, the processing flow in the imaging device 101 according to Embodiment 3 will be explained. In the following explanation, processing that is the same as the processing according to Embodiment 1 will be omitted. First, at S400, the image acquisition unit 1800 acquires the live view image obtained by imaging unit 207 and the live view image transmitted from imaging device 1601. The data of the live view image acquired by the image acquisition unit 1800 is stored in the storage medium 204. The live view image acquired by the image acquisition unit 1800 is also displayed on the display unit 205 via the display control unit 309. Figure 19 is a diagram showing an example of a live view image 1900 displayed on the display unit 205 according to Embodiment 3. The live view image 1900 includes the live view image 500 shown in Figure 5 and the live view image 1920 obtained by receiving from imaging device 1601. The live view image 1920 includes an image 1921 of the subject, person 103.
[0066] After S400, the imaging device 101 executes the process of S401. Next, in S402, the image generation unit 1807 generates a skeletal landmark image by imaging the skeletal landmark data acquired in S401. The skeletal landmark image generated by the image generation unit 1807 is displayed on the display unit 205 via the display control unit 309. The image generation unit 1807 may also generate an image by superimposing the generated skeletal landmark image onto the live view image acquired in S400. In this case, the image generated by the image generation unit 1807 is displayed on the display unit 205 via the display control unit 309. Figure 20 shows an example of a skeletal landmark image 2000 displayed on the display unit 205 according to Embodiment 3. The skeletal landmark image 2000 includes a skeletal landmark image 710 generated by superimposing the skeletal landmark image 700 and the live view image 500, and a live view image 1920 obtained by receiving from the imaging device 1601.
[0067] In this embodiment, skeletal estimation based on the live view image 1920, i.e., the live view image received from the imaging device 1601, is not performed, but the embodiment is not limited to this. For example, the skeletal estimation unit 302 may perform skeletal estimation processing using a skeletal estimation DL model on the live view image received from the imaging device 1601. In this case, the image generation unit 1807 may also generate a landmark image by imaging the skeletal landmark data obtained as a result of the skeletal estimation processing. Alternatively, the image generation unit 1807 may generate an image by superimposing the generated landmark image onto the live view image received from the imaging device 1601 and display it on the display unit 205 via the display control unit 309.
[0068] After S402, the imaging device 101 executes the process in S403. Figure 21 is a diagram showing an example of a pose instruction input screen 2100 displayed on the display unit 205 according to Embodiment 3. The pose instruction input screen 2100 includes an image corresponding to the pose instruction input screen 800 shown in Figure 8(b). The photographer 104 selects a skeletal landmark as the target landmark by touching any of the multiple skeletal landmarks displayed on the skeletal landmark image 700 shown as an example in Figure 20.
[0069] When skeletal landmark 801 is selected as the target landmark, the imaging device 101 displays skeletal landmark 802 for pose instruction near skeletal landmark 801. The imaging device 101 also displays slider bars 815 as GUI components for pose instruction, which are used to input the amount of movement in each axis direction. The imaging device 101 displays an image 2110 on the pose instruction input screen 2100, which is created by superimposing skeletal landmark 2111 corresponding to skeletal landmark 801 and skeletal landmark 2112 for pose instruction corresponding to skeletal landmark 802 onto the live view image 1920. The positions of skeletal landmarks 2111 and 2112 can be calculated by pre-calibrating the positional relationship between the imaging device 1601 and the imaging device 101.
[0070] When the skeletal landmark 802 used for posing instructions moves due to input from the photographer 104, the corresponding skeletal landmark 2112 used for posing instructions also moves along with it. Similarly, when the skeletal landmark 2112 used for posing instructions moves due to input from the photographer 104, the corresponding skeletal landmark 802 used for posing instructions also moves along with it. This allows the photographer 104 to intuitively grasp the three-dimensional position of the skeletal landmarks used for posing instructions and move or specify their position.
[0071] Following S403, in S404, the image generation unit 1807 generates a pose instruction image using the live view image, pose instruction data, and skeletal landmark data acquired in S400, S401, or S403. The pose instruction image generated by the image generation unit 1807 is displayed on the display unit 205 via the display control unit 309 as a display image. Next, in S405, the transmission unit 308 transmits the display image data generated in S404 to the display device 102 via the network 105. The display device 102 receives the display image data transmitted in S405 and displays the display image on the display unit 215 for the subject, person 103.
[0072] Figure 22 is a diagram showing an example of a display image shown in the display unit 205 or display unit 215 according to Embodiment 3. Specifically, the display image 2200 shown in Figure 22(a) includes the display image 900 shown in Figure 9(a) and the display image 2210 generated based on the live view image 1920. Specifically, the display image 2200 is the same as the pose instruction input screen 2100 shown in Figure 21, with the GUI components for pose instruction removed. That is, the target landmark 2211 and the skeletal landmark 2212 for pose instruction correspond to the target landmark 901 and the skeletal landmark 902 for pose instruction.
[0073] The display image 2220 shown in Figure 22(b) includes the display image 910 shown in Figure 9(b) and the display image 2230 generated based on the live view image 1920. The display image 2220 shows only the target landmark and the pose-instructing skeletal landmark from the estimated skeletal landmarks. The target landmark 2231 and the pose-instructing skeletal landmark 2232 correspond to the target landmark 901 and the pose-instructing skeletal landmark 902. The display image 2240 shown in Figure 22(c) includes the display image 920 shown in Figure 9(c) and the display image 2250 generated based on the live view image 1920. The display image 2240 shows the body lines 2251 superimposed on the live view image 1920, representing the movement of the skeletal landmarks according to the pose instructions, rather than the target landmark and the pose-instructing skeletal landmark.
[0074] With the shooting system configured as described above, the subject can recognize how to change their posture. Furthermore, with the shooting system configured as described above, the subject can check how to change their posture by viewing images from multiple directions. As a result, the shooting system configured as described above can assist the photographer in taking appropriate photographs.
[0075] [Other embodiments] The technology of this disclosure can also be implemented by supplying a program that implements one or more of the functions of the embodiments described above to a system or device via a network or storage medium, and by a process in which one or more processors in the computer of that system or device read and execute the program. Furthermore, the technology of this disclosure can also be implemented by a circuit such as an ASIC that implements one or more functions.
[0076] Furthermore, within the scope of this disclosure, the technologies described herein allow for free combination of each embodiment, modification of any component of each embodiment, or omission of any component in each embodiment.
[0077] [Structure of this disclosure] This disclosure includes the following configurations, methods, and programs.
[0078] <Configuration 1> A skeletal estimation means for estimating skeletal landmarks in the image of a person being photographed, included in a captured image, Instruction acquisition means for acquiring instructions to change the pose based on the user's specification of the position of the skeletal landmark after movement, Image generation means for generating a pose instruction image corresponding to the change instruction, A transmission means for transmitting the signal of the pose instruction image to a display device positioned in a location where the person being photographed can be seen, A photographic device characterized by having the following features.
[0079] <Configuration 2> The change instruction acquired by the instruction acquisition means is based on the user's input of the direction in which to move the skeletal landmark to be moved. The imaging device described in configuration 1, characterized by the above.
[0080] <Structure 3> The input for the direction of movement of the skeletal landmark is performed in at least one of the following directions: the vertical and horizontal directions in the captured image, and a direction perpendicular to the plane corresponding to the captured image. The imaging device described in configuration 2, characterized by the above.
[0081] <Structure 4> If the user inputs the orthogonal direction as the direction in which to move the skeletal landmark, The image generation means generates the pose instruction image in which the amount of movement to be moved in the direction toward the front or the direction toward the back in the orthogonal direction is expressed. The imaging device described in configuration 3, characterized by the above.
[0082] <Composition 5> The image generation means generates the pose instruction image in which the representation of at least one of the shape, color, and transparency of the skeletal landmark has been changed when the skeletal landmark to be moved has been moved to the position specified by the user. A photographic device according to any one of configurations 1 to 4 characterized by the above.
[0083] <Composition 6> The instruction acquisition means determines, based on the range of motion of the skeleton and the angle and length between the skeletal landmarks, whether it is possible to move only the skeletal landmark to be moved to the position specified by the user after the move. A photographic device according to any one of configurations 1 to 5, characterized by the above.
[0084] <Composition 7> If the instruction acquisition means determines that it is not possible to move only the skeletal landmark to be moved to the position after the move, it acquires candidate positions to which the skeletal landmark can be moved, based on the range of motion of the skeleton and the angle and length between the skeletal landmarks. The image generation means generates the pose instruction image based on the candidate. The imaging device described in configuration 6, characterized by the above.
[0085] <Structure 8> The image generation means generates the pose instruction image in which at least the skeletal landmark of the object to be moved is represented from among the estimated skeletal landmarks. A photographic device according to any one of configurations 1 to 7, characterized by the above.
[0086] <Composition 9> The image generation means generates the pose instruction image in which only the skeletal landmark of the target of movement among the estimated skeletal landmarks is represented. The imaging device according to configuration 8, characterized by the above.
[0087] <Composition 10> The image generation means generates a pose instruction image that represents the posture of the person being photographed when the skeletal landmark to be moved is moved to the position specified by the user. A photographic device according to any one of configurations 1 to 9 characterized by the above.
[0088] <Composition 11> Display control means for displaying the signal of the pose instruction image on a display device positioned at a location visible to the user, The imaging device according to configuration 1, further comprising the above.
[0089] <Composition 12> The image generation means generates pose instruction images that have different characteristics from each other, with respect to the pose instruction image to be displayed on a display device positioned where the user can see it and the pose instruction image to be transmitted to a display device positioned where the person being photographed can see it. The imaging device described in configuration 11, characterized by the above.
[0090] <Method> A skeletal estimation process that estimates skeletal landmarks in the image of the subject of the photograph included in the captured image, An instruction acquisition step to acquire a pose change instruction based on the user's specification of the position of the skeletal landmark after movement, An image generation step that generates a pose instruction image corresponding to the aforementioned change instruction, A transmission step of transmitting the signal of the pose instruction image to a display device positioned in a location where the person being photographed can be seen, A method for controlling a photographic device, characterized by including the following:
[0091] <Program> A program for causing a computer to function as an imaging device described in any one of configurations 1 to 12.
[0092] <System> A photographic device as described in any one of configurations 1 to 12, A display device positioned in a location where the person being photographed can be seen, A shooting system having the following features. [Explanation of Symbols]
[0093] 101 Imaging device 302 Skeleton Estimation Unit 305 Instruction acquisition part 307 Image Generation Unit 308 Transmitter
Claims
1. A skeletal estimation means for estimating skeletal landmarks in the image of a person being photographed, included in a captured image, Instruction acquisition means for acquiring instructions to change the pose based on the user's specification of the position of the skeletal landmark after movement, Image generation means for generating a pose instruction image corresponding to the change instruction, A transmission means for transmitting the signal of the pose instruction image to a display device positioned in a location where the person being photographed can be seen, A photographic device characterized by having the following features.
2. The change instruction acquired by the instruction acquisition means is based on the user's input of the direction in which to move the skeletal landmark to be moved. The imaging apparatus according to claim 1, characterized by the following:
3. The input for the direction of movement of the skeletal landmark is performed in at least one of the following directions: the vertical and horizontal directions in the captured image, and a direction perpendicular to the plane corresponding to the captured image. The imaging apparatus according to claim 2, characterized by the following:
4. If the user inputs the orthogonal direction as the direction in which to move the skeletal landmark, The image generation means generates the pose instruction image in which the amount of movement to be moved in the direction toward the front or the direction toward the back in the orthogonal direction is expressed. The imaging apparatus according to claim 3, characterized by the following:
5. The image generation means generates the pose instruction image in which the representation of at least one of the shape, color, and transparency of the skeletal landmark has been changed when the skeletal landmark to be moved has been moved to the position specified by the user. The imaging apparatus according to claim 1, characterized by the following:
6. The instruction acquisition means determines, based on the range of motion of the skeleton and the angle and length between the skeletal landmarks, whether it is possible to move only the skeletal landmark to be moved to the position specified by the user after the move. The imaging apparatus according to claim 1, characterized by the following:
7. If the instruction acquisition means determines that it is not possible to move only the skeletal landmark to be moved to the position after the move, it acquires candidate positions to which the skeletal landmark can be moved, based on the range of motion of the skeleton and the angle and length between the skeletal landmarks. The image generation means generates the pose instruction image based on the candidate. The imaging apparatus according to claim 6, characterized by the following:
8. The image generation means generates the pose instruction image in which at least the skeletal landmark of the object to be moved is represented from among the estimated skeletal landmarks. The imaging apparatus according to claim 1, characterized by the following:
9. The image generation means generates the pose instruction image in which only the skeletal landmark of the target of movement among the estimated skeletal landmarks is represented. The imaging apparatus according to claim 8, characterized by the following:
10. The image generation means generates a pose instruction image that represents the posture of the person being photographed when the skeletal landmark to be moved is moved to the position specified by the user. The imaging apparatus according to claim 1, characterized by the following:
11. Display control means for displaying the signal of the pose instruction image on a display device positioned at a location visible to the user, The photographic apparatus according to claim 1, further comprising the following:
12. The image generation means generates pose instruction images that have different characteristics from each other, with respect to the pose instruction image to be displayed on a display device positioned where the user can see it and the pose instruction image to be transmitted to a display device positioned where the person being photographed can see it. The imaging apparatus according to claim 11, characterized by the following:
13. A skeletal estimation process that estimates skeletal landmarks in the image of the subject of the photograph included in the captured image, An instruction acquisition step to acquire a pose change instruction based on the user's specification of the position of the skeletal landmark after movement, An image generation step that generates a pose instruction image corresponding to the aforementioned change instruction, A transmission step of transmitting the signal of the pose instruction image to a display device positioned in a location where the person being photographed can be seen, A method for controlling a photographic device, characterized by including the following:
14. A program for causing a computer to function as a photographic device according to any one of claims 1 to 12.
15. A photographic device according to any one of claims 1 to 12, A display device positioned in a location where the person being photographed can be seen, A shooting system having the following features.