Image processing device, method for operating image processing device, and program for operating image processing device

WO2026160161A1PCT designated stage Publication Date: 2026-07-30FUJIFILM CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
FUJIFILM CORP
Filing Date
2026-01-07
Publication Date
2026-07-30

Smart Images

  • Figure JP2026000298_30072026_PF_FP_ABST
    Figure JP2026000298_30072026_PF_FP_ABST
Patent Text Reader

Abstract

This image processing device comprises a processor. The processor outputs a first template video, and, in cases where a first condition is satisfied while the first template video is being outputted, outputs a second template video.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device, method for operating the image processing device, and operating program for the image processing device.

[0001] The technology disclosed herein relates to an image processing device, a method for operating an image processing device, and an operating program for an image processing device.

[0002] The article "Introduction of FUJIFILM Japan INSTAX Biz Implementation Case Study [Toyota Alvark Tokyo Co., Ltd.] / Fujifilm <Internet URL: https: / / www.youtube.com / watch?v=JNZU1Jkl7lY> 2023 / 06 / 26" describes a service that generates a composite image by combining a pre-prepared template containing images of basketball players with photos taken by fans who visited the basketball game venue, and then hands the fan an instant film with the composite image printed on it on the spot.

[0003] Japanese Patent Publication No. 2004-080417 describes a digital camera with an image synthesis function, characterized by comprising an imaging unit, a display unit that combines and displays a moving image captured by the imaging unit and a moving image for synthesis, and a storage means that associates and stores in memory the constituent still images of the moving image obtained when the shutter is released and the constituent still images of the moving image for synthesis.

[0004] Japanese Patent Application Laid-Open No. 2005-223513 describes an image capturing apparatus characterized by including a photographing booth having a photographing space for accommodating a subject, photographing means for photographing a moving image of the subject in the photographing space, voice collecting means for collecting voice in the photographing space, decorative image data storage means, moving image synthesizing means, storage medium storage means, printing means, and discharging means. The decorative image data storage means stores predetermined background moving image data and predetermined foreground moving image data in advance. The moving image synthesizing means generates synthesized moving image data based on the subject moving image photographed by the photographing means, the background moving image data and the foreground moving image data stored in advance in the decorative image data storage means. The storage medium storage means synchronizes the synthesized moving image data generated by the moving image synthesizing means and the collected voice data based on the collected voice collected by the voice collecting means, and stores them in an information storage medium readable by a computer. The printing means prints a still image based on at least one still image data included in the synthesized moving image data synthesized by the moving image synthesizing means on a printing medium that can be attached to the information storage medium. The discharging means discharges the information storage medium and the printing medium.

[0005] Japanese Patent No. 7364956 describes an information processing apparatus having a processor, the processor determines the time to start synthesis for a stored first video, controls the start of recording of a second video photographed by a camera, synthesizes the first video and the second video based on the time to start synthesis of the first video, and displays an image to be displayed at the time to start synthesis of the first video during photographing by the camera.

[0006] One embodiment according to the technology of the present disclosure provides an image processing apparatus, an operation method of the image processing apparatus, and an operation program of the image processing apparatus that can raise the user's mood.

[0007] The image processing apparatus of the present disclosure includes a processor, and the processor outputs a first template video and outputs a second template video when a first condition is satisfied while the first template video is being output.

[0008] Preferably, when the first condition is met, the processor stops outputting the first template video and starts outputting the second template video.

[0009] Preferably, the processor acquires the live view, generates a composite video of at least a second template video and the live view, and outputs the generated composite video.

[0010] Preferably, the processor generates a composite video of the first template video and the live view, and outputs the generated composite video.

[0011] The second template video should preferably include decorative content.

[0012] The content preferably includes at least one of the following: a person or a character.

[0013] The first template video is preferably a video that contains at least some content.

[0014] It is preferable for the processor to determine that the first condition has been met when it has acquired specific information.

[0015] There are multiple types of specific information, and it is preferable that the processor determines which of the multiple second template videos to output, depending on the type of specific information acquired.

[0016] The processor acquires the live view, and it is preferable that the specific information relates to the video or audio acquired by the live view.

[0017] The processor acquires a live view of the subject, and the specific information is preferably the state of the subject in the video acquired by the live view or audio related to the subject.

[0018] The subject is the user, and specific information is preferably conveyed through user gestures or vocalizations.

[0019] Preferably, the processor acquires a live view, outputs at least one frame from a plurality of frames constituting the second template video, acquires a composite image of the frame and at least one frame from the plurality of frames constituting the live view, and outputs the composite image to a medium.

[0020] Preferably, the processor generates a composite video of at least a first template video and a live view, records the composite video, and outputs a code image relating to the recording destination of the composite video to the medium in addition to the composite image.

[0021] Output to a medium is preferably printing on a print medium.

[0022] The printing medium is preferably instant film.

[0023] The processor preferably outputs at least one of video and audio related to the timing of acquiring the composite image.

[0024] The processor preferably acquires a live view of the subject and acquires a composite image when the state of the subject shown in the live view satisfies the second condition.

[0025] Preferably, the processor acquires multiple composite images, derives image quality evaluation values ​​for the acquired composite images, and uses the composite image whose image quality evaluation value satisfies the third condition to select the composite image to output to the medium.

[0026] Preferably, the processor generates a composite video of the second template video and the live view, derives image quality evaluation values ​​for multiple frames constituting the composite video, and uses the frames whose image quality evaluation values ​​satisfy the fourth condition to select the composite image to be output to the medium.

[0027] The frame of the second template video to be used as the composite image is preferably the final frame of the second template video.

[0028] Preferably, the processor generates a first partially composite video of a first template video and a live view corresponding to the playback time of the first template video, and a second partially composite video of a second template video and a live view corresponding to the playback time of the second template video, and records a composite video composed of the first partially composite video and the second partially composite video.

[0029] The operation method of the image processing device of this disclosure includes outputting a first template video, and outputting a second template video if a first condition is met while the first template video is being output.

[0030] The operating program for the image processing device of this disclosure causes a computer to perform a process that includes outputting a first template video, and outputting a second template video if a first condition is met while the first template video is being output.

[0031] This figure shows the shooting and printing system and the implementation of a shooting and printing service using the shooting and printing system. This flowchart shows the service flow by the shooting and printing system. This figure shows the front view of the shooting equipment. This figure shows the first template video. This figure shows the generation of the first partially composite video of the live view and the first template video. This figure shows the generation of the second partially composite video of the live view and the second template video. This is a block diagram of the computers constituting the distribution server and the shooting equipment. This is a block diagram of the CPU processing unit of the distribution server. This figure shows the contents of the template video DB. This is a block diagram of the CPU processing unit of the shooting equipment. This figure shows the player selection screen. This figure shows the processing of the image recognition unit and image processing unit, etc. This figure shows a display that notifies the timing of acquisition of the composite image. This is a flowchart of the processing procedure of the shooting equipment. This figure shows the processing of the RW control unit of the second embodiment. This figure shows a composite image in which the link destination information of the composite video is embedded in a 2D code. This figure shows the processing of the third embodiment. This figure shows the processing of the fourth embodiment. This is a flowchart of another example of the service flow by the shooting and printing system. This figure shows the first template video which includes a character's self-introduction video as content. This figure shows a composite image in which text in real space is displayed inverted. This diagram shows the process of displaying the second template video inverted during shooting, and then inverting the composite image after shooting.

[0032] [First Embodiment] As an example, as shown in Figure 1, the shooting and printing system 2 comprises a distribution server 10, a shooting device 11, and an instant printer 12. The shooting and printing system 2 is a system for providing a service in which the shooting device 11 captures a still image CAI (see Figure 3) of customer C, which is then combined with the final frame FF (see both Figure 3) of the second template video TV2 to generate a composite image COI, the composite image COI is printed onto instant film 13 using the instant printer 12, and the instant film 13 is handed to customer C on the spot. The shooting device 11 is an example of an "image processing device" related to the technology of this disclosure. Customer C is an example of a "subject" and "user" related to the technology of this disclosure. The instant film 13 is an example of a "medium" and "printing medium" related to the technology of this disclosure.

[0033] Customer C is a fan of a basketball team who has come to an event venue, in this case a basketball game venue. A photo booth 15 is set up at the game venue. A flag 16 with the name of the basketball team written on it is attached to the wall of the photo booth 15. After telling the staff member S stationed at the photo booth 15 which basketball player P (see Figure 2) is his favorite, Customer C stands in the designated spot in front of the flag 16. Customer C is then photographed by the camera 11. Basketball player P is an example of a "person" related to the technology of this disclosure.

[0034] As an example, as shown in Figure 2, the service provided by the shooting and printing system 2 generally follows this flow. First, the distribution server 10, in response to a request from staff S, distributes and outputs the first template video TV1 of basketball player P, which customer C has conveyed to staff S at the shooting booth 15, to the shooting device 11. The shooting device 11 displays the first template video TV1 on the touch panel display 27 (see Figure 3) (step ST10). The first template video TV1 is a video that is displayed prior to the second template video TV2. The first template video TV1 includes a self-introduction video IV, which contains footage of basketball player P introducing himself. The self-introduction video IV is an example of "content" related to the technology of this disclosure. Displaying the first template video TV1 on the touch panel display 27 is an example of "outputting the first template video" related to the technology of this disclosure.

[0035] If customer C makes a specific gesture while the first template video TV1 is being streamed, the streaming server 10 stops streaming the first template video TV1 and instead starts streaming the second template video TV2. The camera 11 stops displaying the first template video TV1 and instead displays the second template video TV2 on the touch panel display 27. Specifically, if customer C makes a gesture of raising both hands while the first template video TV1 is being streamed (step ST11A), the streaming server 10 streams the second template video TV2A to the camera 11, and the camera 11 displays the second template video TV2A on the touch panel display 27 (step ST12A). On the other hand, if customer C makes a gesture of raising their right hand while the first template video TV1 is being distributed (step ST11B), the distribution server 10 distributes the second template video TV2B to the shooting device 11, and the shooting device 11 displays the second template video TV2B on the touch panel display 27 (step ST12B). In this way, the shooting device 11 determines which second template video TV2 to display on the touch panel display 27 from among multiple second template videos TV2A, TV2B, ..., depending on the type of gesture made by customer C. Displaying the second template video TV2 on the touch panel display 27 is an example of "distributing a second template video" according to the technology of this disclosure. Note that the specific gesture may be communicated to customer C verbally by staff S, or it may be announced to customer C with a flyer or the like while they are waiting in line for the shooting booth 15, or it may be displayed on the shooting screen 30 (see Figure 3).

[0036] The second template video TV2 includes a play video PV that shows basketball player P in action. More specifically, the play video PVA of the second template video TV2A is a video showing basketball player P dribbling. Also, the play video PVB of the second template video TV2B is a video showing basketball player P taking a free throw. Similar to the self-introduction video IV, the play video PV is also an example of "content" related to the technology disclosed herein.

[0037] When the second template video TV2 (play video PV) reaches its final frame FF, the shooting device 11 takes a picture of the captured image CAI (step ST13). The shooting device 11 generates a composite image COI of the captured image CAI and the final frame FF (step ST14). The shooting device 11 sends the composite image COI to the instant printer 12. The instant printer 12 prints the composite image COI onto the instant film 13 (step ST15). This completes the service for one customer C.

[0038] As an example, as shown in Figure 3, a tablet terminal is used as an example of the imaging device 11. The imaging device 11 includes a first camera unit 25 (see Figure 1) and a second camera unit 26, which are composed of a lens and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor, a touch panel display 27, and a speaker 28, etc.

[0039] The first camera unit 25 is located on the back of the shooting device 11, opposite the front of the device, where the touch panel display 27 is located. The first camera unit 25 is activated when the photographer views the live view LV displayed on the touch panel display 27 while taking a picture. In other words, the first camera unit 25 is a so-called rear camera. In contrast, the second camera unit 26 is located on the front of the shooting device 11. The second camera unit 26 is activated when the subject, in this case customer C, views the live view LV displayed on the touch panel display 27 while taking a picture, i.e., when taking a selfie. In other words, the second camera unit 26 is a so-called front camera.

[0040] The camera 11 is held and fixed on a stand 29 (see Figure 1) at a set position a certain distance from the center of the wall of the shooting booth 15, with its front facing the subject, customer C. Therefore, customer C can view the touch panel display 27 when the captured image CAI is being taken.

[0041] The touch panel display 27 displays the shooting screen 30. The shooting screen 30 displays a composite video CV of the first template video TV1 or the second template video TV2 (the second template video TV2 is exemplified in Figure 3) and the live view LV of the real space, such as the shooting booth 15 and customer C, which is currently being captured by the second camera unit 26 of the shooting device 11. While the second template video TV2 is in its final frame FF, customer C adjusts the composition, such as the positional relationship with basketball player P, while viewing the shooting screen 30 and following the appropriate advice of staff S. When the second template video TV2 reaches its final frame FF, the second camera unit 26 captures the captured image CAI. The shooting device 11 can be any device having a camera unit and a display, and may be a digital SLR camera, a digital still camera, a smartphone, or a notebook personal computer.

[0042] The first template video TV1 and the second template video TV2 are prepared in advance by the event organizer, in this case the organizer of a basketball game, for example, a public relations representative of the company that owns the basketball team. As an example, as shown in Figure 4, the first template video TV1 includes, in addition to the aforementioned self-introduction video IV, a two-dimensional code (here, a QR (Quick Response) code (registered trademark)) 35, a logo 36, text 37, and a pattern 38. The two-dimensional code 35 represents a link to access a bonus video of the basketball team or basketball player P, or a link to access a composite image COI, etc. The two-dimensional code 35 is an example of a "code image" relating to the technology of this disclosure. Note that one two-dimensional code 35 may represent multiple link destinations. Also, the two-dimensional code 35 is not limited to the QR code exemplified, but may be a barcode.

[0043] The logo 36 is, for example, a combination of the logo of a basketball team and the back number of the basketball player P. The text 37 is, for example, a character string that describes the name, height, weight, school of origin, etc. of the basketball player P. The pattern 38 is, for example, a spade mark. The pattern 38 may be a heart mark, a diamond mark, a circle, a cross mark, a zigzag pattern, a wave pattern, a vine pattern, or the like.

[0044] The self-introduction video IV is arranged in the right half area of the first template video TV1. The self-introduction video IV is arranged with the parts other than the basketball player P in a transparent state. For this reason, the live view LV (captured image CAI) is reflected around the basketball player P.

[0045] The two-dimensional code 35, the logo 36, the text 37, and the pattern 38 are arranged side by side in the vertically long area at the left end of the first template video TV1. The area where the two-dimensional code 35 and the like are arranged is not transparent. For this reason, the live view LV (captured image CAI) is not reflected in this area. Such arrangements of the self-introduction video IV and the like are merely examples and can be freely changed for each first template video TV1. Note that the second template video TV2 has almost the same configuration as the first template video TV1, except that the self-introduction video IV is changed to the play video PV, so the illustration and description are omitted.

[0046] The composite video CV of the first template video TV1 or the second template video TV2 and the live view LV becomes a video as if the basketball player P exists beside the customer C. Note that the live view LV is an RGB (red, green, blue) color video. Also, the contents such as the self-introduction video IV and the play video PV, and thus the first template video TV1 and the second template video TV2 are also RGB color videos.

[0047] The distribution server 10 is, for example, a server computer, a workstation, or the like. The distribution server 10 and the imaging device 11 are connected so as to be able to communicate with each other. Similarly, the imaging device 11 and the instant printer 12 are also connected so as to be able to communicate with each other. The connection form between the distribution server 10 and the imaging device 11, as well as the connection form between the imaging device 11 and the instant printer 12, is, for example, a WAN (Wide Area Network) such as the Internet or a public communication network. Alternatively, it is a short-range wireless communication such as Bluetooth (registered trademark). Note that a plurality of imaging devices 11 may be connected to the distribution server 10.

[0048] As an example, as shown in FIG. 5, the imaging device 11 synthesizes a live view LV corresponding to the reproduction time zone of the first template video TV1 into the first template video TV1, and generates a first partial composite video CV1 as a composite video CV. The imaging device 11 displays a shooting screen 30 including the generated first partial composite video CV1 on the touch panel display 27. Therefore, the customer C can visually recognize the first partial composite video CV1 through the shooting screen 30.

[0049] Similarly, as an example, as shown in FIG. 6, the imaging device 11 synthesizes a live view LV corresponding to the reproduction time zone of the second template video TV2 into the second template video TV2, and generates a second partial composite video CV2 as a composite video CV. The imaging device 11 displays a shooting screen 30 including the generated second partial composite video CV2 on the touch panel display 27. Therefore, the customer C can visually recognize the second partial composite video CV2 through the shooting screen 30. In FIG. 6, a state of generating a composite image COI of a final frame FF, which is one of a plurality of frames of the second template video TV2, and a captured image CAI, which is one of a plurality of frames of the live view LV, is shown.

[0050] As an example, as shown in Figure 7, the computer comprising the distribution server 10 and the imaging equipment 11 includes storage 45, memory 46, CPU (Central Processing Unit) 47, communication unit 48, display 49, and input device 50. These are interconnected via a bus line 51.

[0051] The storage 45 is a hard disk drive built into the computer comprising the distribution server 10 and the imaging equipment 11, or connected via cable or network. Alternatively, the storage 45 is a disk array consisting of multiple hard disk drives installed in series. The storage 45 stores control programs such as the operating system, various application programs (hereinafter referred to as AP (Application Program)), and various data associated with these programs. A solid-state drive may be used instead of a hard disk drive.

[0052] Memory 46 is work memory for the CPU 47 to execute processing. The CPU 47 loads the program stored in storage 45 into memory 46 and executes processing according to the program. In this way, the CPU 47 comprehensively controls each part of the computer. CPU 47 is an example of a "processor" related to the technology of this disclosure. Note that memory 46 may be built into the CPU 47.

[0053] The communication unit 48 is a network interface responsible for controlling the transmission of various types of information with external devices. The display 49 displays various screens. These screens are equipped with GUI (Graphical User Interface) operation functions. The computers comprising the distribution server 10 and the imaging device 11 receive operation instructions from the input devices 50 through the various screens. The input devices 50 include keyboards, mice, touch panels, and microphones for voice input. In the case of the imaging device 11, the touch panel display 27 serves as both the display 49 and the input device 50.

[0054] In the following explanation, each component of the computer constituting the distribution server 10 (storage 45 and CPU 47) will be denoted with the subscript "A," and each component of the computer constituting the imaging device 11 (storage 45 and CPU 47) will be denoted with the subscript "B" to distinguish them.

[0055] As an example, as shown in Figure 8, the operating program 55 is stored in the storage 45A of the distribution server 10. The storage 45A also stores a template video database (hereinafter referred to as DB (Data Base)) 56, etc. As an example, as shown in Figure 9, the template video DB 56 stores one first template video TV1 and multiple second template videos TV2 for each basketball player P. The corresponding gestures are registered in the second template videos TV2.

[0056] Returning to Figure 8, when the operating program 55 is started, the CPU 47A of the distribution server 10 works in cooperation with the memory 46 and the like to function as a request reception unit 60, a read / write (hereinafter referred to as RW (Read / Write)) control unit 61, an image acquisition unit 62, and a distribution control unit 63.

[0057] The request receiving unit 60 receives various requests from the camera equipment 11. These requests include a first designation request 83 (see Figure 11) that specifies basketball player P, and a second designation request 84 (see Figure 12) that specifies a second template video TV2. Each request includes an Identification Data (ID) to uniquely identify the camera equipment 11 that sent the request. The request receiving unit 60 outputs the requests to the RW control unit 61 and the distribution control unit 63.

[0058] The RW control unit 61 controls the storage of various data to the storage 45A and the reading of various data from the storage 45A. For example, the RW control unit 61 stores the first template video TV1 and the second template video TV2 in the template video DB 56 of the storage 45A.

[0059] When the request receiving unit 60 receives the first designated request 83, the RW control unit 61 reads the first template video TV1 of basketball player P specified in the first designated request 83 from the template video DB 56. The RW control unit 61 outputs the read first template video TV1 to the distribution control unit 63.

[0060] When a second designation request 84 is input from the request reception unit 60, the RW control unit 61 reads the second template video TV2 of basketball player P specified in the second designation request 84 from the template video DB 56. The RW control unit 61 outputs the read second template video TV2 to the distribution control unit 63.

[0061] The image acquisition unit 62 acquires the composite image COI generated by the imaging device 11. The image acquisition unit 62 outputs the composite image COI to the RW control unit 61. The RW control unit 61 stores the composite image COI in the storage 45A. As a modification, instead of sending the composite image COI from the imaging device 11 to the instant printer 12, the composite image COI stored in the storage 45A may be sent to the instant printer 12 (this is a modification; the following is a continuation of the first embodiment).

[0062] As an example, as shown in Figure 10, the storage 45B of the imaging device 11 stores an operating program 65. The operating program 65 is installed in the imaging device 11 by a staff member S or operator. The operating program 65 is an application program (AP) that causes the computer constituting the imaging device 11 to function as an "image processing device" according to the technology of this disclosure. In other words, the operating program 65 is an example of an "operating program for an image processing device" according to the technology of this disclosure.

[0063] When the operation program 65 is started, the CPU 47B of the imaging device 11 works in cooperation with the memory 46 and the like to function as a request transmission unit 70, a video reception unit 71, an image acquisition unit 72, an image recognition unit 73, an image processing unit 74, a display control unit 75, a transmission unit 76, and an audio control unit 77.

[0064] The request transmission unit 70 transmits various requests to the distribution server 10 in response to various operation instructions from staff S via the touch panel display 27. These requests include the first designation request 83 and the second designation request 84 mentioned above.

[0065] The video receiving unit 71 receives the first template video TV1, which is output from the distribution server 10 in response to the first designation request 83. The video receiving unit 71 also receives the second template video TV2, which is output from the distribution server 10 in response to the second designation request 84. The video receiving unit 71 outputs the received first template video TV1 and second template video TV2 to the image processing unit 74.

[0066] The image acquisition unit 72 acquires the live view LV that is sequentially transmitted from the second camera unit 26 at a predetermined frame rate. The image acquisition unit 72 outputs the live view LV to the image recognition unit 73 and the image processing unit 74. The image acquisition unit 72 also acquires the captured image CAI taken by the second camera unit 26. The image acquisition unit 72 outputs the captured image CAI to the image processing unit 74.

[0067] The image recognition unit 73 recognizes a specific gesture of customer C displayed on the live view LV from the image acquisition unit 72 using well-known image recognition technology. The image recognition unit 73 outputs the gesture recognition result 78 to the request transmission unit 70 (see also Figure 12). The recognition result 78 is an example of "specific information" relating to the technology of this disclosure. Furthermore, the case in which the image recognition unit 73 recognizes a specific gesture of customer C is an example of "the first condition being met" relating to the technology of this disclosure.

[0068] The image processing unit 74 combines the live view LV with the first template video TV1 to generate the first partially combined video CV1. The image processing unit 74 outputs the first partially combined video CV1 to the display control unit 75.

[0069] The image processing unit 74 combines the live view LV with the second template video TV2 to generate the second partially combined video CV2. The image processing unit 74 outputs the second partially combined video CV2 to the display control unit 75 (see also Figure 12). The image processing unit 74 also outputs the playback status of the second template video TV2 to the audio control unit 77.

[0070] Furthermore, the image processing unit 74 generates a composite image COI from the captured image CAI from the image acquisition unit 72 and the final frame FF of the second template video TV2. The image processing unit 74 outputs the composite image COI to the display control unit 75 and the transmission unit 76.

[0071] The display control unit 75 generates various screens and controls the display of the generated screens on the touch panel display 27. These screens include the aforementioned shooting screen 30 and a confirmation screen for allowing customer C to check the quality of the composite image COI.

[0072] The transmission unit 76 sends the synthesized image COI from the image processing unit 74 to the distribution server 10 and the instant printer 12. The audio control unit 77 controls the audio output to the speaker 28.

[0073] In response to operational instructions from staff member S via the touch panel display 27, the player selection screen 80, as shown in Figure 11, is displayed on the touch panel display 27 under the control of the display control unit 75. The player selection screen 80 displays basketball players P side by side. In addition, the player selection screen 80 has checkboxes 81 next to each basketball player P that allow the user to selectively select one of the basketball players P.

[0074] Staff member S checks the checkbox 81 for the basketball player P that customer C has specified, and then selects the designation button 82. In response to this designation button 82 selection, the request transmission unit 70 sends the first designation request 83 for basketball player P to the distribution server 10. The first designation request 83 is received by the request reception unit 60 of the distribution server 10 and output from the request reception unit 60 to the RW control unit 61. The RW control unit 61 reads the first template video TV1 for basketball player P corresponding to the first designation request 83 from the template video DB 56 and outputs it to the distribution control unit 63. The distribution control unit 63 outputs the first template video TV1 to the shooting device 11.

[0075] As an example, as shown in Figure 12, the request transmission unit 70 sends a second designation request 84 corresponding to the recognition result 78 to the distribution server 10. The second designation request 84 is received by the request reception unit 60 of the distribution server 10 and output from the request reception unit 60 to the RW control unit 61. The RW control unit 61 reads the second template video TV2 of basketball player P corresponding to the second designation request 84 from the template video DB 56 and outputs it to the distribution control unit 63. The distribution control unit 63 outputs the second template video TV2 to the shooting device 11. In Figure 12, an example is shown in which the image recognition unit 73 recognizes a gesture of customer C raising both hands.

[0076] As an example, as shown in Figure 13, when the second template video TV2 reaches its final frame FF, the display control unit 75 displays a countdown display 85 on the shooting screen 30 to indicate the timing for capturing the captured image CAI. The countdown display 85 starts 3 seconds before capturing the captured image CAI, and the numbers decrease by 1 second each time: "3", "2", "1". When it is time to capture the captured image CAI, "Smile!" is displayed. In addition, the audio control unit 77 outputs a countdown audio 86 from the speaker 28, as shown in the speech bubble, in conjunction with the countdown display 85 to indicate the timing for capturing the captured image CAI. In the shooting device 11, when the output of the countdown display 85 and the countdown audio 86 becomes "Smile!", an instruction to capture the captured image CAI is given, and the captured image CAI is captured by the second camera unit 26. The countdown display 85 and countdown sound 86 are examples of "video or audio related to the timing of acquiring a composite image" in the technology of this disclosure. The output of the countdown display 85 and countdown sound 86 may be started before the second template video TV2 reaches its final frame FF. The output of the countdown display 85 and countdown sound 86 may be set to "Smile!" at the moment the second template video TV2 reaches its final frame FF.

[0077] Next, the operation of the above configuration will be explained with reference to the flowchart shown in Figure 14 as an example. As shown in Figure 8, the CPU 47A of the distribution server 10 functions as a request reception unit 60, RW control unit 61, image acquisition unit 62, and distribution control unit 63 when the operation program 55 is activated. Also, as shown in Figure 10, the CPU 47B of the imaging device 11 functions as a request transmission unit 70, video reception unit 71, image acquisition unit 72, image recognition unit 73, image processing unit 74, display control unit 75, transmission unit 76, and audio control unit 77 when the operation program 65 is activated.

[0078] Customer C visits the photo booth 15 to receive the photo printing service provided by the photo printing system 2. Customer C then tells staff member S their favorite basketball player P. Staff member S operates the touch panel display 27 of the photo equipment 11 to display the player selection screen 80 shown in Figure 11. Staff member S checks the checkbox 81 for customer C's favorite basketball player P, and then selects the selection button 82. This sends the first selection request 83 from the request transmission unit 70 to the distribution server 10 (step ST100).

[0079] In the distribution server 10, the request reception unit 60 receives the first designation request 83. The first designation request 83 is output from the request reception unit 60 to the RW control unit 61. The RW control unit 61 then reads the first template video TV1 of basketball player P corresponding to the first designation request 83 from the template video DB 56. The first template video TV1 is output from the RW control unit 61 to the distribution control unit 63, and under the control of the distribution control unit 63, it is output for distribution to the camera equipment 11.

[0080] In the camera 11, the first template video TV1 is received by the video receiving unit 71 (step ST110). The first template video TV1 is output from the video receiving unit 71 to the image processing unit 74.

[0081] In the image acquisition unit 72, the live view LV from the second camera unit 26 is acquired. The live view LV is output from the image acquisition unit 72 to the image recognition unit 73 and the image processing unit 74. In the image processing unit 74, as shown in Figure 5, the live view LV is combined with the first template video TV1 to generate the first partially combined video CV1 (step ST120). The first partially combined video CV1 is output from the image processing unit 74 to the display control unit 75.

[0082] The display control unit 75 generates a shooting screen 30 including the first partially synthesized video CV1. The shooting screen 30 including the first partially synthesized video CV1 is displayed on the touch panel display 27 under the control of the display control unit 75 and made available for viewing by customer C (step ST130).

[0083] Customer C stands in the designated spot in front of the flag 16 in the shooting booth 15 and adjusts the composition while looking at the shooting screen 30 displayed on the touch panel display 27 of the shooting equipment 11. When the composition is appropriate, staff member S speaks to customer C and has customer C perform a specific gesture.

[0084] As shown in Figure 12, the image recognition unit 73 recognizes a specific gesture of customer C displayed on the live view LV (YES in step ST140). The gesture recognition result 78 is output from the image recognition unit 73 to the request transmission unit 70. Then, the request transmission unit 70 sends a second designation request 84 corresponding to the recognition result 78 to the distribution server 10 (step ST150).

[0085] In the distribution server 10, the request reception unit 60 receives the second designation request 84. The second designation request 84 is output from the request reception unit 60 to the RW control unit 61. The RW control unit 61 then reads the second template video TV2 corresponding to the gesture recognition result 78 from the template video DB 56. The second template video TV2 is output from the RW control unit 61 to the distribution control unit 63, and under the control of the distribution control unit 63, it is output for distribution to the camera equipment 11.

[0086] In the camera 11, the second template video TV2 is received by the video receiving unit 71 (step ST160). The second template video TV2 is output from the video receiving unit 71 to the image processing unit 74.

[0087] In the image processing unit 74, as shown in Figure 6, the live view LV is combined with the second template video TV2 to generate the second partially combined video CV2 (step ST170). The second partially combined video CV2 is output from the image processing unit 74 to the display control unit 75.

[0088] The display control unit 75 generates a shooting screen 30 including the second partially synthesized video CV2. The shooting screen 30 including the second partially synthesized video CV2 is displayed on the touch panel display 27 under the control of the display control unit 75 and made available for viewing by customer C (step ST180).

[0089] When the play video PV of the second template video TV2 reaches the final frame FF (YES in step ST190), the countdown display 85 and countdown audio 86 shown in Figure 13 are output under the control of the display control unit 75 and the audio control unit 77 (step ST200). In the shooting device 11, after the output of the countdown display 85 and countdown audio 86, the second camera unit 26 takes a captured image CAI. The captured image CAI is acquired by the image acquisition unit 72 (step ST210). The captured image CAI is transmitted from the image acquisition unit 72 to the image processing unit 74.

[0090] The image processing unit 74 generates a composite image COI of the captured image CAI and the final frame FF of the second template video TV2 (step ST220). The composite image COI is output from the image processing unit 74 to the display control unit 75 and the transmission unit 76.

[0091] The display control unit 75 generates a confirmation screen including the composite image COI. The confirmation screen is displayed on the touch panel display 27 under the control of the display control unit 75 and made available for viewing by customer C. If customer C's consent is obtained on the confirmation screen, the composite image COI is transmitted by the transmission unit 76 to the distribution server 10 and the instant printer 12 (step ST230). In the distribution server 10, the composite image COI is acquired by the image acquisition unit 62 and output from the image acquisition unit 62 to the RW control unit 61. Then, under the control of the RW control unit 61, it is stored in the storage 45A. The composite image COI is also printed on the instant film 13 by the instant printer 12. The instant film 13 with the composite image COI printed on it is handed to customer C by staff S. If customer C's consent is not obtained on the confirmation screen, the captured image CAI is re-shot.

[0092] As explained above, the CPU 47B of the shooting device 11 functions as a display control unit 75. The display control unit 75 displays the shooting screen 30, including the first template video TV1, on the touch panel display 27. If the image recognition unit 73 recognizes a specific gesture of customer C while the first template video TV1 is being displayed, the display control unit 75 displays the shooting screen 30, including the second template video TV2, on the touch panel display 27. Customer C can initially view the first template video TV1, and then view the second template video TV2, which is displayed after being switched by their own gesture. As a result, customer C can build up excitement for taking the captured image CAI after viewing the second template video TV2.

[0093] As shown in Figure 2, when the image recognition unit 73 recognizes a specific gesture of customer C, the display control unit 75 stops displaying the shooting screen 30 including the first template video TV1, and then starts displaying the shooting screen 30 including the second template video TV2. This allows customer C to clearly recognize that the display has switched from the first template video TV1 to the second template video TV2. This allows customer C to focus their attention on the second template video TV2. Alternatively, the display of the first template video TV1 may be continued on a small screen or the like without stopping the display of the first template video TV1, while the display of the second template video TV2 is started.

[0094] The image acquisition unit 62 acquires the live view LV. As shown in Figure 6, the image processing unit 74 generates a second partially composite video CV2, which is a composite video CV of at least the second template video TV2 and the live view LV. As shown in Figures 2 and 3, the display control unit 75 displays the shooting screen 30, which includes the generated second partially composite video CV2, on the touch panel display 27. This gives customer C the opportunity to view the second partially composite video CV2, making it possible to uplift customer C's mood.

[0095] Furthermore, as shown in Figure 5, the image processing unit 74 generates a first partially composite video CV1, which is a composite video CV of the first template video TV1 and the live view LV. The display control unit 75 displays the shooting screen 30, including the generated first partially composite video CV1, on the touch panel display 27. This gives customer C the opportunity to view the first partially composite video CV1, which can enhance customer C's mood.

[0096] As shown in Figure 2, the second template video TV2 includes a gameplay video PV. This makes it possible to further enhance the mood of customer C.

[0097] The gameplay video (PV) includes basketball player P. This makes it possible to further excite customer C.

[0098] Furthermore, as shown in Figure 2, the first template video TV1 is also a video that includes a self-introduction video IV. Therefore, it is possible to further enhance the mood of customer C.

[0099] As shown in Figure 10, when the request transmission unit 70 obtains the recognition result 78 of a specific gesture of customer C from the image recognition unit 73, it determines that the first condition for displaying the second template video TV2 has been met and sends a second designation request 84 to the distribution server 10. As a result, the shooting screen 30 including the second template video TV2 can be displayed on the touch panel display 27 without missing the timing.

[0100] As shown in Figure 2, there are multiple types of gestures. The display control unit 75 determines which of the multiple second template videos TV2 to display, according to the type of gesture. This gives customer C a sense of excitement, wondering which second template video TV2 will be displayed, whether it will be their favorite, or perhaps even a rare and valuable second template video TV2 that they have never seen before. It also increases customer C's desire to collect different second template videos, encouraging them to try different gestures next time to see different second template videos TV2. Consequently, it can lead to an increase in the number of customers C who visit the shooting booth 15, and ultimately, the number of customers C who come to the match venue.

[0101] The specific information concerns the video captured by the Live View LV. More specifically, the specific information is the state of customer C in the video captured by the Live View LV, and the state of customer C is determined by the gestures made by customer C. Therefore, the gestures made by customer C in real time in the shooting booth 15 can be used as a trigger for displaying the second template video TV2. Because customer C's gestures cause the first template video TV1 to switch to the second template video TV2, they can further enhance their mood.

[0102] As shown in Figures 3 and 6, the image processing unit 74 generates a composite image COI from the final frame FF of the second template video TV2 and the captured image CAI. The transmission unit 76 sends the composite image COI to the instant printer 12, which prints the composite image COI onto the instant film 13. As a result, customer C can enjoy the instant film 13 with the composite image COI printed on it, further enhancing customer C's mood. In particular, if the medium is a printed medium and the printed medium is instant film 13, the instant film 13 can be easily handed to customer C on the spot, increasing the immediacy of the photo print service. Note that the medium is not limited to a printed medium. For example, the composite image COI may be sent to customer C's mobile terminal as an email attachment. Also, the printed medium is not limited to instant film 13, but may be paper media such as postcards or A4 paper.

[0103] As shown in Figure 13, the display control unit 75 and the audio control unit 77 output a countdown display 85 and a countdown audio 86 indicating the timing for acquiring the captured image CAI, and consequently the composite image COI. Therefore, customer C can grasp the timing for capturing the captured image CAI. This reduces the chance of customer C missing the timing for capturing the captured image CAI. Note that instead of outputting both the countdown display 85 and the countdown audio 86, only one of them may be output.

[0104] The frame of the second template video TV2, which will be used as the composite image COI, is the final frame FF of the second template video TV2. Therefore, customer C can fully enjoy watching the second template video TV2 before proceeding to capture the captured image CAI.

[0105] [Second Embodiment] As an example, as shown in Figure 15, in the second embodiment, the RW control unit 61 records a composite video CV, which consists of a first partially composite video CV1 and a second partially composite video CV2 generated by the image processing unit 74, to the storage 45A. By recording the composite video CV in this way, the customer C can be given the opportunity to review the composite video CV later. The composite video CV shows a series of events during filming, such as the customer C's reactions to watching the first template video TV1 and the second template video TV2. For this reason, the composite video CV can be used as a trigger for the customer C to recall the excitement they felt during filming.

[0106] In the second embodiment, as shown in Figure 16 as an example, it is preferable to embed the link destination information of the synthesized video CV into the two-dimensional code 35. This allows customer C to easily access the synthesized video CV.

[0107] [Third Embodiment] In the first embodiment described above, the captured image CAI, which is captured in sync with the countdown display 85 and the countdown sound 86 "Smile!", is used to generate the composite image COI, but the embodiment is not limited to this.

[0108] As an example, as shown in Figure 17, in the third embodiment, the CPU 47B functions as a first determination unit 90 and a shooting instruction unit 91, in addition to the processing units 70 to 76 of the first embodiment (only the image acquisition unit 72 is shown in Figure 17). The first determination unit 90 receives the live view LV from the image acquisition unit 72. The first determination unit 90 also receives the first determination condition 92. The first determination condition 92 is stored in the storage unit 45B and read from the storage unit 45B to be input to the first determination unit 90. The first determination condition 92 is a condition for determining whether the state of customer C is suitable for shooting. The first determination condition 92 includes multiple determination items, such as whether customer C is smiling, whether customer C's gaze is directed towards the second camera unit 26, whether customer C does not have red eyes, whether customer C's face is not dark, and whether customer C is still.

[0109] The first determination unit 90 analyzes the live view LV frame when, for example, the countdown display 85 and the countdown sound 86 output "Smile!". It then determines whether the customer C shown in the frame satisfies all of the determination items of the first determination condition 92. The first determination unit 90 outputs a first determination result 93 to the shooting instruction unit 91 indicating whether the customer C satisfies all of the determination items of the first determination condition 92. The shooting instruction unit 91 transmits a shooting instruction 94 to the second camera unit 26 only if the first determination result 93 from the first determination unit 90 indicates that the state of customer C satisfies all of the determination items of the first determination condition 92. The second camera unit 26 receives the shooting instruction 94 and takes a captured image CAI. When the state of customer C satisfies all of the determination items of the first determination condition 92, this is an example of "when the state of the subject satisfies the second condition" according to the technology of this disclosure.

[0110] Thus, in the third embodiment, the captured image CAI, and consequently the composite image COI, is acquired when the state of customer C displayed on the live view LV satisfies all of the judgment items of the first judgment condition 92. This prevents capturing customer C in an unsuitable state, such as not smiling or not making eye contact. It eliminates the need to reshoot the captured image CAI. It enhances the value of the captured image CAI, and consequently the composite image COI. Note that the shooting instruction unit 91 may send a shooting instruction 94 to the shooting device 11 only if the state of customer C satisfies a set number of judgment items among the judgment items of the first judgment condition 92.

[0111] [Fourth Embodiment] As an example, as shown in Figure 18, in the fourth embodiment, the CPU 47B functions as an evaluation value derivation unit 100 and a second determination unit 101, in addition to the processing units 70 to 76 of the first embodiment (only the display control unit 75 is shown in Figure 18). The evaluation value derivation unit 100 receives from the image processing unit 74 multiple frames F_CV2 of the second partially synthesized video CV2, between the times when the output of the countdown display 85 and the countdown sound 86 becomes "3", "2", "1", and "Smile!". The evaluation value derivation unit 100 also receives an evaluation value derivation model 102. The evaluation value derivation model 102 is stored in storage 45B and read from storage 45B to be input to the evaluation value derivation unit 100. The evaluation value derivation model 102 is a trained model composed of a convolutional neural network or the like. The evaluation value derivation unit 100 inputs each of the multiple frames F_CV2 of the second partial composite video CV2 into the evaluation value derivation model 102, and outputs an image quality evaluation value 103 for each frame F_CV2 from the evaluation value derivation model 102. The image quality evaluation value 103 is a value that comprehensively evaluates the appropriateness of the image's exposure, shutter speed, white balance, and sharpness, as well as the quality of the customer C's condition (whether she is smiling, making eye contact, not blurred, etc.). For example, the image quality evaluation value 103 ranges from 0 to 100. The evaluation value derivation unit 100 outputs the image quality evaluation value 103 to the second determination unit 101. The multiple frames F_CV2 are an example of the "multiple composite images" and "multiple frames" related to the technology of this disclosure. Note that the multiple frames F_CV2 may also be multiple frames after the countdown display 85 and countdown sound 86 output "Smile!".

[0112] The second determination unit 101 receives the second determination condition 104. The second determination condition 104 is stored in the storage 45B and is read from the storage 45B and input to the second determination unit 101. The second determination condition 104 is, for example, that the image quality evaluation value 103 is 90 or higher. The second determination condition 104 is an example of the "third condition" and "fourth condition" related to the technology of this disclosure. The second determination unit 101 selects a frame F_CV2 from among a plurality of frames F_CV2 in accordance with the second determination condition 104, in which the image quality evaluation value 103 is 90 or higher. The second determination unit 101 outputs the second determination result 105, which includes the selected frame F_CV2, to the display control unit 75.

[0113] The display control unit 75 displays a list of frames F_CV2 included in the second determination result 105 and generates a selection screen 106 that allows the user to selectively select one frame F_CV2. The display control unit 75 displays the selection screen 106 on the touch panel display 27.

[0114] Customer C selects one frame F_CV2 from the list of frames F_CV2 displayed on the selection screen 106 to be used as a composite image COI to be printed on the instant film 13. Alternatively, the selection screen 106 may allow selection of multiple frames F_CV2, generating multiple composite image COIs and printing multiple composite image COIs onto multiple instant films 13.

[0115] Thus, in the fourth embodiment, the evaluation value derivation unit 100 derives image quality evaluation values ​​103 for a plurality of frames F_CV2 that constitute the second partially composited video CV2. The second determination unit 101 determines which of the plurality of frames F_CV2 have image quality evaluation values ​​103 that satisfy the second determination condition 104. The display control unit 75 then uses the frame F_CV2 whose image quality evaluation value 103 satisfies the second determination condition 104 to select the composite image COI to be printed on the instant film 13. As a result, the frame F_CV2 whose image quality evaluation value 103 satisfies the second determination condition 104 can be used as the composite image COI, and the value of the composite image COI can be increased.

[0116] In the first embodiment described above, customer C's gestures were used as an example of specific information, but the system is not limited to this. As an example, the system may be shown in Figure 19. That is, if customer C says "GO FUZIX!" while the first template video TV1 is being displayed (step ST11X), the distribution server 10 outputs the second template video TV2A to the camera 11, and the display control unit 75 displays the second template video TV2A on the touch panel display 27 (step ST12A). On the other hand, if customer C says "GREAT!" while the first template video TV1 is being displayed (step ST11Y), the distribution server 10 outputs the second template video TV2B to the camera 11, and the display control unit 75 displays the second template video TV2B on the touch panel display 27 (step ST12B). Thus, the specific information may be voice related to customer C, or more specifically, voice spoken by customer C. In this case, the CPU 47A functions as a voice recognition unit instead of the image recognition unit 73.

[0117] Similar to the first embodiment described above, the voice spoken by customer C in real time in the shooting booth 15 can be used as a trigger for the distribution output of the second template video TV2. Because customer C's voice switches the first template video TV1 to the second template video TV2, they can further enhance their excitement.

[0118] Furthermore, the specific information is not limited to the gestures or vocalizations of the example customer C. It may also be the gestures or vocalizations of staff member S. In this case, staff member S would be an example of a "subject" related to the technology of this disclosure. In addition, a template video switching button may be provided on the shooting screen 30, and when the switching button is operated by staff member S, the first template video TV1 may be switched to the second template video TV2. In this case, the operation signal of the switching button would be the specific information. In addition, when the playback time of the self-introduction video IV of the first template video TV1 reaches the set time, the first template video TV1 may be switched to the second template video TV2. In this case, the signal indicating that the playback time of the self-introduction video IV has reached the set time would be the specific information.

[0119] When switching the display from the first template video TV1 to the second template video TV2, a roulette display effect may be performed to randomly display multiple second template video TV2s, and then a second template video TV2 corresponding to specific information may be displayed.

[0120] The content of the first template video TV1 is not limited to the example basketball player P's self-introduction video IV. It could also be an introduction video of a basketball team, or a highlight video of a previous game. It could also be a video featuring multiple basketball players P. The content of the second template video TV2 is also not limited to the example basketball player P's gameplay video PV. It could also be an interview video of basketball player P, or a video of basketball player P as a child. The first template video TV1 may contain content related to the first basketball player P, and the second template video TV2 may contain content related to the second basketball player P. In this case, by continuing to display the first template video TV1 while displaying the second template video TV2, the video could start with only the first basketball player P, but then both the first and second basketball players P may appear and perform together.

[0121] The person is not limited to the example basketball player P. It could be a baseball player, soccer player, American football player, or other athlete. It could also be an entertainer such as an actor, talent, or idol. There can be multiple people. Furthermore, as shown in Figure 20 as an example, the self-introduction video IV in the first template video TV1, it is not limited to a person; it could also be a game character CH. It could also be an animal, a plant, or a building such as a temple, shrine, or tower, or even a vehicle such as a train, car, or motorcycle. It could also be a work of art such as a painting or sculpture. Multiple types of subjects, such as characters and people, or animals and plants, may be mixed together in one template video.

[0122] Furthermore, depending on the specifications or settings of the second camera unit 26, it may capture the real space with a horizontal inversion. In such cases, for example, as shown in the composite image COI in Figure 21, the basketball team logo on basketball player P's uniform and the basketball team logo on logo 36 are not horizontally inverted, but the basketball team logo on flag 16, which exists in the real space, is horizontally inverted. As a result, a composite image COI with an unnatural appearance is generated. For example, if customer C was wearing a uniform with the name of the basketball team's sponsor on it, and the name of the basketball team's sponsor existed in the real space, that name would also be horizontally inverted. This phenomenon is undesirable from the standpoint of respecting sponsors.

[0123] Therefore, as shown in Figure 22 as an example, when the image processing unit 74 generates the second partially composited video CV2, it horizontally flips the second template video TV2. This flipping of the second template video TV2 is performed when staff member S selects the flip button provided on the shooting screen 30. The image processing unit 74 generates a composite image COI by combining the captured image CAI and the final frame FF of the second template video TV2, and then applies the flipping process to the composite image COI. In this way, it is possible to generate a composite image COI that looks good and does not feel unnatural, in which the basketball team logo on the basketball player P's uniform and the basketball team logo on the flag 16 that exists in the real world are not horizontally flipped. If the name of the basketball team's sponsor exists in the real world, the sponsor can also be taken into consideration. In addition to the second template video TV2, the first template video TV1 may also be horizontally flipped to generate the first partially composited video CV1.

[0124] The CAI image may be captured by staff member S. Alternatively, customer C may be given a remote controller with a shutter release button and allowed to capture the CAI image. Furthermore, while a special booth at an event venue, such as the photography booth 15 at a basketball game venue, has been given as an example of a situation in which the CAI image may be captured, the situation is not limited to this. The technology disclosed herein may also be applied to situations in which a general user captures the CAI image using the front camera of their smartphone.

[0125] The CAI image may be a 3D computer graphic or an image created using a machine learning model. Furthermore, customer C is not limited to the single person exemplified. Multiple customers C, such as a couple, parent and child, or friends, may appear in the CAI image.

[0126] The hardware configuration of the computer constituting the imaging device 11 can be modified in various ways. For example, the imaging device 11 can be composed of multiple computers separated as hardware, in order to improve processing power and reliability. For example, the functions of the request receiving unit 70, video receiving unit 71, display control unit 75, transmission unit 76, and audio control unit 77, and the functions of the image recognition unit 73 and image processing unit 74 can be distributed among two computers. In this case, the imaging device 11 is composed of two computers. Alternatively, the distribution server 10 may handle all or part of the functions of the imaging device 11.

[0127] When the distribution server 10 is responsible for the functions of the image recognition unit 73 and the image processing unit 74, it transmits the live view LV and captured image CAI from the shooting device 11 to the distribution server 10. Alternatively, the distribution server 10 may be responsible for the function of the display control unit 75 and generate various screens such as the shooting screen 30 in the distribution server 10. In this case, the distribution server 10 distributes the screen data of the various screens to the shooting device 11 in the form of web distribution screen data created using a markup language such as XML (Extensible Markup Language). The shooting device 11 reproduces the various screens based on the screen data and displays them on the touch panel display 27. In this case, distributing and outputting the screen data of the shooting screen 30, including the first template video TV1 and the second template video TV2, to the shooting device 11 is an example of "outputting the first template video" and "outputting the second template video" according to the technology of this disclosure. Alternatively, other data description languages ​​such as JSON (JavaScript® Object Notation) may be used instead of XML.

[0128] Thus, the hardware configuration of the computer in the imaging device 11 can be appropriately changed according to the required performance, such as processing power, safety, and reliability. Furthermore, not only the hardware, but also application programs such as the operating programs 55 and 65 can, of course, be duplicated or distributed and stored on multiple storage devices for the purpose of ensuring safety and reliability.

[0129] In each of the above embodiments, the processing of each processing unit, such as the request receiving unit 60, RW control unit 61, image acquisition unit 62, distribution control unit 63, request transmission unit 70, video receiving unit 71, image acquisition unit 72, image recognition unit 73, image processing unit 74, display control unit 75, transmission unit 76, audio control unit 77, first determination unit 90, shooting instruction unit 91, evaluation value derivation unit 100, and second determination unit 101, is executed on any computer. Furthermore, any computer may execute these processes using a processor as hardware, a program as software, or a combination thereof. In that case, the processor is configured to cooperate with the program to execute the various processes in each of the above embodiments, and can function as each unit or means in each of the above embodiments. Also, the execution order of the processing by the processor is not limited to the order described and may be changed as appropriate. Any computer may be a general-purpose computer, a computer for a specific application, a workstation, or any other system capable of executing each of the processes.

[0130] A processor may consist of one or more hardware components, and the type of hardware is not limited. For example, a processor may consist of hardware such as the example CPUs 47A and 47B, programmable logic devices such as MPUs (Micro Processing Units) and FPGAs (Field Programmable Gate Arrays), dedicated circuits for performing specific processing such as ASICs (Application Specific Integrated Circuits), GPUs (Graphic Processing Units), or NPUs (Neural Processing Units). Furthermore, the type of hardware may be a combination of different types of hardware. When multiple hardware components are configured to perform one or more processes of a processor, these multiple hardware components may reside in physically separate devices or in the same device. Furthermore, in any embodiment, the order of each process performed by the processor is not limited to the order described above and may be changed as appropriate. The hardware is composed of an electrical circuit (circuitry) or the like, which is a combination of circuit elements such as semiconductor elements.

[0131] Furthermore, the program may be firmware or software such as microcode. Alternatively, the program may be, for example, a set of program modules, each function of which may be implemented by a processor configured to perform its respective function. The program may be program code or multiple code segments stored on one or more non-temporary computer-readable media (e.g., storage media or other storage devices). The program may be divided and stored on multiple non-temporary computer-readable media located in physically separate devices. Program code or code segments may represent any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. Program code or code segments may be connected to other code segments or hardware circuits by sending and receiving information, data, arguments, parameters, or memory contents.

[0132] From the above description, the technology described in the following supplementary information can be understood.

[0133] [Addendum 1] An image processing device comprising a processor, wherein the processor outputs a first template video, and outputs a second template video if a first condition is met while the first template video is being output. [Addendum 2] The image processing device according to Addendum 1, wherein the processor stops outputting the first template video and starts outputting the second template video if the first condition is met. [Addendum 3] The image processing device according to Addendum 1 or Addendum 2, wherein the processor acquires a live view, generates a composite video of at least the second template video and the live view, and outputs the generated composite video. [Addendum 4] The image processing device according to Addendum 3, wherein the processor generates a composite video of the first template video and the live view, and outputs the generated composite video. [Addendum 5] The image processing device according to any one of Addendum 1 to Addendum 4, wherein the second template video includes decorative content. [Addendum 6] The image processing device according to Addendum 5, wherein the content includes at least one of a person and a character. [Addendum 7] The image processing apparatus according to Addendum 5 or Addendum 6, wherein the first template video is a video containing at least the content. [Addendum 8] The image processing apparatus according to any one of Addendum 1 to Addendum 7, wherein the processor determines that the first condition has been met when specific information is acquired. [Addendum 9] The image processing apparatus according to Addendum 8, wherein there are multiple types of specific information, and the processor determines which second template video to output from among multiple second template videos according to the type of specific information acquired. [Addendum 10] The image processing apparatus according to Addendum 8 or Addendum 9, wherein the processor acquires a live view, and the specific information relates to the video or audio acquired by the live view. [Addendum 11] The image processing apparatus according to Addendum 10, wherein the processor acquires a live view of a subject, and the specific information is the state of the subject in the video acquired by the live view or audio related to the subject.[Addendum 12] The image processing apparatus according to Addendum 11, wherein the subject is a user, and the specific information is a gesture by the user or a voice uttered by the user. [Addendum 13] The image processing apparatus according to any one of Addendums 1 to 12, wherein the processor acquires a live view, outputs at least one frame from a plurality of frames constituting the second template video, acquires a composite image of the frame and at least one frame from a plurality of frames constituting the live view, and outputs the composite image to a medium. [Addendum 14] The image processing apparatus according to Addendum 13, wherein the processor generates a composite video of at least the first template video and the live view, records the composite video, and outputs a code image relating to the recording destination of the composite video to the medium in addition to the composite image. [Addendum 15] The image processing apparatus according to Addendum 13 or Addendum 14, wherein the output to the medium is printing to a print medium. [Addendum 16] The image processing apparatus according to Addendum 15, wherein the print medium is instant film. [Addendum 17] The image processing apparatus according to any one of Addendum 13 to Addendum 16, wherein the processor outputs at least one of video and audio relating to the timing of acquiring the composite image. [Addendum 18] The image processing apparatus according to any one of Addendum 13 to Addendum 17, wherein the processor acquires a live view of a subject, and acquires the composite image when the state of the subject shown in the live view satisfies the second condition. [Addendum 19] The image processing apparatus according to any one of Addendum 13 to Addendum 18, wherein the processor acquires a plurality of composite images, derives image quality evaluation values ​​for the plurality of acquired composite images, and uses the composite image whose image quality evaluation value satisfies the third condition to select the composite image to be output to the medium.[Addendum 20] The image processing apparatus according to any one of Addendum 13 to 19, wherein the processor generates a composite video of the second template video and the live view, derives image quality evaluation values ​​for a plurality of frames constituting the composite video, and uses the frames from the plurality of frames whose image quality evaluation values ​​satisfy the fourth condition for selection of a composite image to be output to the medium. [Addendum 21] The image processing apparatus according to any one of Addendum 13 to 20, wherein the frame of the second template video to be used as the composite image is the final frame of the second template video. [Addendum 22] The image processing apparatus according to any one of Addendum 13 to 21, wherein the processor generates a first partial composite video of the first template video and the live view corresponding to the playback time of the first template video, and a second partial composite video of the second template video and the live view corresponding to the playback time of the second template video, and records a composite video composed of the first partial composite video and the second partial composite video.

[0134] The technology of this disclosure can be appropriately combined with the various embodiments and / or variations described above. Furthermore, it is understood that various configurations can be adopted without departing from the spirit of the invention, and the invention is not limited to the embodiments described above. In addition, the technology of this disclosure extends not only to programs, but also to storage media for non-temporarily storing programs, and to computer program products containing programs.

[0135] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0136] In this specification, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0137] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

Claims

1. An image processing device comprising a processor, wherein the processor outputs a first template video, and outputs a second template video if a first condition is met while the first template video is being output.

2. The image processing apparatus according to claim 1, wherein the processor stops outputting the first template video and starts outputting the second template video when the first condition is met.

3. The image processing apparatus according to claim 1, wherein the processor acquires a live view, generates a composite video of at least the second template video and the live view, and outputs the generated composite video.

4. The image processing apparatus according to claim 3, wherein the processor generates a composite video of the first template video and the live view, and outputs the generated composite video.

5. The image processing apparatus according to claim 1, wherein the second template video includes decorative content.

6. The image processing apparatus according to claim 5, wherein the content includes at least one of a person and a character.

7. The image processing apparatus according to claim 5, wherein the first template video is a video containing at least the content.

8. The image processing apparatus according to claim 1, wherein the processor determines that the first condition is met when it acquires specific information.

9. The image processing apparatus according to claim 8, wherein there are multiple types of the specific information, and the processor determines which second template video to output from among the multiple second template videos according to the type of specific information acquired.

10. The image processing apparatus according to claim 8, wherein the processor acquires a live view, and the specific information relates to video or audio acquired by the live view.

11. The image processing apparatus according to claim 10, wherein the processor acquires a live view of a subject, and the specific information is the state of the subject in the video acquired by the live view or sound related to the subject.

12. The image processing apparatus according to claim 11, wherein the subject is a user, and the specific information is a gesture or a voice uttered by the user.

13. The image processing apparatus according to claim 1, wherein the processor acquires a live view, outputs at least one frame from a plurality of frames constituting the second template video, acquires a composite image of the frame and at least one frame from a plurality of frames constituting the live view, and outputs the composite image to a medium.

14. The image processing apparatus according to claim 13, wherein the processor generates a composite video of at least the first template video and the live view, records the composite video, and outputs a code image relating to the recording destination of the composite video to the medium in addition to the composite image.

15. The image processing apparatus according to claim 13, wherein the output to the medium is printing to a printing medium.

16. The image processing apparatus according to claim 15, wherein the printing medium is instant film.

17. The image processing apparatus according to claim 13, wherein the processor outputs at least one of video and audio relating to the timing of acquiring the composite image.

18. The image processing apparatus according to claim 13, wherein the processor acquires a live view of a subject, and acquires the composite image when the state of the subject shown in the live view satisfies the second condition.

19. The image processing apparatus according to claim 13, wherein the processor acquires a plurality of composite images, derives image quality evaluation values ​​for the acquired plurality of composite images, and uses the composite image whose image quality evaluation value satisfies the third condition to select a composite image to output to the medium.

20. The image processing apparatus according to claim 13, wherein the processor generates a composite video of the second template video and the live view, derives image quality evaluation values ​​for a plurality of frames constituting the composite video, and uses the frames from the plurality of frames whose image quality evaluation values ​​satisfy the fourth condition to select a composite image to be output to the medium.

21. The image processing apparatus according to claim 13, wherein the frame of the second template video to be used as the composite image is the final frame of the second template video.

22. The image processing apparatus according to claim 13, wherein the processor generates a first partially composited video of the first template video and the live view corresponding to the playback time period of the first template video, and a second partially composited video of the second template video and the live view corresponding to the playback time period of the second template video, and records a composite video composed of the first partially composited video and the second partially composited video.

23. A method for operating an image processing device, which includes outputting a first template video, and outputting a second template video if the first condition is met while the first template video is being output.

24. An operating program for an image processing device that causes a computer to perform a process including outputting a first template video, and outputting a second template video if the first condition is met while the first template video is being output.