Information processing device, information processing method, and program
Patent Information
- Application Number
- PCT/JP2026/003180
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-01-29
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026003180_01102026_PF_FP_ABST
Abstract
Description
Information processing apparatus, information processing method, and program
[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program.
[0002] Japanese Unexamined Patent Publication No. 2021-114072 discloses an image editing apparatus including a storage means, a labeled image extracting means, and an editing means. A labeled image is an image obtained by adding a label including attribute information of an editing target and information about the image to an image containing the editing target. The storage means stores the labeled image. The labeled image extracting means extracts an image that satisfies a predetermined condition from the labeled images. The editing means edits the extracted image based on the predetermined condition.
[0003] Examples of the predetermined condition include, if the editing target is a person, the presence of a specific person or a predetermined story. The specific person may be, for example, a single person or a group of a plurality of persons. Examples of the predetermined story include various modes such as album style, document style, story style, and diary style. The story in each mode can be automatically generated using, for example, artificial intelligence or the like.
[0004] One embodiment according to the present disclosure provides an information processing apparatus, an information processing method, and a program that can obtain image content in which story properties are taken into consideration without requiring a user to perform complicated editing work on a plurality of images.
[0005] A first aspect according to the present disclosure is an information processing apparatus including a processor, wherein the processor acquires a plurality of images, acquires at least one of information about the plurality of images and a user instruction, fits two or more images included in the plurality of images, the two or more images matching a story based on at least one of the acquired information about the plurality of images and the acquired user instruction, into a layout matching the two or more images to generate image content, and the story is associated with the two or more images and / or the two or more image content.
[0006] A second aspect of the present disclosure is an information processing apparatus according to the first aspect, wherein the processor generates a story based on at least one of a plurality of images and user instructions.
[0007] A third aspect of the present disclosure is an information processing apparatus according to the first or second aspect, wherein the processor selects two or more images from a plurality of images based on a story.
[0008] A fourth aspect of the present disclosure is an information processing device relating to any one of the first to third aspects, wherein the processor determines a layout based on two or more images.
[0009] A fifth aspect of this disclosure is an information processing device relating to any one of the first to fourth aspects, wherein the story represents the relationship between two or more images and / or two or more image contents, and the relationship includes a time series of two or more images and / or a time series of two or more image contents.
[0010] A sixth aspect of the present disclosure is an information processing device according to any one of the first to fifth aspects, wherein the processor generates a plurality of stories based on at least one of a plurality of images and user instructions.
[0011] The seventh aspect of this disclosure is an information processing device according to the sixth aspect, wherein when multiple stories are generated, the multiple stories have different content due to branching.
[0012] The eighth aspect of this disclosure is an information processing device according to the seventh aspect, wherein the information is common to multiple stories before the branching point, and differs among multiple stories after the branching point.
[0013] A ninth aspect of the present disclosure is an information processing apparatus according to the seventh or eighth aspect, wherein one of a plurality of stories is selected according to an externally given selection instruction, and a processor selects two or more images based on elements contained in the selected story.
[0014] A tenth aspect of the present disclosure is an information processing device relating to any one of the first to ninth aspects, wherein the processor selects two or more images based on elements contained in a story, and the elements are words, phrases, and / or clauses.
[0015] An eleventh aspect of the present disclosure is an information processing device relating to any one of the first to tenth aspects, wherein a story is generated by a first trained model when at least one of information relating to a plurality of images and a user instruction is input to the first trained model.
[0016] A twelfth aspect of this disclosure is an information processing device relating to any one of the first to eleventh aspects, wherein information about a story and multiple images is input to a second trained model, and two or more images are selected from the multiple images based on the inference results obtained by the second trained model.
[0017] A thirteenth aspect of the present disclosure is an information processing device relating to any one of the first to eleventh aspects, wherein two or more images are selected from a plurality of images based on the degree of relevance between a story and a plurality of images.
[0018] A fourteenth aspect of the present disclosure is an information processing device relating to any one of the first to thirteenth aspects, wherein the processor modifies a story in accordance with an externally provided story modification instruction.
[0019] A 15th aspect of the present disclosure is an information processing device relating to any one of the first to 14th aspects, wherein the layout is obtained based on a story.
[0020] A sixteenth aspect of the present disclosure is an information processing device relating to any one of the first to fifteenth aspects, wherein the layout is obtained on a page-by-page basis based on story divisions.
[0021] The seventeenth aspect of this disclosure is an information processing device according to the sixteenth aspect, wherein a story is displayed on a screen, and the screen displays visible information that allows for the identification of divisions in the story displayed on the screen.
[0022] The eighteenth aspect of this disclosure is that multiple images are input into a story-generating AI.
[0023] This is an information processing device relating to any one of the first to seventeenth embodiments, including a virtual image generated by [the method described above].
[0024] The 19th aspect of this disclosure is an information processing device relating to any one of the first to seventeenth aspects, wherein the story is a sentence or a fixed phrase having a sentence-like structure.
[0025] A 20th aspect of the present disclosure is an image processing method comprising acquiring a plurality of images, acquiring at least one of information relating to the plurality of images and / or user instructions, and outputting image content of two or more images that fit a story based on the plurality of acquired images and at least one of the acquired user instructions, wherein the story relates to two or more images and / or two or more image contents.
[0026] A 21st aspect of the present disclosure is a program for causing a computer to perform a process, the process comprising acquiring a plurality of images, acquiring at least one of information relating to the plurality of images and / or user instructions, and outputting image content of two or more images that fit a story based on the acquired plurality of images and at least one of the acquired user instructions, wherein the story relates to two or more images and / or two or more image contents.
[0027] This is a conceptual diagram showing an example of how a user uses an information processing device and an example of the configuration of the information processing device. This is a conceptual diagram showing an example of the content of machine learning performed to obtain a caption generation AI. This is a conceptual diagram showing an example of the content of machine learning performed to obtain a story generation AI. This is a conceptual diagram showing an example of the content of machine learning performed to obtain a feature information generation AI. This is a conceptual diagram showing an example of the content of the third training data. This is a conceptual diagram showing an example of the content of processing performed by the processor using the caption generation AI. This is a conceptual diagram showing an example of the content of processing performed by the processor using the story generation AI. This is a conceptual diagram showing an example of the content of processing performed by the processor using the feature information generation AI. This is a conceptual diagram showing an example of the content of feature information. This is a conceptual diagram showing an example of rules used to select multiple captured images for a photobook and determine the layout. This is a conceptual diagram showing an example of a photobook generated by the processor 18 and displayed on the screen. This is a flowchart showing an example of the flow of the photobook generation process. This is a conceptual diagram showing an example of the content of machine learning performed to obtain a story generation AI that generates multiple stories. This is a conceptual diagram showing an example of the content of machine learning performed to obtain a story generation AI that generates the first story and multiple stories branched from the first story. This is a flowchart showing an example of the flow of the photobook generation process, including the process in which the processor modifies the story according to story modification instructions. This is a conceptual diagram illustrating an example of the process in which a virtual image is generated based on a story by a generative AI and stored in an NVM. It is also a conceptual diagram illustrating an example of the configuration of an information processing system.
[0028] Hereinafter, an example of an embodiment of the information processing device, information processing method, and program related to this disclosure will be described with reference to the attached drawings.
[0029] First, let's explain the terminology used in the following explanation.
[0030] CPU stands for "Central Processing Unit". GPU stands for "Graphics Processing Unit". GPGPU stands for "General-Purpose computing on Graphics Processing Units". APU stands for "Accelerated Processing Unit". TPU stands for "Tensor Processing Unit". NPU stands for "Neural Processing Unit". DSP stands for "Digital Signal Processor". RAM stands for "Random Access Memory". DRAM stands for "Dynamic Random Access Memory". NVM stands for "Non-volatile memory". ROM stands for "Read Only Memory". EEPROM stands for "Electrically Erasable Programmable Read Only Memory". MRAM stands for "Magnetoresistive Random Access Memory". ReRAM stands for "Resistive Random Access Memory". FRAM (registered trademark) stands for "Ferroelectric Random Access Memory". ASIC stands for "Application Specific Integrated Circuit". FPGA stands for "Field-Programmable Gate Array". CD-ROM stands for "Compact Disc Read Only Memory". DVD-ROM stands for "Digital Versatile Disc Read Only Memory". SSD stands for "Solid State Drive". HDD stands for "Hard Disk Drive". USB stands for "Universal Serial Bus".EL stands for "Electro-Luminescence". UI stands for "User Interface". I / F stands for "Interface". AI stands for "Artificial Intelligence". LAN stands for "Local Area Network". WAN stands for "Wide Area Network". 5G stands for "5th Generation Mobile Communication System". LLM stands for "Large Language Model". BLIP stands for "Bootstrapping Language-Image Pre-training". OFA stands for "One-For-All". CLIP stands for "Contrastive Language-Image Pretraining". ALIGN stands for "A Large-scale ImaGe and Noisy-text embedding". BLIP-2 is an abbreviation for "Bootstrapped Language-Image Pretraining 2". CG is an abbreviation for "Computer Graphics". Exif is an abbreviation for "Exchangeable Image File Format".
[0031] In the following description, a signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, a processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU, GPU, GPGPU, NPU, APU, TPU, or DSP.
[0032] In the following description, signed RAM refers to volatile memory that temporarily stores information and is used as work memory by the processor. An example of RAM is DRAM.
[0033] In the following description, signed NVM refers to non-volatile memory in which the stored information is retained even when the power is turned off, and is used for storing programs and data. Examples of NVM include flash memory, EEPROM, ROM, MRAM, ReRAM, or FRAM.
[0034] In the following description, signed storage refers to non-volatile storage devices that retain stored information even when the power is turned off, and is used for storing programs and data. Examples of storage include SSDs, HDDs, and magnetic tape drives.
[0035] In the following description, a signed external interface (I / F) is responsible for the exchange of various types of information between multiple devices connected to each other. An example of an external interface is the USB interface. External interfaces include communication interfaces. Communication interfaces include communication processors and antennas, etc., and are responsible for communication between multiple computers. Examples of communication standards applicable to communication interfaces include wireless communication standards such as 5G, Wi-Fi®, or Bluetooth®.
[0036] In this specification, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0037] Figure 1 shows an example of the configuration of the information processing device 10. As shown in Figure 1, the information processing device 10 is used by a user 12. The information processing device 10 comprises a processor 18, an NVM 20, RAM 22, and an external I / F 23. The processor 18, NVM 20, RAM 22, and external I / F 23 are connected to a bus 24. The processor 18 processes various types of information using the NVM 20 and RAM 22.
[0038] In this embodiment, the information processing device 10 is an example of the "information processing device" and "computer" as described in this disclosure. Also, in this embodiment, the processor 18 is an example of the "processor" as described in this disclosure. Also, in this embodiment, the user 12 is an example of the "external" as described in this disclosure.
[0039] An external interface 23 is connected to a UI device 16. The UI device 16 receives instructions from the user 12 and presents various information to the user 12.
[0040] The UI system device 16 includes a reception device 16A and a display 16B. An example of the reception device 16A is a keyboard, mouse, and / or touch panel. An example of the display 16B is a liquid crystal display or an EL display.
[0041] The reception device 16A receives instructions from the user 12. The processor 18 acquires the instructions received by the reception device 16A via the external I / F 23 and operates according to the acquired instructions. The display 16B displays the processing results, etc., of the processor 18 under the control of the processor 18.
[0042] Furthermore, storage 28 is connected to the external interface 23. The external interface 23 is responsible for the exchange of various information between the processor 18 and the storage 28. That is, the processor 18 writes various information to the storage 28 via the external interface 23 and reads various information from the storage 28 via the external interface 23.
[0043] Furthermore, a network 30 is connected to the external I / F 23. The external I / F 23 manages the transmission and reception of various types of information between an external device on the network 30 (such as a personal computer, server, and / or smart device) and the processor 18. Examples of the network 30 include WAN, LAN, and the like. The information processing apparatus 10 may be connected to the network 30 via a wireless method or a wired method. The processor 18 transmits various types of information to an external device on the network 30 via the external I / F 23, and acquires various types of information from an external device on the network 30 via the external I / F 23.
[0044] Conventionally, when the user 12 creates a digital photobook using image processing software or the like, it is necessary for the user 12 to manually select a plurality of images, arrange the selected plurality of images in chronological order, and consider the relationship between images so as to create a narrative.
[0045] However, as the number of images increases, the work becomes more complicated, and a lot of time and effort are required to create a photobook. In addition, the quality of the finished photobook greatly depends on the sense, knowledge, and experience of the creator of the photobook. Therefore, depending on the creator of the photobook, the relevance between the images included in the photobook may be poor. Poor relevance between images results in a photobook that lacks narrative.
[0046] In view of such circumstances, in the present embodiment, photobook generation processing is executed by the processor 18. A photobook generation program 32 is stored in the NVM 20. The processor 18 reads the photobook generation program 32 from the NVM 20, and executes the read photobook generation program 32 on the RAM 22. The photobook generation processing is implemented by the processor 18 executing the photobook generation program 32. In the present embodiment, the photobook generation program 32 is an example of the "program" according to the present disclosure.
[0047] NVM 20 stores a caption generation AI 34, a story generation AI 36, and a feature information generation AI 38. In the present embodiment, the caption generation AI 34 and the story generation AI 36 are examples of the "first trained model" according to the present disclosure, and the feature information generation AI 38 is an example of the "second trained model" according to the present disclosure.
[0048] The caption generation AI 34, the story generation AI 36, and the feature information generation AI 38 are used by the processor 18 in photobook generation processing as described below. In the following, in order to facilitate understanding of the present disclosure, description will be given with an example embodiment in which a digital-format photobook 84 (see FIG. 11) obtained based on a plurality of images captured at a wedding ceremony such as a wedding ceremony and a reception is generated by the photobook generation processing.
[0049] In the photobook generation processing, the processor 18 acquires a plurality of images. The processor 18 also acquires at least one of information related to the plurality of images and a user instruction. Then, the processor 18 generates image content in which two or more images included in the plurality of images, the two or more images matching a story based on at least one of the acquired information related to the plurality of images and the acquired user instruction, are fitted into a layout matching the two or more images. The story used herein is a story related to two or more images and / or two or more pieces of image content.
[0050] In the photobook generation processing, the processor 18 generates a story based on at least one of information related to the plurality of images and a user instruction. Here, for generating the story, the processor 18 uses the caption generation AI 34 and the story generation AI 36.
[0051] In the photobook generation process, the processor 18 selects two or more images from multiple images based on the generated story. Then, the processor 18 determines the layout based on the two or more selected images. Here, the processor 18 uses a feature information generation AI 38 to select two or more images and determine the layout.
[0052] Next, we will explain an example of machine learning performed to obtain the caption generation AI 34, the story generation AI 36, and the feature information generation AI 38, with reference to Figures 2 to 5. Figure 2 shows an example of machine learning performed to obtain the caption generation AI 34. Figure 3 shows an example of machine learning performed to obtain the story generation AI 36. Furthermore, Figures 4 and 5 show examples of machine learning performed to obtain the feature information generation AI 38.
[0053] As shown in Figures 2 to 4, machine learning to obtain the caption generation AI 34, story generation AI 36, and feature information generation AI 38 is performed by the learning execution device 40. The learning execution device 40 is implemented by a computer having a processor 42 and an NVM 44, etc. Since the hardware configuration of the learning execution device 40 is basically the same as the hardware configuration of the information processing device 10, a description of the hardware configuration of the learning execution device 40 is omitted here.
[0054] As shown in Figure 2, the NVM 44 stores multiple sets of first training data 46. The first training data 46 is a dataset of first example data 48 and first correct answer data 50. The first example data 48 includes an example image 48A. In this embodiment, an image showing a scene from a bridal ceremony is used as an example of an example image 48A. The example image 48A may be an image obtained by taking pictures at various actual bridal ceremonies, or it may be a CG generated by a generation AI or the like that simulates an image obtained by taking pictures at various actual bridal ceremonies.
[0055] The first correct answer data 50 includes a caption label 50A, which is a label indicating the caption of the example image 48A (i.e., the scene depicted in the example image 48A).
[0056] Assuming that multiple first training data sets 46 configured in this manner are stored in the NVM 44, the processor 42 in the learning execution device 40 acquires the first training data sets 46 from the NVM 44. The processor 42 then performs machine learning using the first training data sets 46.
[0057] In this case, for example, the processor 42 generates a caption generation AI 34 by optimizing the model 52 using backpropagation based on a plurality of first training data 46. For example, the model 52 may be BLIP, OFA, or a combination of CLIP and an existing LLM (e.g., GPT-3, GPT-3.5-turbo, or GPT-4). That is, the processor 42 generates a caption generation AI 34 by performing fine tuning on a conventionally known generation AI using a plurality of first training data 46.
[0058] Specifically, the processor 42 inputs the first example data 48 into the model 52 and compares the output result from the model 52 with the first correct answer data 50 to adjust multiple optimization variables within the model 52 in order to minimize the error. Examples of multiple optimization variables include weights that indicate the strength of connections between neurons (in other words, connection weights), and biases (in other words, offset values) that control the activation of neurons (in other words, values used to adjust the output of neurons).
[0059] The model 52 is optimized by the processor 42 repeatedly performing a learning process using multiple first training data 46. The caption generation AI 34 generated by this optimization of the model 52 is stored in the NVM 20 of the information processing device 10 (see Figure 1). As will be described in more detail later, the caption generation AI 34 stored in the NVM 20 is used by the processor 18 (see Figure 1).
[0060] As shown in Figure 3, the NVM 44 stores multiple sets of second training data 53. The second training data 53 is a dataset of second example data 54 and second correct answer data 56. The second example data 54 includes a group of captions 54A. The group of captions 54A consists of multiple captions 54A1, each with different content. The multiple captions 54A1 correspond to the multiple example images 48A included in the multiple first training data 46 shown in Figure 2. Each of the multiple captions 54A1 is generated by the caption generation AI 34 (see Figure 2) when each of the multiple example images 48A (see Figure 2) is input to the caption generation AI 34 (see Figure 2).
[0061] The second correct answer data 56 includes story 56A and multiple scene numbers 56B. Story 56A consists of multiple sentences. In the example shown in Figure 3, sentences 56A1 and 56A2 are shown as examples of multiple sentences. Each of sentences 56A1 and 56A2 is assigned a scene number 56B. Scene number 56B is a number that identifies a scene. The scene number 56B assigned to sentence 56A1 is a number that identifies the scene that sentence 56A1 expresses, and the scene number 56B assigned to sentence 56A2 is a number that identifies the scene that sentence 56A2 expresses. For example, if sentence 56A1 expresses a wedding ceremony scene, sentence 56A1 is assigned scene number 56B "1". Also, for example, if sentence 56A2 expresses a wedding reception scene, scene number 56B "2" is assigned.
[0062] Assuming that multiple second training data sets 53 configured in this manner are stored in the NVM 44, the processor 42 in the learning execution device 40 acquires the second training data sets 53 from the NVM 44. The processor 42 then performs machine learning using the second training data sets 53.
[0063] In this case, for example, the processor 42 generates a story generation AI 36 by optimizing the model 58 using backpropagation based on a plurality of second training data 53. An example of the model 58 is an existing LLM (e.g., GPT-3, GPT-3.5-turbo, or GPT-4). That is, the processor 42 generates a story generation AI 36 by performing fine tuning on a conventionally known generation AI using a plurality of second training data 53.
[0064] Specifically, in a similar manner to generating the caption generation AI 34 from the model 52 shown in Figure 2, the processor 42 inputs the second example data 54 into the model 58, compares the output result from the model 58 with the second correct answer data 56, and adjusts multiple optimization variables within the model 58 to minimize the error.
[0065] The model 58 is optimized by the processor 42 repeatedly performing a learning process using multiple second training data 53. The story generation AI 36 generated by this optimization of the model 58 is stored in the NVM 20 of the information processing device 10 (see Figure 1). As will be described in more detail later, the story generation AI 36 stored in the NVM 20 is used by the processor 18 (see Figure 1).
[0066] As shown in Figure 4, the NVM 44 stores multiple sets of third training data 59. The third training data 59 is a dataset of third example data 60 and third correct answer data 62.
[0067] As shown in Figure 5, the third example data 60 includes a group of example images 60A, a story 60B, and multiple scene numbers 60C. The group of example images 60A consists of multiple example images 60A1. An example of multiple example images 60A1 is the multiple example images 48A included in the multiple first training data 46 shown in Figure 2. The story 60B is a story generated by the caption generation AI 34 and the story generation AI 36 based on the multiple example images 60A1. That is, the caption generation AI 34 generates multiple captions based on the multiple example images 60A1, and the story generation AI 36 generates the story 60B based on the multiple captions generated by the caption generation AI 34.
[0068] Story 60B consists of multiple sentences. In the example shown in Figure 5, sentences 60B1 and 60B2 are shown as examples of multiple sentences. Each of sentences 60B1 and 60B2 is assigned a scene number 60C. Similar to scene number 56B shown in Figure 3, scene number 60C is a number that identifies a scene. The scene number 60C assigned to sentence 60B1 is a number that identifies the scene that sentence 60B1 expresses, and the scene number 60C assigned to sentence 60B2 is a number that identifies the scene that sentence 60B2 expresses.
[0069] The third correct answer data 62 includes multiple correct images 62A1, multiple scene labels 62B, multiple scene numbers 62C, and multiple associated scores 62D. The multiple correct images 62A1 are identical to the multiple example images 60A1. Each of the multiple example images 60A1 is assigned a scene label 62B, a scene number 62C, and an associated score 62D.
[0070] The scene label 62B attached to example image 60A1 is a label that indicates the content of the scene. For example, if the scene in example image 60A1 is "The bride and groom exchange vows in a church," then an example of a scene label 62B attached to example image 60A1 would be a label named "Wedding Ceremony." Also, for example, if the scene in example image 60A1 is "Toasting with guests," then an example of a scene label 62B attached to example image 60A1 would be a label named "Wedding Reception." Furthermore, for example, if the scene in example image 60A1 is "A casual party," then an example of a scene label 62B attached to example image 60A1 would be a label named "After-party."
[0071] Scene number 62C is an identification number assigned to each scene and is used to identify which scene it is. For example, if the content of correct image 62A1 is the "wedding ceremony" scene, then correct image 62A1 will be assigned scene number 62C "1". If the content of correct image 62A1 is the "wedding reception" scene, then correct image 62A1 will be assigned scene number 62C "2". If the content of correct image 62A1 is the "after-party" scene, then correct image 62A1 will be assigned scene number 62C "3".
[0072] The association score 62D is a numerical representation of the strength of the association between the content of the scene in the ground truth image 62A1 and the story 60B. An example of the association score 62D is the attention score calculated using story 60B as the query and ground truth image 62A1 as the key.
[0073] The association score 62D may be determined by an annotator or generated by a generating AI (not shown in the illustration). For example, the association score 62D is expressed as a number in the range of 0.00 to 1.00. A higher association score 62D assigned to the correct image 62A means that the content of the scene in the correct image 62A1 is strongly related to the story 60B. For example, if the story 60B is a story about a bridal ceremony and the content of the scene in the correct image 62A1 is "a scene of the bride and groom gazing at each other during the wedding ceremony," then the association score 62D is "0.95". Also, for example, if the story 60B is a story about a bridal ceremony and the content of the scene in the correct image 62A1 is "a scene of the back of the groom's friend at the wedding reception," then the association score 62D is "0.40". Furthermore, if, for example, the content of the scene shown in the correct image 62A1 is unrelated to the bridal ceremony (for example, a construction site visible from the window of the bridal venue), then the related score 62D will be "0.00".
[0074] Assuming that multiple third-level training data sets 59 configured in this manner are stored in the NVM 44, as shown in Figure 4, the processor 42 in the learning execution device 40 acquires the third-level training data sets 59 from the NVM 44. The processor 42 then performs machine learning using the third-level training data sets 59.
[0075] In this case, for example, the processor 42 generates a feature information generation AI 38 by optimizing the model 63 using backpropagation based on multiple third training data 59. Examples of existing multimodal relevance analysis models include CLIP, ALIGN, and / or BLIP-2. The processor 42 generates the feature information generation AI 38 by performing fine tuning on the existing multimodal relevance analysis model using multiple third training data 59.
[0076] Specifically, in a similar manner to generating the caption generation AI 34 from the model 52 shown in Figure 2, the processor 42 inputs the third example data 60 into the model 63, compares the output result from the model 63 with the third correct answer data 62, and adjusts multiple optimization variables within the model 63 to minimize the error.
[0077] The model 63 is optimized by the processor 42 repeatedly performing a learning process using multiple third training data 59. The feature information generation AI 38 generated by this optimization of the model 63 is stored in the NVM 20 of the information processing device 10 (see Figure 1). As will be described in more detail later, the feature information generation AI 38 stored in the NVM 20 is used by the processor 18 (see Figure 1).
[0078] Next, we will explain an example of inference performed by the caption generation AI 34, story generation AI 36, and feature information generation AI 38 in the photobook generation process, with reference to Figures 6 to 9. Figure 6 shows an example of inference performed by the caption generation AI 34. Figure 7 shows an example of inference performed by the story generation AI 36. Figures 8 and 9 show examples of inference performed by the feature information generation AI 38.
[0079] As shown in Figure 6, in the information processing device 10, the NVM 20 stores a group of captured images 64. The group of captured images 64 consists of a plurality of captured images 64A. Each of the plurality of captured images 64A is an image obtained by capturing each scene of the bridal ceremony with a camera. In this embodiment, the group of captured images 64 is an example of the "multiple images" according to the disclosure.
[0080] The processor 18 acquires each of the multiple captured images 64A from the NVM 20. The processor 18 inputs the captured images 64A acquired from the NVM 20 to the caption generation AI 34, causing the caption generation AI 34 to generate a caption 66 representing the scene depicted in the captured images 64A. The processor 18 acquires the caption 66 generated by the caption generation AI 34 and stores it in the NVM 20. This process is repeatedly executed, and as shown in Figure 7 as an example, multiple captions 66 with different content are stored in the NVM 20. Note that this is merely an example of how multiple captions 66 are stored in the NVM 20, and the multiple captions 66 may also be stored in the storage 28 and / or external devices on the network 30.
[0081] As shown in Figure 7, in the information processing device 10, the processor 18 acquires each of the multiple captions 66 from the NVM 20. The processor 18 inputs the multiple captions 66 acquired from the NVM 20 to the story generation AI 36, causing the story generation AI 36 to generate a story 68 and multiple scene numbers 69.
[0082] Story 68 is a text, consisting of multiple sentences. In the example shown in Figure 7, sentences 68A and 68B are shown as examples of multiple sentences. Each of sentences 68A and 68B is assigned a scene number 69. Scene number 69 is a number that identifies a scene. The scene number 69 assigned to sentence 68A identifies the scene that sentence 68A describes, and the scene number 69 assigned to sentence 68B identifies the scene that sentence 68B describes. For example, if sentence 68A describes a wedding ceremony scene, sentence 68A is assigned scene number 69 "1". Also, for example, if sentence 68B describes a wedding reception scene, scene number 69 is assigned "2".
[0083] The processor 18 stores a story-attached image group 70 (see Figure 8), which associates the story 68 and multiple scene numbers 69 generated by the story generation AI 36 with the captured image group 64 (see Figure 6), in the NVM 20. Note that this is merely an example of how the story-attached image group 70 is stored in the NVM 20; the story-attached image group 70 may also be stored in the storage 28 and / or in an external device on the network 30.
[0084] As shown in Figure 8, the NVM 20 stores a group of images with a story 70. The story 68 included in the group of images with a story 70 is a story related to the group of captured images 64 included in the group of images with a story 70. In the example shown in Figure 8, the group of captured images 64 is a plurality of captured images 64A obtained by capturing each scene of the wedding ceremony and reception, and the story 68 is a story related to the wedding ceremony and reception. That is, the sentence 68A included in the story 68 is a story related to the wedding ceremony, and the sentence 68B included in the story 68 is a story related to the reception.
[0085] The processor 18 acquires a group of images 70 with a story from the NVM 20. The processor 18 then inputs the group of images 70 with a story to the feature information generation AI 38, causing the feature information generation AI 38 to generate feature information 74. In this embodiment, the feature information 74 is an example of "information relating to multiple images" and "inference results" according to this disclosure.
[0086] The feature information 74 includes the captured image group 64 and information indicating the features of the captured image group 64. In the example shown in Figure 9, as an example of information indicating the features of the captured image group 64, a plurality of scene labels 76, a plurality of scene numbers 78, and a plurality of association scores 80 are shown. Each of the plurality of captured images 64A included in the captured image group 64 is assigned a scene label 76, a scene number 78, and an association score 80. In this embodiment, the association score 80 is an example of "similarity" as described in this disclosure.
[0087] The scene label 76 assigned to the captured image 64A is a label that indicates the content of the scene. For example, if the scene in the captured image 64A is "the bride and groom exchanging vows in a church," then an example of a scene label 76 assigned to the captured image 64A would be a label named "wedding ceremony." Also, for example, if the scene in the captured image 64A is "toasting with guests," then an example of a scene label 76 assigned to the captured image 64A would be a label named "wedding reception." Furthermore, for example, if the scene in the captured image 64A is "a casual party," then an example of a scene label 76 assigned to the captured image 64A would be a label named "after-party."
[0088] Scene number 78 is an identification number assigned to each scene and is used to identify which scene it is. For example, if the content of captured image 64A is the "wedding ceremony" scene, captured image 64A will be assigned scene number 78 "1". If the content of captured image 64A is the "wedding reception" scene, captured image 64A will be assigned scene number 78 "2". If the content of captured image 64A is the "after-party" scene, captured image 64A will be assigned scene number 78 "3". In this way, multiple scene numbers 78 are assigned to multiple captured images 64A, making it possible to identify the timeline of multiple captured images 64A from the multiple scene numbers 78.
[0089] The association score 80 is a numerical representation of the strength of the association between the content of the scene captured in the image 64A and the story 68 (see Figure 8). An example of the association score 80 is the attention score calculated using the story 68 as the query and the image 64A as the key. For example, the association score 80 is expressed as a numerical value in the range of 0.00 to 1.00. A higher association score 80 assigned to the image 64A indicates a stronger association between the content of the scene captured in the image 64A and the story 68.
[0090] For example, if story 68 is a story depicting a bridal ceremony, and the scene captured in image 64A is "a scene of the bride and groom gazing at each other during the wedding ceremony," then the association score 80 is "0.95." Also, for example, if story 68 is a story depicting a bridal ceremony, and the scene captured in image 64A is "a scene showing the back of the groom's friend at the wedding reception," then the association score 80 is "0.40." Furthermore, for example, if the scene captured in image 64A is unrelated to the bridal ceremony (for example, a construction site visible from the window of the wedding venue), then the association score 80 is "0.00."
[0091] Figure 10 shows an example of Rule 82, which defines the criteria for selecting the captured image 64A to be used in the photobook 84 (see Figure 11), and the criteria for determining the layout of the selected captured image 64A. For example, Rule 82 defines the correspondence between the pages of the photobook 84 (see Figure 11), the scene label 76, the scene number 78, the associated score 80, the captured image 64A, the layout 82A, and the rule explanation 82B. In the example shown in Figure 10, the layout 82A refers to the position and size of the captured image 64A within the page. Note that this is merely an example, and the layout 82A may also include the pattern and / or color of the mounting paper.
[0092] The processor 18 selects a plurality of captured images 64A to be used in the photobook 84 (see Figure 11) based on the feature information 74 by executing rule-based processing using rule 82. In this embodiment, the processor 18 selects a plurality of captured images 64A to be used in the photobook 84 (see Figure 11) based on the rule 82 shown in Figure 10 by referring to the feature information 74. Specifically, the processor 18 selects a plurality of captured images 64A that satisfy the priority and selection conditions defined by rule 82 from the plurality of captured images 64A included in the feature information 74, based on the scene label 76, scene number 78, and association score 80 assigned to the plurality of captured images 64A included in the feature information 74.
[0093] In other words, the processor 18 refers to the scene label 76, scene number 78, and association score 80 assigned to each captured image 64A, and preferentially selects the captured image 64A that has a high association with the story 68 (see Figure 8) based on rule 82. For example, among the captured images 64A assigned the same scene number 78, the processor 18 selects the captured image 64A with the highest association score 80 as the main image, and selects the captured image 64A with the next highest association score 80 as the auxiliary image.
[0094] To explain with a more specific example, for instance, an image 64A whose relevance score 80 exceeds a certain score range (for example, an image 64A with a relevance score 80 of 0.80 or higher) is judged to be the most important image for the story 68 and is selected as the main image. Also, an image 64A whose relevance score 80 is within a certain score range (for example, an image 64A with a relevance score 80 of 0.40 or higher but less than 0.80) is judged to be an image that is moderately relevant to the story 68 and is selected as a supplementary image. Furthermore, an image 64A whose relevance score 80 falls below a certain score range (for example, an image 64A with a score less than 0.40) is judged to be an image that is almost irrelevant to the story 68 and is selected as an image to be excluded or reduced. As a result, images 64A that match the story 68 and have a high relevance score 80 are preferentially used in the photobook 84 (see Figure 11).
[0095] Once the multiple captured images 64A to be used in the photobook 84 (see Figure 11) are determined in this way, the processor 18 further determines a layout 82A that is suitable for the multiple captured images 64A to be used in the photobook 84 (see Figure 11) by performing rule-based processing using rule 82.
[0096] Rule 82 includes constraints on the number of captured images 64A corresponding to the scene number 78 on each page, rules for prioritizing the placement of captured images 64A, and page transition rules. For example, Rule 82 includes rules such as placing the captured image 64A with the highest association score 80 largely in the center and placing auxiliary captured images 64A around it, reducing the size of captured images 64A with an association score 80 below a threshold (e.g., 0.40), carrying them over to the next page, or deleting them, and placing captured images 64A with different scene numbers 78 on different pages. The layout 82A used in the photobook 84 (see Figure 11) is determined according to the Rule 82 thus created.
[0097] To give a more specific example, captured images 64A whose relevant score 80 exceeds a certain score range are placed in the center of the page or in a prominent position, captured images 64A whose relevant score 80 is within a certain score range are used at the edge of the page or as background, and captured images 64A whose relevant score 80 falls below a certain score range are excluded or reduced in size.
[0098] The processor 18 uses rule 82 to determine a layout 82A that appropriately arranges the selected multiple captured images 64A on each page of the photobook 84 (see Figure 11). The layout 82A is formed on a page-by-page basis based on the divisions of the story 68 (see Figure 8). This is achieved by determining the layout 82A according to the page, scene label 76, scene number 78, and associated score 80 defined in rule 82.
[0099] As an example, as shown in Figure 11, the processor 18 generates a photobook 84 based on a plurality of captured images 64A selected according to rule 82 and a layout 82A determined according to rule 82. The photobook 84 contains a plurality of image contents 85. The processor 18 generates the image contents 85 by fitting the plurality of captured images 64A selected according to rule 82 into the page at the appropriate position and size. In other words, the processor 18 generates the image contents 85 by fitting the plurality of captured images 64A selected according to rule 82 into the layout 82A determined according to rule 82.
[0100] Here, the image content 85 is generated on a page-by-page basis by fitting multiple captured images 64A into a layout 82A determined according to rule 82. In the example shown in Figure 11, image content 85A is generated as the first page of the photobook 84, and image content 85B is generated as the second page of the photobook 84. That is, in the photobook 84 shown in Figure 11, multiple captured images 64A that fit the story 68 are arranged corresponding to each page (in the example shown in Figure 11, the first page and the second page, respectively).
[0101] For the purpose of facilitating understanding of this disclosure, a two-page image content 85 is provided as an example. However, this is merely an example, and the photobook 84 may contain three or more pages of image content 85, or it may contain only one page of image content 85.
[0102] The processor 18 displays the photobook 84 on screen 16B1 of the display 16B. Image content 85 is displayed on screen 16B1 page by page. The story 68 is also displayed on screen 16B1.
[0103] Of the story 68, the segmented story 68A, which is the part related to the image content 85, is displayed within the image content 85 on a page-by-page basis. That is, the segmented story 68A1, which is the part of the story 68 related to the multiple captured images 64A included in the image content 85A, is displayed within the image content 85A, and the segmented story 68A2, which is the part of the story 68 related to the multiple captured images 64A included in the image content 85B, is displayed within the image content 85B.
[0104] Story 68 represents the relationships (e.g., chronological order) of image contents 85A and 85B. Divided stories 68A1 and 68A2 are stories obtained by dividing story 68 along chronological order. Divided story 68A1 represents the relationships of multiple captured images 64A within image content 85A (e.g., the relationships of multiple subjects depicted in multiple captured images 64A and / or the chronological order of multiple captured images 64A), and divided story 68A2 represents the relationships of multiple captured images 64A within image content 85B (e.g., the relationships of multiple subjects depicted in multiple captured images 64A and / or the chronological order of multiple captured images 64A). In this embodiment, the multiple captured images 64A within image content 85 are an example of "two or more images" according to this disclosure, and image contents 85A and 85B are examples of "two or more image contents" according to this disclosure.
[0105] Within the image content 85, page numbers P1 and P2 are displayed to visually recognize page breaks corresponding to the progression of story 68. In the example shown in Figure 11, page number P1 is displayed in image content 85A, and page number P2 is displayed in image content 85B. Page numbers P1 and P2 are visible information that allows for the identification of breaks in story 68.
[0106] In this way, the processor 18 determines the page layout of the photobook 84 based on rule 82 and displays the generated image content 85 (image content 85A and 85B in the example shown in Figure 11) on the screen 16B1 according to the determined page layout.
[0107] In the example shown in Figure 11, image contents 85A and 85B are displayed side by side (i.e., in a list format). However, this is merely one example, and the display of image contents 85A and 85B may be switched in a sliding manner from one to the other in response to instructions received by the receiving device 16A (see Figure 1). Alternatively, multiple image contents 85 may be displayed one by one in response to instructions from the user 12, as if turning the pages of a bound photobook.
[0108] Furthermore, although the example shown in Figure 11 illustrates a form in which the photobook 84 is displayed on screen 16B1, the photobook 84 may also be stored by the processor 18 in the storage 28 and / or an external device on the network 30 when predetermined conditions are met (for example, when the condition that a storage instruction has been received by the receiving device 16A is met).
[0109] Next, an example of the flow of the photobook generation process executed by the processor 18 of the information processing device 10 will be described with reference to Figure 12. The flow of the photobook generation process shown in Figure 12 is an example of the "image processing method" according to this disclosure.
[0110] In the photobook generation process shown in Figure 12, first, in step ST10, the processor 18 acquires a group of captured images 64 from the NVM 20 (see Figure 6). After the process in step ST10 is executed, the photobook generation process proceeds to step ST12.
[0111] In step ST12, the processor 18 inputs each captured image 64A included in the captured image group 64 to the caption generation AI 34, causing the caption generation AI 34 to generate a caption 66 (see Figure 6). The processor 18 stores the caption 66 obtained for each captured image 64A in the NVM 20 (see Figure 6). After the processing in step ST12 is executed, the photobook generation process moves on to step ST14.
[0112] In step ST14, the processor 18 retrieves all captions 66 from the NVM 20 (see Figure 7). After the processing in step ST14 is completed, the photobook generation process moves on to step ST16.
[0113] In step ST16, the processor 18 inputs all the captions 66 acquired in step ST14 to the story generation AI 36, causing the story generation AI 36 to generate a story 68 and multiple scene numbers 69 (see Figure 7). The processor 18 then stores the story-attached image group 70 (see Figure 8), which associates the story 68 and multiple scene numbers 69 generated by the story generation AI 36 with the captured image group 64 (see Figure 6), into the NVM 20. After the processing in step ST16 is completed, the photobook generation process moves on to step ST18.
[0114] In step ST18, the processor 18 acquires the story-based image group 70 from the NVM 20 (see Figure 8). After the processing in step ST18 is completed, the photobook generation process moves on to step ST20.
[0115] In step ST20, the processor 18 inputs the story-related image group 70 acquired in step ST18 to the feature information generation AI 38, causing the feature information generation AI 38 to generate feature information 74 (see Figures 8 and 9). After the processing in step ST20 is completed, the photobook generation process moves on to step ST22.
[0116] In step ST22, the processor 18 performs rule-based processing using rule 82 to select multiple captured images 64A for each page from the multiple captured images 64A included in the feature information 74 generated in step ST20. Here, the multiple captured images 64A selected by the processor 18 are an example of "two or more images" according to this disclosure. After the processing in step ST22 is executed, the photobook generation process moves to step ST24.
[0117] In step ST24, the processor 18 determines a layout 82A that matches the multiple captured images 64A selected in step ST22 by performing rule-based processing using rule 82. After the processing in step ST24 is completed, the photobook generation process moves on to step ST26.
[0118] In step ST26, the processor 18 generates image content 85 for each page by fitting the multiple captured images 64A selected in step ST22 into the layout 82A determined in step ST24 (see Figure 11). The processor 18 then displays the image content 85 on screen 16B1 (see Figure 11). The processor 18 also displays the divided story 68A and page number 86 within the image content 85 displayed on screen 16B1 (see Figure 11). After the processing in step ST26 is completed, the photobook generation process ends.
[0119] As described above, in the information processing device 10, image content 85 is generated by fitting multiple captured images 64A that match a story 68 based on feature information 74, which is an example of information related to the captured image group 64, into a matching layout 82A (see Figures 6 to 11). The story 68 is a story related to the multiple captured images 64A and the multiple image content 85 within the image content 85. Here, since the determination of the story 68 and the layout 82A is performed without requiring detailed specifications from the user 12, the amount of work involved in selecting and arranging multiple captured images 64A for the photobook 84 from the captured image group 64 is greatly reduced. As a result, image content 85 that takes storytelling into consideration can be obtained without requiring the user 12 to perform complicated editing work on the captured image group 64.
[0120] Furthermore, the information processing device 10 generates a story 68 based on the captured image group 64 (see Figures 6 and 7). This makes it easier to obtain the story 68 compared to when the user 12 creates the story 68 by referring to the captured image group 64.
[0121] Furthermore, the information processing device 10 selects multiple captured images 64A for the photobook 84 from the group of captured images 64 based on the story 68. Therefore, multiple captured images 64A with high consistency with the content of the story 68 are selected, and a story-driven image content 85 can be obtained. As a result, when creating the photobook 84, the user 12 can save time while utilizing the multiple captured images 64A in a more attractive way.
[0122] Furthermore, in the information processing device 10, the layout 82A is determined based on multiple captured images 64A selected from the captured image group 64 for the photobook 84 by rule-based processing using rule 82. Therefore, it becomes easier to adjust the appearance of the photobook 84 to match the story 68, and the effort required for layout adjustments by the user 12 can be greatly reduced.
[0123] Furthermore, in the information processing device 10, the story 68 expresses the relationships between multiple captured images 64A and multiple image contents 85 within the image content 85. These relationships include the time series of the multiple captured images 64A selected in the manner described above, and the time series of the image contents 85A and 85B. Therefore, multiple captured images 64A can be arranged in a natural flow within the pages of the photobook 84. As a result, the consistency of the story 68 within the photobook 84 is enhanced, and a more user-friendly photobook 84 can be provided to the user 12.
[0124] Furthermore, in the information processing device 10, words such as "wedding ceremony" or "wedding reception," which are elements included in the story 68, are assigned to the captured image 64A as scene labels 76 and scene numbers 78 (see Figure 9). Then, based on the scene labels 76 and scene numbers 78 assigned to the captured image 64A, multiple captured images 64A for the photobook 84 are selected (see Figures 9 to 11). Therefore, by handling the text elements that constitute the story 68 in finer units (here, as an example, words), the information processing device 10 can evaluate the affinity between the story 68 and the captured images 64A in a more detailed manner and then select multiple captured images 64A for the photobook 84.
[0125] Furthermore, since the layout 82A is obtained based on the story 68 in the information processing device 10, the user 12 can easily create image content 85 with a strong narrative while minimizing the effort required to individually consider the layout 82A. Specifically, since the arrangement of pages and images is determined according to the structure of the story 68, it becomes easier to naturally determine a layout 82A that follows the flow of the story 68. As a result, the user 12 can achieve a related arrangement of images without having to be aware of the details of the story 68, significantly reducing the burden of editing work.
[0126] Here, we have given an example of how multiple images 64A for the photobook 84 are selected based on words that make up story 68, but this is merely one example. For example, multiple images 64A for the photobook 84 may be selected based on phrases such as noun phrases, verb phrases, adjective phrases, adverbial phrases, and / or prepositional phrases that make up story 68 (for example, phrases such as "beautiful flowers," "go out to play," "very beautiful," "last night," and / or "on the desk"). In this case, a scene label and scene number that can identify the phrase should be used as the scene label 76 and scene number 78. In this way, the affinity between story 68 and the images 64A can be evaluated on a phrase-by-phrase basis that makes up story 68, and then multiple images 64A for the photobook 84 can be selected.
[0127] Alternatively, multiple captured images 64A for the photobook 84 may be selected based on clauses, which are larger semantic units than phrases (for example, a unit containing a subject and predicate, such as "The dog is running"). In this case, a scene label 76 and a scene number 78 that can identify the clause should be used. By doing so, the affinity between the story 68 and the captured images 64A can be evaluated on a clause-by-clause basis to make up the story 68, and then multiple captured images 64A for the photobook 84 can be selected.
[0128] Furthermore, the information processing device 10 generates a story 68 using the caption generation AI 34 and the story generation AI 36 (see Figure 7). Therefore, compared to when the user 12 creates a story 68 while understanding the content of the captured image group 64, the story 68 can be obtained more easily. In addition, a story 68 with stable accuracy can be provided to the user 12 regardless of the user's writing ability.
[0129] Furthermore, the information processing device 10 selects multiple captured images 64A for the photobook 84 from the captured image group 64 based on the feature information 74 (see Figures 8 and 9), which is the inference result of the feature information generation AI 38. This makes it easier to select multiple captured images 64A for the photobook 84 from the captured image group 64 based on knowledge and experience alone. Also, it is possible to select multiple captured images 64A that match the story 68 with higher accuracy compared to when the user 12 selects multiple captured images 64A for the photobook 84 based on knowledge and experience alone.
[0130] Furthermore, the information processing device 10 selects multiple captured images 64A for the photobook 84 from the group of captured images 64 based on the relevant score 80. Therefore, compared to the case where the user 12 selects multiple captured images 64A for the photobook 84 from the group of captured images 64 based solely on knowledge and experience, it is possible to achieve efficient selection of multiple captured images 64A that are highly consistent with the story 68.
[0131] Furthermore, in the information processing device 10, the layout 82A of the image content 85 included in the photobook 84 is obtained on a page-by-page basis based on the divisions of the story 68. This allows for a natural correspondence between the flow of the story 68 and the page structure. As a result, a user-friendly and well-organized layout 82A can be provided to the user 12.
[0132] Furthermore, the information processing device 10 displays the story 68 on screen 16B1 (see Figure 11). On screen 16B1, page numbers P1 and P2 are displayed as an example of visible information that allows identification of the divisions of the story 68 displayed on screen 16B1. Therefore, the user 12 can easily visually grasp the divisions of the story 68, making it easier for the user 12 to confirm and manipulate the structure of the story 68 when editing and viewing it.
[0133] In the above embodiment, the story 68 was described using an example of a written form, but this is merely one example. For example, instead of story 68, a fixed phrase story (in other words, a set phrase) with a written structure may be used. Examples of fixed phrase stories with a written structure include "beginning, middle, end," "introduction, conflict, resolution," "cause and effect," "unexpected development," or "climax and resolution." In this way, even if a fixed phrase story with a written structure is used instead of story 68, the same effects as in the above embodiment can be obtained.
[0134] In the above embodiment, one consistent story 68 was illustrated, but this is merely one example. For example, instead of story 68, multiple stories may be generated by the processor 18 of the information processing device 10. For example, a story generation AI 88 shown in Figure 13 may be used to generate multiple stories.
[0135] As shown in Figure 13, the machine learning performed by the learning execution device 40 to obtain the story generation AI 88 uses the second training data 90. The second training data 90 differs from the second training data 53 (see Figure 3) described in the above embodiment in that it has second ground truth data 92 instead of second ground truth data 56. The second ground truth data 92 differs from the second ground truth data 56 in that it has multiple stories. Similar to the above embodiment, each sentence included in each of the multiple stories is assigned a scene number corresponding to scene number 56B. In the example shown in Figure 13, as an example of multiple stories, an emotional story 92A, a comical story 92B, and a documentary-style story 92C are shown, and each sentence included in each of these stories is assigned a scene number corresponding to scene number 56B.
[0136] Machine learning using the second training data 90, which includes the second correct answer data 92 configured in this way, is performed on the model 58, thereby optimizing the model 58 and obtaining the story generation AI 88.
[0137] In this case, in the photobook generation process shown in Figure 12, the processor 18 uses the story generation AI 88 in step ST16, and the story generation AI 88 generates multiple stories (here, as an example, an emotional story, a comical story, and a documentary-style story). Then, in the photobook generation process shown in Figure 12, the processor 18 executes each of the processes from step ST18 onward for each of the multiple stories.
[0138] As a result, multiple different stories are generated from the image group 64, allowing the user 12 to be provided with a photobook 84 that offers diverse storytelling and variations using the same material.
[0139] It should be noted that while an example of a configuration in which multiple different stories are generated from the captured image group 64 has been given here, this is merely one example. For example, multiple stories may be generated by an existing generation AI such as LLM (not shown) based on instructions given by the user 12 to the information processing device 10 via the reception device 16A (for example, a prompt such as "Generate multiple different stories of a bridal ceremony"). Alternatively, multiple stories may be generated by an existing generation AI such as LLM (not shown) based on the captured image group 64 and instructions given by the user 12 to the information processing device 10 via the reception device 16A (for example, a prompt such as "Generate multiple different stories of a bridal ceremony using the captured image group").
[0140] Alternatively, one of the multiple stories may be selected according to a selection instruction received by the reception device 16A.
[0141] Furthermore, while we have listed emotional stories, comical stories, and documentary-style stories as examples of the multiple stories that can be generated by the story generation AI 88, these are merely examples, and other stories such as the groom's story, the bride's story, and the guests' stories may also be used.
[0142] Furthermore, if multiple stories are generated, these stories may have different content due to branching. In this case, for example, as shown in Figure 14, a story generation AI 94 is used.
[0143] The second training data 96 is used for machine learning to obtain the story generation AI 94. The second training data 96 differs from the second training data 90 shown in Figure 13 in that it includes second ground truth data 98 instead of second ground truth data 92. The second ground truth data 98 contains multiple stories whose content changes depending on the branching.
[0144] In the example shown in Figure 14, the first half story 98A, which is the story of the first half of the bridal ceremony; the second half emotional story 98B, which is the touching story of the second half of the bridal ceremony; and the second half humorous story 98C, which is the humorous story of the second half of the bridal ceremony are exemplified. Here, when machine learning is performed on the model 58 using the second training data 96, which includes the second correct answer data 98, the model 58 is optimized and the story generation AI 94 is obtained.
[0145] When branching generates multiple different stories, an identification number is used as the scene number corresponding to scene numbers 56B, 62C, 69, and 78 described in the above embodiment, which can identify the chronological order of the scenes and the branching of the scenes (for example, the branch "the emotional scene in the second half" and the branch "the humorous scene in the second half"). For example, a sequential number indicates the chronological order of the scenes, and by adding a branch number to the sequential number, the branch number indicates the branching of the scenes.
[0146] For example, if the latter half of the wedding reception branches into a touching scene and a humorous scene, the text and images that make up the story (for example, the correct image 62A and the captured image 64A, etc.) will be assigned the scene number "201" if they refer to the touching scene in the latter half of the wedding reception, and "202" if they refer to the humorous scene.
[0147] Thus, when the story before the branching point is common to multiple stories, and the story after the branching point differs among multiple stories (i.e., when branching results in multiple stories with different content), the correspondence between the scene labels corresponding to scene labels 62B and 76 described in the above embodiment and the scene numbers corresponding to scene numbers 56B, 62C, 69, and 78 described in the above embodiment can be defined as shown in Table 1 below, for example. Note that since scene labels are names that represent the content of a scene, they can be used in common even in different stories if the scenes have the same content.
[0148]
[0149] By associating scene numbers, scene labels, and scene content as shown in Table 1, the photobook generation process shown in Figure 12 can handle cases where multiple stories with different content are generated due to branching (for example, the first half story, the second half emotional story, and the second half humorous story). In other words, even when multiple stories with different content are generated due to branching, the photobook generation process shown in Figure 12 can be handled when the processor 18 refers to the rules shown in Table 1, and the processor 18 sequentially executes the processes from step ST18 onwards for each story, both before and after branching.
[0150] In this way, when multiple stories with different content are generated by the story generation AI 94 through branching, diverse developments can be created while reusing the same materials, making it easier to flexibly change the story according to the user's preferences. For example, simply adding the emotional and humorous developments of the second half to the first half of the story can provide the user 12 with multiple different story experiences. In addition, since scenes from the first half other than the branching parts can be used as common elements, the burden of editing work is reduced. Therefore, it is possible to generate various variations of stories from the same materials, presenting the user 12 with multiple choices and developments, and also to improve the efficiency of material reuse.
[0151] Here, we have given an example where story 68 branches from the first half into an emotional story and a humorous story in the second half. However, this is merely one example, and the emotional story and the humorous story in the second half may merge into a single story (for example, the story of the wedding reception ending). In this case as well, by using scene numbers and scene labels, it is possible to identify whether it is the story before the merger, one of the multiple branched stories, or the story after the merger. Therefore, in the same manner as in the above embodiment, by fitting multiple captured images 64A that are suitable for each story into a suitable layout 82A, an image content 85 containing multiple captured images 64A related to each story, and multiple image content 85 related to each story can be obtained.
[0152] Furthermore, multiple stories with different content obtained through branching may be displayed on the screen 16B1 by the processor 18, and in this state, one of the multiple stories may be selected by the processor 18 according to a selection instruction given by the user 12 to the information processing device 10 via the receiving device 16A. In this case, an image content 85 including multiple captured images 64A related to the selected story should be generated by the processor 18 in the same manner as in the above embodiment. By doing so, only captured images 64A that are highly compatible with the content of the story selected according to the user 12's instruction can be included in the image content 85, so that the user 12 can be provided with an image content 85 that is close to the user 12's intention.
[0153] In the above embodiment, an example was given in which the story 68 generated by the story generation AI 36 is fixed, but this is merely one example, and the story 68 may be modified by the processor 18 according to instructions given from an external source.
[0154] In this case, for example, the photobook generation process shown in Figure 15 is executed by the processor 18. The photobook generation process shown in Figure 15 differs from the photobook generation process shown in Figure 12 in that it includes the processes of step ST100, step ST102, and step ST104 between the processes of step ST16 and step ST18.
[0155] In step ST100 of the photobook generation process shown in Figure 15, the story 68 included in the story-based image group 70 stored in the NVM 20 is displayed on screen 16B1. After the process of step ST100 is executed, the photobook generation process proceeds to step ST102.
[0156] In step ST102, the processor 18 determines whether a story modification instruction (for example, an instruction from the user 12) has been received by the receiving device 16A. If the story modification instruction has not been received by the receiving device 16A in step ST102, the determination is denied and the photobook generation process proceeds to step ST18. If the story modification instruction has been received by the receiving device 16A in step ST102, the determination is affirmed and the photobook generation process proceeds to step ST104.
[0157] In step ST104, the processor 18 modifies story 68 according to the story modification instructions. After the processing in step ST104 is completed, the photobook generation process proceeds to step ST18.
[0158] In this way, by modifying the story 68 according to the story modification instructions received by the reception device 16A, the story 68 can be updated while minimizing the effort required to regenerate the story. Furthermore, by updating the story 68 according to the user 12's intentions, the content of the photobook 84 (i.e., the arrangement, size, and relationship of the multiple captured images 64A within the image content 85, as well as the relationships between the multiple image contents 85, etc.) can be brought closer to what the user 12 intended.
[0159] In the above embodiment, an example of how the captured image group 64 is used by the information processing device 10 was given, but this is merely one example. For example, as shown in Figure 16, the image group 100 may be used by the information processing device 10 instead of the captured image group 64.
[0160] Image group 100 differs from captured image group 64 in that it includes at least one virtual image 101. The virtual image is generated by generation AI 102. Generation AI 102 is obtained by fine-tuning an existing generation AI (e.g., DALL・E 2 or Stable Diffusion) through machine learning using a dataset of ground truth stories, which are stories that serve as correct data, and example images that show scenes included in the ground truth stories.
[0161] The processor 18 inputs the story 68 to the generation AI 102, causing the generation AI 102 to generate a virtual image 101. The virtual image 101 is a virtual image (for example, computer graphics) that represents a scene included in the story 68.
[0162] As a result, even for scenes that have not actually been photographed or are imaginary scenes, a virtual image 101 that fits the story 68 is generated by the generation AI 102, further expanding the range of story expression in the photobook 84.
[0163] In this context, the generated AI 102 is an example of the "generated AI" related to this disclosure, and the virtual image 101 is an example of the "virtual image" related to this disclosure.
[0164] In the above embodiment, an example was given in which a story 68 is generated based on the captured image group 64. However, this is merely one example, and the story 68 may be generated based on instructions from the user 12 (for example, a prompt received by the reception device 16A). Alternatively, the story 68 may be generated based on the captured image group 64 and instructions from the user 12 (for example, a prompt received by the reception device 16A). Furthermore, multiple stories may also be generated based on the captured image group 64 and instructions from the user 12. In addition, at least one story may be generated based on metadata of the captured image group 64 (for example, data in Exif format), which is an example of information about the captured image group 64, and / or instructions from the user 12.
[0165] In the above embodiment, an example was given in which story 68 is generated by a two-stage configuration consisting of a caption generation AI 34 and a story generation AI 36, but this is merely one example. A generation AI that enables the generation of story 68 by fine-tuning a multimodal model integrating image understanding and language generation (e.g., BLIP-2, Flamingo, or PaLI, etc.) with a "dataset that interprets sequence of images as a story" may be used instead of the caption generation AI 34 and story generation AI 36.
[0166] In the above embodiment, an example of how the photobook generation process is performed by the information processing device 10 was described, but this disclosure is not limited thereto, and at least some of the processes included in the photobook generation process may be performed by a device provided outside the information processing device 10. An example of this case will be described below with reference to Figure 17.
[0167] Figure 17 is a conceptual diagram showing an example of the configuration of the information processing system 104. The information processing system 104 is an example of an "information processing device" according to the present disclosure. The information processing system 104 differs from the information processing device 10 described in the above embodiment in that it has an external device 106.
[0168] The external device 106 is connected to the information processing device 10 via the network 30 in a communicative manner. An example of the external device 106 is at least one server that directly or indirectly sends and receives data with the information processing device 10 via the network 30. The external device 106 receives processing execution instructions from the processor 18 of the information processing device 10 via the network 30. The external device 106 then executes the processing according to the received processing execution instructions and transmits the processing results to the information processing device 10 via the network 30. In the information processing device 10, the processor 18 receives the processing results transmitted from the external device 106 via the network 30 and executes processing using the received processing results.
[0169] Examples of processing execution instructions include an instruction to have the external device 106 execute at least a part of the photobook generation process.
[0170] One example of at least a part of the photobook generation process (i.e., the process to be executed by the external device 106) is the processing by the caption generation AI 34 (see Figure 6). In this case, the external device 106 executes the processing by the caption generation AI 34 according to the processing execution instructions given from the processor 18 via the network 30, and transmits the information including the caption 66 as a result of the processing to the information processing device 10 via the network 30. In the information processing device 10, the processor 18 receives the information including the caption 66 and performs the same processing as in the above embodiment using the caption 66.
[0171] A second example of at least a part of the photobook generation process (i.e., the process to be executed by the external device 106) is the processing by the story generation AI 36 (see Figure 7). In this case, the external device 106 executes the processing by the story generation AI 36 according to the processing execution instructions given from the processor 18 via the network 30, and transmits the information processing device 10 with information including the story 68 and a plurality of scene numbers 69 as a result of the processing. In the information processing device 10, the processor 18 receives the information including the story 68 and a plurality of scene numbers 69, and executes the same processing as in the above embodiment using the story 68 and the plurality of scene numbers 69.
[0172] A third example of at least a part of the photobook generation process (i.e., the process to be executed by the external device 106) is the processing by the feature information generation AI 38 (see Figure 8). In this case, the external device 106 executes the processing by the feature information generation AI 38 according to the processing execution instructions given from the processor 18 via the network 30, and transmits the information including the feature information 74 as the processing result to the information processing device 10. In the information processing device 10, the processor 18 receives the information including the feature information 74 and executes the same processing as in the above embodiment using the feature information 74.
[0173] A fourth example of at least a part of the photobook generation process (i.e., the process to be executed by the external device 106) is rule-based processing using rule 82 (see Figure 10). In this case, the external device 106 executes rule-based processing using rule 82 in accordance with the processing execution instructions given from the processor 18 via the network 30, and as a result transmits information to the information processing device 10 including a plurality of captured images 64A selected by the external device 106 and a layout 82A determined by the external device 106. In the information processing device 10, the processor 18 receives the information including the plurality of captured images 64A selected by the external device 106 and the layout 82A determined by the external device 106, and executes the same processing as in the above embodiment using the plurality of captured images 64A selected by the external device 106 and the layout 82A determined by the external device 106.
[0174] The external device 106 may be implemented through cloud computing. Cloud computing is merely one example; the external device 106 may also be implemented through network computing such as fog computing, edge computing, or grid computing.
[0175] In the above embodiment, the captured image 64A is a still image, but the captured image 64A may be a moving image.
[0176] In the above embodiments, each process is executed on any computer. Furthermore, any computer may execute these processes using a processor as hardware, a program as software, or a combination thereof. In this case, the processor is configured to work in cooperation with the program to execute the various processes in the above embodiments, and can function as a unit or means in the above embodiments. Also, the execution order of the processes by the processor is not limited to the order described and may be changed as appropriate. Any computer may be a general-purpose computer, a computer designed for a specific purpose, a workstation, or any other system capable of executing each process.
[0177] A processor may consist of one or more hardware components, and the type of hardware is not limited. For example, a processor may consist of hardware such as a CPU, MPU, FPGA or other programmable logic device, ASIC or other dedicated circuitry for executing specific processes, GPU, and / or NPU. The type of hardware may also be a combination of different types of hardware. When multiple hardware components are configured to execute one or more processes of a processor, the multiple hardware components may reside in physically separate devices or in the same device. Furthermore, in any embodiment, the order of each process performed by the processor is not limited to the order described above and may be changed as appropriate. Hardware is composed of electrical circuits (circuitry) that combine circuit elements such as semiconductor elements.
[0178] Furthermore, the program may be software such as firmware or microcode. Alternatively, the program may be, for example, a group of program modules, each function of which may be implemented by a processor configured to perform its respective function. The program may also be program code and / or multiple code segments stored on one or more non-temporary computer-readable media (e.g., storage). The program may be divided and stored on multiple non-temporary computer-readable media located in physically separate devices. Program code or code segments may represent any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, instructions, data structures, or program statements. Program code or code segments may be connected to other code segments or hardware circuits by sending and receiving information, data, arguments, parameters, or memory contents.
[0179] Furthermore, while the above embodiment illustrates a configuration in which the photobook generation program 32 is pre-stored in the NVM 20 (i.e., installed), this disclosure is not limited thereto. The photobook generation program 32 may be provided in a form stored on a storage medium such as a CD-ROM, DVD-ROM, and / or USB memory. Alternatively, the photobook generation program 32 may be provided in a form that can be downloaded from an external device via a network.
[0180] This disclosure covers all program products. Program products include all forms of products for providing programs. For example, program products include programs provided via networks such as the Internet, and non-temporary computer-readable storage media such as CD-ROMs, DVDs, and USB memory sticks on which programs are stored.
[0181] The photobook generation process described above is merely one example. Therefore, it goes without saying that you may remove unnecessary steps, add new steps, or change the processing order, as long as you do not deviate from the main purpose.
[0182] The descriptions and illustrations presented above are detailed explanations of the parts related to this disclosure and are merely examples of this disclosure. For example, the above explanation of the structure, function, operation, and effect is an example of the structure, function, operation, and effect of the parts related to this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace parts of the descriptions and illustrations presented above, as long as you do not deviate from the spirit of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the parts related to this disclosure, explanations of common technical knowledge, etc., that do not require special explanation to enable the implementation of this disclosure have been omitted from the descriptions and illustrations presented above.
[0183] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
Claims
1. An information processing device comprising a processor, the processor acquires a plurality of images, acquires at least one of information relating to the plurality of images and user instructions, and generates image content in which two or more images that fit the two or more images included in the plurality of images into a layout that fits the two or more images, the story relating to the two or more images and / or the two or more image contents.
2. The information processing apparatus according to claim 1, wherein the processor generates the story based on at least one of the information relating to the plurality of images and the user instructions.
3. The information processing apparatus according to claim 1, wherein the processor selects two or more images from the plurality of images based on the story.
4. The information processing apparatus according to claim 1, wherein the processor determines the layout based on the two or more images.
5. The information processing apparatus according to claim 1, wherein the story represents the relationship between the two or more images and / or the two or more image contents, and the relationship includes a time series of the two or more images and / or the two or more image contents.
6. The information processing apparatus according to claim 1, wherein the processor generates a plurality of stories based on at least one of the information relating to the plurality of images and the user instructions.
7. The information processing apparatus according to claim 6, wherein, when multiple stories are generated, the multiple stories have different content due to branching.
8. The information processing apparatus according to claim 7, wherein the part before the branch is common to the multiple stories, and the part after the branch is different to the multiple stories.
9. The information processing apparatus according to claim 7, wherein one of a plurality of stories is selected according to an externally given selection instruction, and the processor selects two or more images based on elements contained in the selected story.
10. The information processing apparatus according to claim 1, wherein the processor selects two or more images based on elements contained in the story, and the elements are words, phrases, and / or clauses.
11. The information processing apparatus according to claim 1, wherein at least one of the information relating to the plurality of images and the user instruction is input to the first trained model, and the story is generated by the first trained model.
12. The information processing apparatus according to claim 1, wherein information relating to the story and the plurality of images is input to a second trained model, and two or more images are selected from the plurality of images based on the inference results obtained by the second trained model.
13. The information processing apparatus according to claim 1, wherein two or more images are selected from the plurality of images based on the degree of relevance between the story and the plurality of images.
14. The information processing apparatus according to claim 1, wherein the processor modifies the story in accordance with a story modification instruction given from an external source.
15. The information processing apparatus according to claim 1, wherein the layout is obtained based on the story.
16. The information processing apparatus according to claim 1, wherein the layout is obtained on a page-by-page basis based on the division of the story.
17. The information processing apparatus according to claim 16, wherein the story is displayed on a screen, and the screen displays visible information that allows for the identification of the divisions of the story displayed on the screen.
18. The information processing apparatus according to claim 1, wherein the plurality of images include virtual images generated when the story is input to the generating AI.
19. The information processing apparatus according to claim 1, wherein the story is a sentence or a fixed phrase having a sentence-like structure.
20. An image processing method comprising: acquiring multiple images; acquiring at least one of information relating to the multiple images and / or user instructions; and outputting image content in which two or more images that fit a story based on the acquired multiple images and at least one of the acquired user instructions are fitted into a layout that fits the two or more images, wherein the story relates to the two or more images and / or the two or more image contents.
21. A program for causing a computer to perform a process, the process comprising: acquiring a plurality of images; acquiring at least one of information relating to the plurality of images and a user instruction; and outputting image content in which two or more images that fit a story based on the acquired plurality of images and at least one of the acquired user instruction are fitted into a layout that fits the two or more images, the story being a program relating to the two or more images and / or the two or more image contents.