Method for controlling information processing apparatus, information processing apparatus, and program
The control method for an information processing apparatus addresses the challenge of reducing user labor in instructing generative AI models by generating secondary prompts based on initial prompts, enhancing the efficiency of processing multiple images.
Patent Information
- Application Number
- JP2023207114
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-19
AI Technical Summary
Existing technologies face challenges in reducing user labor when instructing prompts for generative AI models to perform processing or editing on multiple images according to user intentions.
A control method for an information processing apparatus that includes a holding step for a first prompt related to processing a first image and a generation step for generating a second prompt based on the first prompt, allowing for efficient processing of multiple images.
This method significantly reduces user labor in instructing prompts for generative AI models, enabling more efficient processing and editing of multiple images according to user intentions.
Smart Images

Figure 2025091701000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for controlling an information processing apparatus, the information processing apparatus, and a program.
Background Art
[0002] Generally, when editing layout data with a layout such as a poster or a leaflet, first, a template material closer to the completed image is selected. Then, for the selected template, contents such as text and images are added or deleted, and the positions and sizes of the respective contents are changed to adjust the overall layout. In recent years, in the editing of the above-described layout data, it has become possible to use so-called generative AI technologies when adding or editing contents. Examples of such generative AI technologies include Stable Diffusion (registered trademark), ChatGPT (registered trademark), and GAN (adversarial generative algorithm).
[0003] When editing layout data, there may be a case where the image content included in the template is edited while maintaining the layout of the template. For example, in a case where a poster for an event guide in a hot season is changed to a poster for an event guide in a cold season, if the portrait in the poster is wearing light clothing, it is required to change it to thick clothing suitable for the cold season. In such a situation, when there are a plurality of portrait images in the poster, it may be required to change the clothing of all the people. In a situation where editing according to a common "context" is required for a plurality of image contents in this way, it may be troublesome for the user to give instructions for prompts that are inputs to the generative AI for each content. Against such a background, various technologies for reducing the labor involved in instructing a prompt that serves as input for generative AI have been under consideration. Patent Document 1 discloses a technology for extracting profile information of a subject in image content and information on what kind of image it is from the image content, and automatically creating a prompt for generating the image content based on the information.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, although the technology disclosed in Patent Document 1 can generate an image similar to the original image, it is difficult to perform processing or editing on an existing image in accordance with the user's intention. Therefore, there is a demand for the realization of a technology that can further reduce the labor involved in instructing a prompt that serves as input to a model constructed based on machine learning, such as generative AI, for performing processing or editing on each of a plurality of target images in accordance with the user's intention.
[0006] In view of the above problems, an object of the present invention is to reduce the labor of the user involved in instructing a prompt for performing processing on a plurality of images using a model constructed based on machine learning in a more suitable manner.
Means for Solving the Problems
[0007] The control method of the information processing apparatus according to the present invention includes a holding step of holding a first prompt including an instruction related to the processing of the first image, which is received as an input when the first model, which has been learned to perform processing on an input image based on the input prompt, outputs a processed image by performing processing on the first image, and a generation step of generating a second prompt including at least an instruction related to the processing of a specified second image based on the first prompt.
Advantages of the Invention
[0008] According to the present invention, it becomes possible to reduce the user's labor related to the instruction of the prompt for performing processing on a plurality of images in a more suitable manner by using a model constructed based on machine learning.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Mode for Carrying Out the Invention
[0010] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the present specification and drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant description is omitted.
[0011] <Overview of Generative AI Technology> First, in order to make the features of the information processing system according to an embodiment of the present disclosure easier to understand, an overview of generative AI technology will be described below. Generative AI technology is a technology that uses generative models such as Stable Diffusion, ChatGPT, and Generative Adversarial Network (GAN). For example, FIG. 7 is a diagram showing an overview of generative AI technology. In generative AI technology, an input image and an input prompt 700 such as an image or text are used, and a generative model 701 that outputs products 702 such as characters, images, and videos with a high probability of matching the "context" represented by the input image and input prompt 700 is utilized. The relationship between the input value and the "context" is acquired when the generative model 701 learns using a large amount of images and texts. Also, the generative model 701 can change the generated product 702 by mainly changing the initial value generated from random numbers during the generation of the product 702.
[0012] <First Embodiment> An example of an information processing system according to the first embodiment of the present disclosure will be described below. FIG. 1 is a diagram showing an example of the system configuration of the information processing system according to the present embodiment, and shows an example when configured as a printing system involving layout data editing for an image output device. In the printing system shown in FIG. 1, editing of layout data and transmission of a print job to the image output device are executed by an externally connected client (for example, a terminal device such as a PC). The image output device corresponds to a device that generates and outputs a printed matter by forming an image on a recording medium such as paper. Also, when creating a print job, for example, an editing operation of print settings may be performed on the screen of the above client. Hereinafter, the system configuration of the printing system illustrated in FIG. 1 will be described in more detail.
[0013] As shown in FIG. 1, the client 102 is connected to each of the server 104, the image output device 100, and the image output device 101 via the network 103. The client 102 is used, for example, for editing layout data such as posters and leaflets, and requests the server 104 to perform some editing and data processing and rendering processing on the layout data. The server 104 schematically shows a server that performs various processes such as editing, data processing, and rendering on various data such as layout data. Further, the client 102 generates a print job by applying print settings to the layout data after the editing operation, and transmits the print job to a desired image output device (for example, a specified image output device among the image output devices 100 and 101). In the example shown in FIG. 1, the case where there are two image output devices accessible by the client 102 via the network 103 is shown, but the number of image output devices is not limited, and it may be one or three or more. Similarly, for the client and the server, although each is one in the example shown in FIG. 1, there may be two or more of them.
[0014] As an example of the process related to printing, an example of the case where a print job is transmitted to the image output device 100 from the printing application installed in the client 102 via the printer driver will be described. It is assumed that the printing application and the printer driver have been previously installed in the client 102. The printing application acquires device information of the associated image output device 100 from the printer driver, and printing parameters such as paper type, paper size, and print quality, and edits the print settings within the range of the acquired set of parameters. The printing application generates a print job based on the above print settings and the layout data image subjected to rendering processing by the server 104, and transmits the print job to the image output device 100 via the spool of the printer driver. The image output device 100 executes printing based on the print settings of the print job received from the printing application. In addition, the image output device 100 holds configuration information regarding the ink and paper it handles, and status information such as the idle state and print errors as device information. Further, when printing cannot be executed normally due to factors such as an abnormal state such as insufficient paper remaining or out of ink, or an error in the print settings, the image output device 100 may display a warning message on the main body panel to present to the user the reason why printing cannot be executed normally.
[0015] Referring to FIG. 2, an example of the hardware configuration of the image output device 100 will be described. Note that since the image output device 101 can be realized with a configuration substantially the same as that of the image output device 100, a detailed description thereof will be omitted. The operation of the image output device 100 is controlled by a CPU (Central Processing Unit) 200. The CPU 200 operates based on a control program stored in a program area in a ROM (Read Only Memory) 201, a control program stored in a storage area such as an external memory 208, and the like. The CPU 200 outputs an image signal as output information to a printing unit (printer engine) 207 connected to a printing unit I / F 205 via a system bus 203. The CPU 200 may notify the client 102 of the information within the image output device 100 via the communication by communicating with the client 102 via the input unit 204. Further, the CPU 200 can also receive output data to be output to the printing unit 207 via the input unit 204. The RAM (Random Access Memory) 202 functions as the main memory of the CPU 200 and as a temporary storage area such as a work area. The RAM 202 may be configured such that its memory capacity can be expanded by an optional RAM connected to an expansion port (not shown). Note that the RAM 202 is used for an output information expansion area, an environmental data storage area, a non-volatile memory, etc. The external memory 208 is implemented by a hard disk drive (HDD) or an IC card, etc., and access thereto is controlled by a memory controller 206. The external memory 208 can be connected as an option, and can store font data, an emulation program, form data, information on used ink, information on paper feed paper types and sizes, and main body status information, etc. Also, the operation unit 209 includes a panel realized by a display device such as a liquid crystal display, and is configured to be able to display various information on the panel.
[0016] With reference to FIG. 3, an example of the hardware configuration of an information processing apparatus applicable as the client 102 shown in FIG. 1 will be described. Note that since the server 104 can be realized with a configuration substantially the same as that of the client 102, a detailed description thereof will be omitted. Inside the main body 307 of the computer, a CPU 300, a ROM 301, a RAM 302, a keyboard controller 304, a display controller 305, and a disk controller 306 are included. The CPU 300 reads various programs such as a control program, a system program, and an application program from the external memory 311 via the disk controller 306 into the RAM 302. The CPU 300 performs various data processes and display control of the display 310, etc. by executing the various programs read into the RAM 302. The CPU 300 may read a control program, etc. from the ROM 301. Also, the CPU 300 may be realized as a dedicated circuit such as an ASIC. The CPU 300 and the dedicated circuit correspond to an example of a hardware circuit or a hardware processor. The disk controller 306 controls access to external memories 311 such as HDDs, CD-ROMs, DVD-ROMs, and USBs. The RAM 302 may be configured so that its capacity can be expanded by an optional RAM (not shown) or the like, and is mainly used as a work area for the CPU 300. The keyboard controller 304 controls key inputs from input devices such as the keyboard 308 and the pointing device 309. The display controller 305 controls the display on the display 310. In each embodiment of the present disclosure, unless otherwise specified, it is assumed that the CPU 300 controls each unit connected to the main bus 303 via the main bus 303. Of course, in the server 104, components that are not necessarily essential, such as the display 310, may not be included in the configuration.
[0017] Referring to FIG. 4, an example of the functional configuration of the information processing system according to the present embodiment will be described, particularly focusing on the configurations of the image output device 100, the client 102, and the server 104 described with reference to FIGS. 1 to 3. Note that since the image output device 101 is substantially the same as the image output device 100, a detailed description thereof will be omitted.
[0018] First, an example of the functional configuration of the client 102 and the server 104 will be described. The layout data editing unit 401 adds and deletes contents such as characters and images to be placed on printed materials such as posters and flyers, and adjusts the layout of each content. When performing processing such as color tone correction, clipping, and filling for each content, a request for such processing is made from the layout data editing unit 401 to the data content editing unit 411 of the server 104. The layout data is stored as a cache in the layout data DB 400 of the client 102, or stored in the layout data DB 410 of the server 104 for each client 102 (or for each account if there is a user account).
[0019] The prompt setting unit 402 acquires prompt information corresponding to an instruction related to image processing. The generated image proposal unit 403 requests the image processing unit 412 of the server 104 to perform image processing. The image processing unit 412 of the server 104 requests the image generation unit 416 to generate an image in response to an instruction from the generated image proposal unit 403. At this time, the prompt information acquired by the prompt acquisition unit 413 and the image acquired by the image acquisition unit 414 are used for the request. The image generation unit 416 generates an image using the generation model 415. The generation model 415 corresponds to a learned model that has been learned by machine learning to perform processing on an input image based on the input prompt, and corresponds to an example of the "first model". The generated image is transmitted to the client 102 by the image processing unit 412. Then, the generated image proposal unit 403 of the client 102 proposes the above image transmitted from the image processing unit 412 to the user.
[0020] The print job transmission unit 405 creates a print job and transmits the created print job to the image output device 100. When creating the print job, the print job transmission unit 405 requests the preview image generation unit 417 or the print image generation unit 418 of the server 104 to perform preview of layout data or generation processing of a print image.
[0021] Next, an example of the functional configuration of the image output device 100 will be described. The ROM 201 holds a device information holding unit 421, a print job receiving unit 422, and a print execution unit 423. The print job receiving unit 422 receives a print job transmitted from the client 102. The print execution unit 423 executes print processing based on the received print job. The device information holding unit 421 holds information such as the type and remaining amount of ink installed in the image output device 100, information such as the type and size of registration paper and paper for paper feeding, the main body status information of the image output device 100, and the status information of print jobs. When the image output device 100 to be used is determined in advance, the information held in the device information holding unit 421 may be used to create layout data suitable for the image output device 100. In this case, the target information is acquired from the device information holding unit 421 and may be held on the client 102 or server 104 side in association with the layout data DB 400 or 410.
[0022] Referring to FIG. 5, an example of the layout data stored in the layout data DBs 400 and 410 will be described. The data table shown in FIG. 5 is held for each layout data, for example. The data table shown in FIG. 5 is provided with areas such as an ID 500, content 501, content type 502, layout coordinates 503, setting information 504, and metadata 505, and information regarding the content can be set in each area for each content.
[0023] In the area of the ID 500, identification information (ID) for uniquely identifying the content on the layout data is set. In the area of the content 501, the value of each content such as characters and images laid out on the layout data is set. In the area of the content type 502, information indicating the type of the content set in the area of the content 501 is set. In the area of the layout coordinates 503, a value indicating the position of the content on the layout data is set. In the area of the setting information 504, an attribute value indicating characteristics of each content such as the color and size of the content is set. In addition, settings regarding the entire layout data such as manuscript size and data for variable printing may be retained. In this case, for example, in the area of content 501, "entire" may be retained as a value, in the area of content type 502, the setting type may be retained, and in the area of setting information 504, the set value may be retained. Of course, the above is merely an example. For example, it may be operated in separate files for each type of parameter, or parameters of other types than those exemplified above may be included in the layout data.
[0024] In the area of metadata 505, metadata for each content may be retained. In the present embodiment, for example, as shown by sample 506, the ID of the related image may be set in the area of metadata 505. In addition, when the content type 502 is an image, an explanation of the image may be set in the area of metadata 505. As a specific example, in the example shown as sample 507, an explanation "Explanation: Image of cherry blossom viewing" is set in the area of metadata 505 for the target image. Also, in the example shown as sample 508, an explanation "Type: Person" is set in the area of metadata 505 for the target image. In addition, when the content type 502 is text, information such as the type and purpose of the text to be described may be set in the area of metadata 505. As a specific example, in the example shown as sample 509, information indicating the purpose of the text is set in the area of metadata 505 for the target text.
[0025] In the area of prompt information 510, when the content is edited using a generative AI (for example, generative model 415) that edits the content according to the input prompt, the prompt information input to the generative AI may be retained. Details of the content retained as prompt information will be described separately later with reference to FIG. 11.
[0026] Referring to FIG. 6, an example of a layout data editing screen 600 for the image output device 100, which is displayed on the display 310 of the client 102, will be described. The layout data editing screen 600 includes a template list 601 as an area where candidates for layout templates are displayed as a list. The user can browse and select the template that is closest to the completed form of the layout data via the template list 601. The template selected via the template list 601 is displayed in the layout editing area 604. Information on the templates displayed in the template list 601 may be obtained from the layout data DB 400 or 410 as layout data, or may be obtained from an external network service such as a cloud service or an SNS service.
[0027] The layout data editing unit 401 or the data content editing unit 411 performs editing such as position adjustment, color tone correction, clipping, filling, etc. on each content displayed on the layout editing area 604. The method of specifying the target content is not particularly limited. For example, selection may be made on the content on the layout editing area 604 by a mouse operation, a tap operation, etc., or the content to be added may be specified when adding content.
[0028] When content is edited via the layout data editing screen 600, for example, a processing menu is displayed when the target content is specified, and the user's specification of the editing content is received via the processing menu. FIG. 9 is a diagram showing an example of the processing menu 902 displayed on the layout data editing screen 600. In the example shown in FIG. 9, for the processing menu 902, a color tone correction menu 903 and a processing menu 906 by a generation AI are displayed as candidates for menus for specifying the processing content.
[0029] In the color correction menu 903, for example, there are provided an area 904 for receiving the specification of the content of color correction from the processed image image by preset specification, and an area 905 for receiving the specification related to detailed color correction such as contrast and tone curve. With such a configuration, the user can specify the content of the editing related to the color correction to be applied to the content displayed in the layout editing area 604 via the area 904 and the area 905.
[0030] In the processing menu 906 by the generative AI, there is provided a prompt input area 908 for receiving the specification of the prompt information that is the input to the generative AI. The prompt lock button 907 is a button for switching the availability of changing the prompt information input in the prompt input area 908, and by transitioning to the locked state where changes are not allowed, it is possible to prevent the situation where the prompt information is inadvertently changed. The generation button 909 is a button for receiving an instruction related to the generation of an image based on the prompt information input in the prompt input area 908. When the generation button 909 is pressed, an image is generated by the generated image proposal unit 403 based on the prompt information input in the prompt input area 908, and a preview of the image is displayed in the generated image preview area 910. The confirmation button 911 is a button for receiving an instruction from the user related to the confirmation of the generated image displayed in the generated image preview area 910. When the confirmation button 911 is pressed in a state where a preview of the generated image is displayed in the generated image preview area 910, the content (image) displayed in the layout editing area 604 is replaced with the generated image. Also, at this time, the state of the prompt lock button 907 may be changed to the locked state. By applying such control, when the generated image is confirmed, it is possible to keep the prompt information from being inadvertently changed. Note that the images presented via the generated image preview area 910 by the generated image proposal unit 403 are not limited to one, and a plurality of generated images may be presented. When a plurality of generated images are presented via the generated image preview area 910, the generated images may be finalized by pressing the confirmation button 911 while any one of the plurality of generated images is selected and accepted.
[0031] Note that in the above, as an example of the processing menu 902, an example of the case where menus corresponding to the color tone correction function and the processing function by the generation AI are respectively displayed has been described. However, these are merely examples and do not limit the available processing functions. As a specific example, processing functions such as size adjustment, cropping, filling, and image filters may be available, and the menus corresponding to the processing functions may be displayed in the processing menu 902. Also, in the example shown in FIG. 9, an example of the processing menu when the type of the target content is an image has been shown. However, the processing menu displayed according to the type of content may be selectively switched.
[0032] Here, the example shown in FIG. 6 will be described again. The image addition button 602 is a button for receiving an instruction from the user to add image content to the layout editing area 604. As a specific example, when the image addition button 602 is pressed, a dialog for specifying the file path of the target image content is displayed, and when the file path is specified via the dialog, the import process of the file may be executed. Also, the text addition button 603 is a button for receiving an instruction from the user to add text content. As a specific example, when the text addition button 603 is pressed, a process related to inserting new text content may be executed. Note that the above is merely an example and does not limit the type of target content. Buttons for receiving instructions regarding the addition of the content may be displayed according to the type of target content. Also, an interface for specifying the import source of the content may be provided according to the type of target content, and external cloud service storage, SNS services, etc. may be configured to be specifiable as the import source. Further, an instruction regarding the addition of the content may be received by a drag-and-drop operation of the content on the layout editing area 604.
[0033] The print execution button 605 is a button for receiving an instruction regarding the execution of printing from the user. When the print execution button 605 is pressed, the print job transmission unit 405 creates a print job for the layout data with an image displayed in the layout editing area 604, and transmits the print job to a target image output device (for example, the image output device 100 or 101). The save button 606 is a button for receiving an instruction regarding the saving of the layout data being edited from the user. When the save button 606 is pressed, the layout data with an image displayed in the layout editing area 604 is saved in a predetermined storage area (for example, the layout data DB 400 or 410).
[0034] Referring to FIGS. 8 and 10, an example of the processing of the information processing system according to the present embodiment will be described, particularly focusing on the processing of the client 102 and the server 104 when image processing is performed using a generation AI by an operation via the layout data editing screen 600. FIG. 8 is a flowchart showing an example of the processing of each of the client 102 and the server 104. Also, FIG. 10 is a diagram showing an example of the layout data editing screen 600. Note that the series of processing shown in FIG. 8 is realized, for example, when the CPU 300 in each of the client 102 and the server 104 reads out a program stored in the ROM 301 and executes it in the RAM 302.
[0035] The series of processes shown in FIG. 8 is started by the generated image proposal unit 403 at the timing when the generation button 909 on the layout data editing screen 600 is pressed, for example, on the client 102. In S2001, the generated image proposal unit 403 acquires the image content specified in the layout editing area 604. In S2002, the prompt setting unit 402 acquires the prompt information input in the prompt input area 908. In S2003, the generated image proposal unit 403 requests the image processing unit 412 of the server 104 to generate a new image using the prompt information acquired in S2002. In S2004, the image processing unit 412 generates a new image by the generation-based AI technology using the generation model 415 through the image generation unit 416. Then, the image processing unit 412 transmits the generated image to the generated image proposal unit 403 of the client 102. In S2005, the generated image proposal unit 403 receives the generated image transmitted from the server 104 in S2004 and displays the generated image in the generated image preview area 910 of the layout data editing screen 600.
[0036] In S2006, the generated image proposal unit 403 receives an instruction from the user regarding the confirmation of the generated image via the confirmation button 911 on the layout data editing screen 600. In S2007, the generated image proposal unit 403 determines whether an instruction from the user regarding the confirmation of the generated image has been received (that is, whether the confirmation button 911 has been pressed). If the generated image proposal unit 403 determines in S2007 that an instruction from the user regarding the confirmation of the generated image has not been received, the series of processes shown in FIG. 8 is terminated. On the other hand, if the generated image proposal unit 403 determines in S2007 that an instruction from the user regarding the confirmation of the generated image has been received, the process proceeds to S2008. In S2008, the generated image proposal unit 403 updates the data stored in at least one of the layout data DBs 400 and 410 based on the generated image received from the server 104 in S2005. Further, as exemplified by the content 1001 in FIG. 10, the generated image proposal unit 403 replaces the image content on the layout editing area 604 with the generated image. After that, the generated image proposal unit 403 ends the series of processes shown in FIG. 8.
[0037] As described above, by applying the control described with reference to FIGS. 8 and 10, it becomes possible to receive a prompt instruction from the user and cause the generation AI to process and generate image content using the prompt as an input.
[0038] With reference to FIG. 11, an example of layout data will be described by focusing on the case where the image content is updated using the generation AI by executing the process shown in FIG. 8 for the layout data DB exemplified in FIG. 5. As shown as Sample 1101, the target content is updated with the content described with reference to FIG. 10, that is, the image generated by the generation AI. Further, as shown as Sample 1102, the content type of the target content is updated to "generated AI image". In addition, as shown as Sample 1103, the prompt input to the prompt input area 908 and the state of the prompt lock button 907 are added as prompt information.
[0039] Referring to FIGS. 12 and 13, an example of the processing of the information processing system according to the present embodiment will be described, particularly focusing on the processing when prompt information applied to some contents is diverted to the processing of other contents by an instruction via the layout data editing screen 600. FIG. 12 is a flowchart showing an example of the processing of the client 102. FIG. 13 is a diagram showing an example of the layout data editing screen 600. The series of processes shown in FIG. 12 is realized, for example, when the CPU 300 reads out the program stored in the ROM 301 to the RAM 302 and executes it in the client 102.
[0040] The series of processes shown in FIG. 12 is started by the prompt diversion unit 404, for example, at the timing when an instruction regarding the diversion of the prompt is received via the layout data editing screen 600 in the client 102. FIG. 13 shows an example of the state of the layout data editing screen 600 when an instruction regarding the diversion of prompt information is received. When the designation of the image content 1301 generated by the generation AI is received on the layout editing area 604, the processing menu 902 is displayed. At this time, the prompt input area 908 contains the prompt information applied when generating the image content 1301. In this state, when the action button 1302 is pressed, the action menu 1303 for the prompt is displayed. In the example shown in FIG. 13, the action menu 1303 displays menus of "Divert to Other Contents" and "Divert from Other Contents". The menu presented as "Divert to Other Contents" corresponds to a menu for diverting the displayed prompt for the processing of other image contents. The menu presented as "Divert from Other Contents" corresponds to a menu for diverting the prompt applied during the processing from the image content processed by the generation AI. In the present embodiment, an example will be described when the menu of "Divert to Other Contents" is selected when receiving an instruction regarding the diversion of the prompt in the example shown in FIG. 13. Note that the above is merely an example and does not limit the method of receiving instructions regarding the diversion of prompts. For example, instead of the action menu 1303, a UI such as a tooltip may be used to receive instructions regarding the diversion of prompts.
[0041] In S3001, the prompt diversion unit 404 receives the designation of the destination content on the layout editing area 604. At this time, the prompt diversion unit 404 refers to at least one of the layout data DBs 400 and 410, and after pre-determining whether each content is appropriate as the destination content, the prompt diversion unit 404 may narrow down the content that can be designated on the layout editing area 604.
[0042] For example, the prompt diversion unit 404 may narrow down the content that can be designated to content whose content attribute is an image, image content whose prompt state is not in a locked state in the prompt information, generated AI image content, etc. As another example, the prompt diversion unit 404 may narrow down the content that can be designated according to the type of the source content stored in the metadata and the type of the destination content. As a specific example, when an instruction for clothing such as "wear clothes" is given as a prompt, it is considered appropriate to apply it to images of types such as people and animals. On the other hand, it is not necessarily considered appropriate to apply an instruction for clothing as exemplified above to images of types of figures such as circles, squares, and stars. In view of such a situation, for example, a determination as to whether they are of the same type, a determination as to whether they are synonyms even if they are not of the same type, or a determination as to whether they belong to the same category may be made. Note that the determination as to whether they are synonyms or whether they belong to the same category can be realized, for example, by using a pre-defined dictionary or the distance when a language is vectorized. In addition, the prompt reuse unit 404 may narrow down the specifiable content in view of the content of the prompt itself. As a specific example, the prompt reuse unit 404 may determine whether it is appropriate to apply the prompt by obtaining the probability of occurrence in the same context after vectorizing the prompt and the type of the content to be reused respectively. In addition, when receiving the designation of the content to be reused, the prompt reuse unit 404 may determine whether the designated content is appropriate as the content to be reused, and display a warning according to the determination result. As described above, by determining whether it is appropriate as the content to be reused, it is possible to prevent the occurrence of a situation where, when the prompt is reused, appropriate processing is not performed by the generative AI (for example, the intended processing is not performed).
[0043] In S3002, the prompt reuse unit 404 acquires the original prompt (for example, the prompt applied to the original content), and generates a prompt to be applied to the content to be reused (hereinafter, also referred to as the prompt for the content to be reused) based on the prompt. At this time, the prompt reuse unit 404 may generate the prompt for the content to be reused by, for example, copying the original prompt. As another example, the prompt reuse unit 404 may generate the prompt for the content to be reused by adjusting the original prompt so that appropriate processing is performed on the content to be reused. In this embodiment, for the sake of convenience, it is assumed that the prompt reuse unit 404 generates the prompt for the content to be reused by copying the original prompt. In addition, the original prompt corresponds to an example of the first prompt, and the prompt for the content to be reused generated from the original prompt corresponds to an example of the second prompt. In addition, the image content to be processed by the original prompt (the first prompt) corresponds to an example of the first image, and the image content to be processed by the prompt for the content to be reused (the second prompt) corresponds to an example of the second image.
[0044] In S3003, the prompt diversion unit 404 controls such that the processing menu 902 of the diversion destination content is displayed with the prompt generated in S3002 input to the prompt input area 908. Then, the prompt diversion unit 404 ends the series of processes shown in FIG. 12.
[0045] Referring to FIG. 14, an example of the state of the layout data editing screen 600 will be described with attention paid to the case where the result of prompt diversion is displayed by executing the process shown in FIG. 12. In the example shown in FIG. 14, the image content 1301 is the content from which the prompt is diverted, and the image content 1304 is specified as the content of the prompt diversion destination on the layout editing area 604. In this case, the processing menu 902 is displayed with the generated prompt input to the prompt input area 908. When the generation button 909 is pressed, the series of processes described with reference to FIG. 8 is executed, and the image content 1304 specified as the diversion destination content is subjected to image processing by the generation AI. By applying the above control, for example, the user can easily generate image contents with the same taste, such as changing to thick clothing in the same way, by using the prompt applied to some contents for other contents as well. Also, by displaying the prompt lock button 907 in an unlocked state, additional prompt editing by the user may be allowed. By applying such control, even when the image content generated by directly diverting the prompt is not what the user intends, the user can easily attempt to regenerate the image content after modifying the prompt.
[0046] In addition, in the present embodiment, an example in which one image content is specified as the content of the destination of the prompt diversion has been described. However, a plurality of image contents may be specified. In that case, in at least one of the layout data DBs 400 and 410, the information of the original prompt may be written to the prompt information 510 of each specified image content. Also, without finishing the process up to the display of the result of the prompt diversion, the process related to the image processing using the diverted prompt (for example, the process described with reference to FIG. 8) may be executed as a series of processes. By applying such control, the user can immediately confirm the processing result for the image at the destination via the generated image preview area 910. As described above, according to the information processing system according to the present embodiment, the prompt used for the processing of the content by the generation AI can be diverted to process other content. Therefore, even in a situation where there are a plurality of contents to be edited, the user can cause the generation AI to perform processing on each content without complicated operations such as individually inputting prompts for each content.
[0047] <Second Embodiment> An example of the information processing system according to the second embodiment of the present disclosure will be described below. In the first embodiment, an example in which the prompt for image processing is diverted by being copied in the process of S3002 shown in FIG. 12 has been described. In the present embodiment, an example of the configuration and process for generating a prompt will be described so that appropriate processing (for example, processing reflecting the user's intention) is performed on the content at the destination. In the present embodiment, the description will be made focusing on the parts different from the first embodiment, and detailed description of the parts substantially the same as the first embodiment will be omitted.
[0048] FIG. 15 is a functional block diagram showing an example of the functional configuration of the information processing system according to the present embodiment. As can be seen by comparing with the configuration shown in FIG. 4, the information processing system according to the present embodiment is different from the first embodiment in that a large language model 1501 and a prompt generation unit 1502 are added to the server 104. Details of the large language model 1501 and the prompt generation unit 1502 will be described separately later in conjunction with the description of the processing of the information processing system with reference to FIG. 16.
[0049] Referring to FIG. 16, an example of the processing of the information processing system according to the present embodiment will be described, focusing particularly on the processing of the client 102 and the server 104. Since the processing of S3001 and S3003 is substantially the same as the processing of S3001 and S3003 shown in FIG. 12, detailed description thereof will be omitted.
[0050] In S4000, the prompt diversion unit 404 requests the prompt generation unit 1502 of the server 104 to generate a prompt. In S4001, the prompt generation unit 1502 acquires the source image processing prompt through the prompt acquisition unit 413, and generates an image processing prompt adapted to the destination content using the large language model 1501. The large language model 1501 corresponds to a learned model that has been trained to generate a prompt based on the input instruction by machine learning, and corresponds to an example of the "second model". Then, the prompt generation unit 1502 transmits the generated image processing prompt to the prompt diversion unit 404 of the client 102.
[0051] Here, with reference to FIG. 17, a specific example of prompt generation using a large language model will be described. When a text generation prompt 1700 representing an instruction regarding the text to be generated is given to a large language model (LLM) 1701, the large language model 1701 outputs a text product 1702 with a high probability of conforming to the text generation prompt 1700. The relationship between the input value and the "context" is acquired when the large language model 1701 learns using a large number of sentences. By giving the text generation prompt 1700 appropriate instructions along with the text, various processes can be performed on the text, such as summarizing the text, expanding the content, or translating it.
[0052] In this embodiment, a large language model is used to generate an image processing prompt. For example, when the source for diversion is a human image and the image processing prompt for the source for diversion is "wear gloves and boots", but the target image for diversion is a dog image, it is difficult to say that the phrase "gloves" in the image processing prompt is appropriate. Therefore, the prompt generation unit 1502 may create a text generation prompt for the large language model 1501 and then request the generation of an image processing prompt suitable for the dog image. The text generation prompt in this case may have the following content, for example.
[0053] Please change the application target of the following image processing prompt to be a dog image. “wear gloves and boots”
[0054] The large language model 1501 receives the input of the above text generation prompt and generates an image processing prompt suitable for the dog image, such as "wear boots" or "wear boots on each legs".
[0055] As described above, according to the information processing system according to this embodiment, not only can the prompt for image processing be simply diverted, but it is also possible to generate a processing prompt adapted to the content of the diversion destination.
[0056] <Third Embodiment> An example of the information processing system according to the third embodiment of the present disclosure will be described below. In the first embodiment, an example was described in which the designation of the image content that is the source of the prompt is first received, and then the designation of the content that is the destination of the prompt is received. In this embodiment, an example will be described in which the designation of the image content that is the destination of the prompt is first received, and then the designation of the content that is the source of the prompt is received. In this embodiment, in the example shown in FIG. 13, it is assumed that the image content 1304 is designated on the layout editing area 604, and an instruction related to the diversion of the prompt is received by selecting the "Divert from Other Content" menu in the action menu 1303.
[0057] Referring to FIG. 18, an example of the processing of the information processing system according to this embodiment will be described, focusing particularly on the processing of the client 102. Note that since the processing of S3002 and S3003 is substantially the same as the processing of S3002 and S3003 shown in FIG. 12, a detailed description thereof will be omitted. The series of processes shown in FIG. 18 is started by the prompt diversion unit 404, for example, at the timing when an instruction related to the diversion of the prompt is received via the layout data editing screen 600 on the client 102. In S5001, the prompt diversion unit 404 accepts the designation of the source content on the layout editing area 604. At this time, the prompt diversion unit 404 refers to at least one of the layout data DBs 400 and 410, and after pre-determining whether each content is appropriate as the destination content, it may narrow down the content that can be designated on the layout editing area 604. As a specific example, it is difficult to divert a prompt from content that does not hold prompt information. Therefore, the prompt diversion unit 404 may limit the content that can be designated to content in which prompt information has been input, such as sample 1103 in the example shown in FIG. 11. Further, the prompt diversion unit 404 may narrow down the source content that can be designated according to the types of the source and destination contents of the diversion, the relationship between the type of the destination content and the source prompt, and the like. By applying the above control, it becomes possible to prevent a situation where, when a prompt is diverted, appropriate processing is not performed by the generative AI (for example, the intended processing is not performed).
[0058] Also, when the prompt diversion unit 404 accepts the designation of the content that is the source of the prompt, it may present to the user, as prompt information, the prompt designated for the content and the images before and after processing based on the prompt. FIG. 19 is a diagram showing an example of the layout data according to the present embodiment. When image content is processed by the generative AI, as shown as sample 1901, images before and after processing by the generative AI with the target prompt as the input may be held in association with the prompt information.
[0059] FIG. 20 is a diagram showing an example of the state of the layout data editing screen 600, and shows an example of a case where corresponding prompt information is presented when diverting a prompt. For example, in S5001 of FIG. 15, the prompt diversion unit 404 may display a prompt information pop-up 2001 when an image content 1301 is specified on the layout editing area 604. In the prompt information pop-up 2001, according to the prompt information shown as a sample 1901 in FIG. 19, a prompt 2002 and a processed image 2003 showing the images before and after processing by the target prompt are displayed. By applying such control, the user can determine whether to divert the target prompt after checking the displayed prompt information. Note that a divert button 2005 and a cancel button 2004 are used to determine whether to divert the target prompt. Specifically, when the divert button 2005 is pressed, the diversion of the target prompt is performed, and when the cancel button 2004 is pressed, the diversion of the prompt is cancelled.
[0060] As described above, in the present embodiment, an example has been described in which the designation of the image content that will be the destination of the prompt diversion is first received, and then the designation of the content that will be the source of the prompt diversion is received. In addition, by presenting the images before and after processing by the prompt applied to the source content, the user can determine whether to divert the prompt before the prompt is actually diverted and applied.
[0061] <Other Embodiments> The present invention can also be realized by executing the following processing. That is, software (program) that realizes the functions of the above-described embodiments is supplied to a system or device via a network or various storage media, and a computer (or CPU, MPU, etc.) of the system or device reads and executes the program. In the above-described embodiment, a layout data creation application was given as an example of an application. However, the present invention is not limited to this example, and can be realized and is effective in any application having a similar image layout function. In the above-described embodiment, a personal computer was assumed as the information processing apparatus. However, the present invention is not limited to this example, and can be realized and is effective for any information processing apparatus (terminal) such as a mobile phone, a personal digital assistant, a digital still camera, a digital video camera, a portable music player, a game, a set-top box, an Internet appliance, etc., for which a similar usage method is possible. In the above-described embodiment, Ethernet was used as an example of the network configuration. However, the present invention is not limited to this example, and other arbitrary network configurations such as a wireless LAN, IEEE1394, Bluetooth, etc. may be used. As described above, the preferred embodiments of the present invention have been described in detail. However, the present invention is not limited to such specific embodiments, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.
[0062] Further, the disclosure of the present embodiment includes the following methods, configurations, and programs. (Method 1) A control method for an information processing apparatus, comprising: a holding step of holding a first prompt including an instruction related to the processing of a first image, which is received as an input when a first model, which has been learned to perform processing on an input image based on an input prompt, outputs a processed image by performing processing on the first image; and a generating step of generating a second prompt including at least an instruction related to the processing of a specified second image based on the first prompt. A control method for an information processing apparatus, characterized by including these steps. (Method 2) The control method for an information processing apparatus according to Method 1, further including a receiving step of receiving an instruction related to the reuse of the first prompt, wherein the generating step generates the second prompt based on the first prompt when the receiving step receives an instruction related to the reuse of the first prompt. (Method 3) In the reception step, after receiving the designation of the first image, an instruction related to the appropriation of the first prompt and the designation of the second image are received. The control method of the information processing apparatus according to Method 2, characterized in that. (Method 4) In the reception step, after receiving the designation of the second image, an instruction related to the appropriation of the first prompt and the designation of the first image are received. The control method of the information processing apparatus according to Method 2, characterized in that. (Method 5) In the generation step, based on at least any one of the metadata of the first image and the metadata of the second image, and the first prompt, the second prompt is generated. The control method of the information processing apparatus according to any one of Methods 1 to 4, characterized in that. (Method 6) In the generation step, by inputting the first prompt and an instruction related to the processing of the first prompt to a second model that generates a prompt based on the input instruction, the second model is caused to generate the second prompt. The control method of the information processing apparatus according to any one of Methods 1 to 5, characterized in that. (Method 7) When the designation of the first image is received, it includes a first output control step of controlling so that the first prompt is presented to the user in association with the first image. The control method of the information processing apparatus according to any one of Methods 1 to 6, characterized in that. (Method 8) The holding step includes holding the first prompt and the processed image in association with each other, and when the designation of the first image is received, controlling so that the first image and the processed image are presented to the user as the images before and after processing based on the first prompt. The control method of the information processing apparatus according to any one of Methods 1 to 7, characterized in that. An information processing apparatus, comprising: a holding unit that holds a first prompt including an instruction related to processing of a first image, which is received as an input when a first model, trained to perform processing on an input image based on an input prompt, outputs a processed image by performing processing on the first image; and a generation unit that generates a second prompt including at least an instruction related to processing of a specified second image based on the first prompt. (Program 1) A program for causing a computer to function as an information processing apparatus, the information processing apparatus comprising: a holding unit that holds a first prompt including an instruction related to processing of a first image, which is received as an input when a first model, trained to perform processing on an input image based on an input prompt, outputs a processed image by performing processing on the first image; and a generation unit that generates a second prompt including at least an instruction related to processing of a specified second image based on the first prompt.
Explanation of Signs
[0063] 102 Client 104 Server 400 Layout Data DB 404 Prompt Diversion Unit 415 Generation Model
Claims
1. A control method for an information processing apparatus, comprising: A holding step of holding a first prompt including an instruction related to processing of the first image, which is received as an input when the first model, which has been learned to perform processing on an input image based on an input prompt, outputs a processed image by performing processing on the first image; A generation step of generating a second prompt including at least an instruction related to processing of a specified second image based on the first prompt; A control method for an information processing apparatus, characterized by including the above.
2. Including a reception step of receiving an instruction related to reuse of the first prompt, The control method for an information processing apparatus according to claim 1, wherein in the generation step, when the reception step receives an instruction related to reuse of the first prompt, the second prompt is generated based on the first prompt.
3. The control method for an information processing apparatus according to claim 2, wherein the reception step receives an instruction related to reuse of the first prompt and a specification of the second image after receiving a specification of the first image.
4. The control method for an information processing apparatus according to claim 2, wherein the reception step receives an instruction related to reuse of the first prompt and a specification of the first image after receiving a specification of the second image.
5. The control method for an information processing apparatus according to claim 1, wherein in the generation step, the second prompt is generated based on at least one of the metadata of the first image and the metadata of the second image and the first prompt.
6. The generation step inputs, to a second model that generates a prompt based on an input instruction, the first prompt and an instruction related to the processing of the first prompt, so as to cause the second model to generate the second prompt. The control method for an information processing apparatus according to claim 1 is characterized by this.
7. The control method for an information processing apparatus according to claim 1 includes a first output control step of controlling to present the first prompt to the user in association with the first image when the designation of the first image is received.
8. The holding step holds the first prompt and the processed image in association with each other. When the designation of the first image is received, the control method for an information processing apparatus according to claim 1 includes a second output control step of controlling to present the first image and the processed image to the user as the images before and after processing based on the first prompt. This is the control method for an information processing apparatus according to claim 1, characterized by this.
9. A holding means for holding a first prompt including an instruction related to the processing of the first image, which was received as an input when a first model learned to process an input image based on an input prompt outputs a processed image by processing the first image; A generation means for generating a second prompt including at least an instruction related to the processing of a designated second image based on the first prompt; An information processing apparatus, characterized by having these.
10. A computer, A holding means for holding a first prompt including an instruction related to the processing of the first image, which was received as an input when a first model learned to process an input image based on an input prompt outputs a processed image by processing the first image; Generating means for generating a second prompt including at least an instruction related to the processing of a specified second image based on the first prompt; A program for causing an information processing apparatus to function as an information processing apparatus characterized by having
Citation Information
Patent Citations
Information processing device, method, program, and system
JP7329293B1
Cited By
Computer program, device, and method for causing print execution unit of printer to print image
WO2026028802A1