Device, information processing device, information processing system, information processing method, and program
By integrating the character input unit and the operation method acquisition unit in the device, using the first model acquisition operation method and applying it to the image input device, the shortcomings of instruction input and image processing when generating AI of image input in the prior art are solved, and images suitable for generating a model are realized, and user experience and the quality of generated images are improved.
Patent Information
- Application Number
- JP2023183113
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-25
- Publication Date
- 2025-05-12
AI Technical Summary
The prior art lacks an effective way to input instructions and process images when inputting images to generate AI, resulting in the generated images not meeting the user's intentions.
A device is designed that includes a character input unit and an operation method acquisition unit. The character input unit accepts a character string as input, and the operation method obtains the operation method using the first model (a machine learning model based on a character string), and applies the operation method to the image input device to generate an image suitable for the generation model.
Through this device, the user can easily enter instructions and generate images suitable for the generation model, reducing the user's judgment burden on the operation method and improving the quality of the generated image.
Smart Images

Figure 2025072789000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an apparatus, an information processing device, an information processing system, an information processing method, and a program. [Background technology]
[0002] In recent years, generative AI using machine learning models (generative models) has been developed. Multimodal generative AI that generates products in various formats according to various input information has been put to practical use, such as text generation AI that generates text by inputting text, generative AI that generates text by inputting images, AI that generates videos by inputting text, and generative AI that generates images by inputting text and images. The products referred to here include various types of content such as text, audio, images, and videos.
[0003] Generative AI generates products according to inputted commands, but whether the product matches the user's intentions depends largely on the wording of the commands. Commands are often input in natural language format, and the skill of determining what commands should be input to obtain the intended results is a subject of research known as command engineering.
[0004] Conventionally, devices have been considered that require an operator to handwrite instructions and determine the function to be instructed from the instructions. Summary of the Invention [Problem to be solved by the invention]
[0005] Previous studies on generative AI have focused mainly on generative algorithms and methods of utilizing generative AI in data space, and studies from the perspective of linking with hardware devices are still in their infancy. For example, when combining a device that inputs images with a generative AI, there has been insufficient consideration of how to input instructions to the device and how to operate the device to obtain images suitable for input to the generative AI (generative model) in order to input the images input by the device according to the input instructions to the generative AI.
[0006] The present invention has been made in consideration of the above points, and has an object to enable a device to input images suitable for input into a generative model. [Means for solving the problem]
[0007] In order to solve the above problem, the device has a string input unit that accepts input of a string indicating processing content to be applied to an image input by the device, and an operation method acquisition unit that acquires an operation method of the device using a first model that takes the string as input and outputs an operation method of the device when inputting an image to which the processing content indicated by the string is applied, and the first model is trained so as to reduce an error between data generated by a second model based on a certain string and an image obtained by inputting an original input image into the device based on the operation method output by the first model to which the certain string has been input, or an image to which image processing corresponding to the operation method has been applied to the original input image, and data generated by the second model based on the certain string and the certain image. Effect of the Invention
[0008] The device can be made capable of inputting images suitable for input into a generative model. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 illustrates an example of a configuration of an information processing system according to a first embodiment. [Diagram 2] FIG. 1 is a first diagram showing an example of input data, an instruction command, and generated data. [Diagram 3] FIG. 2 is a second diagram showing an example of input data, instruction commands, and generated data. [Figure 4] FIG. 3 is a third diagram showing an example of input data, instruction commands, and generated data. [Diagram 5] 1 is a diagram illustrating an example of a hardware configuration of a device 10 according to a first embodiment. [Figure 6] 2 is a diagram illustrating an example of a hardware configuration of an information processing device 20 according to the first embodiment. [Figure 7] FIG. 2 illustrates an example of a functional configuration of the information processing system during learning according to the first embodiment. [Figure 8] FIG. 2 is a diagram illustrating an example of the configuration of a learning data storage unit 25. [Figure 9] FIG. 13 is a diagram for explaining the learning of the generative model m2. [Figure 10] 13 is a flowchart illustrating an example of a processing procedure for learning a movement method determination model m1 in the first embodiment. [Figure 11] FIG. 2 is a diagram illustrating an example of a functional configuration of an information processing system during inference according to a first embodiment. [Figure 12] 13 is a flowchart illustrating an example of a processing procedure of a data generation process using a trained motion method determination model m1. [Figure 13] 13 is a diagram illustrating an example of a screen transition in a data generation process using a trained operation method determination model m1. FIG. [Figure 14] 13 is a diagram showing an example of the configuration of an operation method modification table 131. FIG. [Figure 15] 13 is a flowchart illustrating an example of a procedure for modifying an operation method. [Figure 16] FIG. 11 illustrates an example of a functional configuration of an information processing system during learning according to a second embodiment. [Figure 17]13 is a flowchart illustrating an example of a processing procedure for learning a movement method determination model m1 in the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Fig. 1 is a diagram showing an example of the configuration of an information processing system in a first embodiment. In Fig. 1, a device 10 that has functions such as scanning, printing, copying, etc. and can input images is connected to an information processing device 20 via a wired or wireless network.
[0011] The information processing device 20 is one or more computers that use a generation AI to generate data corresponding to a character string (hereinafter referred to as an instruction command) that the device 10 accepts from a user along with the input of image data (hereinafter referred to as "input data") input by the device 10. An instruction command is a character string (text data) that indicates, in natural language, the processing content to be applied to the input data. The generation of data using the generation AI may be provided as a cloud service.
[0012] Some specific examples of input data, instruction commands, and data generated by the information processing device 20 (hereinafter referred to as "generated data") will be described.
[0013] Fig. 2 is a first diagram showing an example of input data, an instruction command, and generated data. Fig. 2 shows an example in which an instruction command is input to extract characters included in an image of the input data. In this case, text data included in the input data is generated as generated data.
[0014] Fig. 3 is a second diagram showing an example of input data, instruction commands, and generated data. Fig. 3 shows an example in which an instruction command is input indicating that the person included in the image of the input data is to be maintained as is and the background is to be changed to a beach. In this case, image data in which the background of the person is replaced from mountains to a beach is generated as the generated data.
[0015] FIG. 4 is a third diagram showing an example of input data, instruction commands, and generated data. FIG. 4 shows an example in which an instruction command is input to print hidden characters written in infrared ink on an image of input data. In this example, the contents of the original are also shown because the image contained in the original differs from the image as input data. The scanner of device 10 is configured to be capable of reading with visible light and infrared light, and device 10 generates input data in which hidden characters are visualized by reading infrared light from the original based on the instruction command. Information processing device 20 extracts a character string (text data) from the input data. Device 10 prints the character string.
[0016] In the case of generating data as shown in Figs. 2 to 4, when the device 10 performs a scan (capturing an image), the operation method (setting conditions related to the scan by the device 10) suitable for data generation by the generation AI differs depending on the instruction command. However, the instruction contents of the instruction command are diverse, and it is difficult to determine the operation method on a rule basis. In addition, it is difficult for the user himself to determine the operation method suitable for data generation instructed by the instruction command. Therefore, in this embodiment, the operation method during scanning according to the instruction command is automatically determined using a machine learning model.
[0017] Fig. 5 is a diagram showing an example of the hardware configuration of the device 10 in the first embodiment. In Fig. 5, the device 10 has hardware such as a controller 11, a scanner 12, a printer 13, a modem 14, an operation panel 15, a network interface 16, and an SD card slot 17.
[0018] The controller 11 has a CPU 111, a RAM 112, a ROM 113, a HDD 114, and a NVRAM 115. The ROM 113 stores various programs, data used by the programs, and the like. The RAM 112 is used as a storage area for loading programs, a work area for the loaded programs, and the like. The CPU 111 realizes various functions by processing the programs loaded into the RAM 112. The HDD 114 stores programs, various data used by the programs, and the like. The NVRAM 115 stores various setting information and the like.
[0019] The scanner 12 is hardware (image reading means) for reading (capturing) image data from an original. The printer 13 is hardware (printing means) for printing print data on a print sheet. The modem 14 is hardware for connecting to a telephone line, and is used for transmitting and receiving image data by FAX communication. The operation panel 15 is hardware equipped with an input means such as a button for receiving input from a user, a display means such as a liquid crystal panel, and the like. The liquid crystal panel may have a touch panel function. In this case, the liquid crystal panel also functions as an input means. The network interface 16 is hardware for connecting to a network such as a LAN (whether wired or wireless). The SD card slot 17 is used to read a program stored in the SD card 80. That is, in the device 10, not only the program stored in the ROM 113 but also the program stored in the SD card 80 can be loaded into the RAM 112 and executed. The SD card 80 may be replaced by another recording medium (for example, a CD-ROM or a USB (Universal Serial Bus) memory, etc.). That is, the type of recording medium corresponding to the position of the SD card 80 is not limited to a specific one. In this case, the SD card slot 17 may be replaced by hardware corresponding to the type of recording medium.
[0020] Fig. 6 is a diagram showing an example of a hardware configuration of the information processing device 20 in the first embodiment. The information processing device 20 in Fig. 6 includes a drive device 200, an auxiliary storage device 202, a memory device 203, a processor 204, and an interface device 205, which are all connected to each other via a bus B.
[0021] A program for implementing processing in the information processing device 20 is provided by a recording medium 201 such as a CD-ROM. When the recording medium 201 storing the program is set in the drive device 200, the program is installed from the recording medium 201 to the auxiliary storage device 202 via the drive device 200. However, the program does not necessarily have to be installed from the recording medium 201, but may be downloaded from another computer via a network. The auxiliary storage device 202 stores the installed program as well as necessary files, data, and the like.
[0022] When an instruction to start a program is received, the memory device 203 reads out the program from the auxiliary storage device 202 and stores it. The processor 204 is a CPU or a GPU (Graphics Processing Unit), or a CPU and a GPU, and executes functions related to the information processing device 20 in accordance with the program stored in the memory device 203. The interface device 205 is used as an interface for connecting to a network.
[0023] Fig. 7 is a diagram showing an example of a functional configuration of the information processing system during learning according to the first embodiment. In Fig. 7, the device 10 has a character string input unit 121, an operation method acquisition unit 122, an operation method setting unit 123, an image input unit 124, and a request unit 125. Each of these units is realized by a process executed by the CPU 111 of one or more programs installed in the device 10.
[0024] The information processing device 20 has a movement method determination unit 21, a data generation learning unit 22, a data generation unit 23, a movement method determination learning unit 24, a movement method determination model m1, and a generation model m2. Each of these units is realized by a process executed by the processor 204 of one or more programs installed in the information processing device 20. The information processing device 20 also uses a learning data storage unit 25. The learning data storage unit 25 can be realized by using, for example, the auxiliary storage device 202, or a storage device connectable to the information processing device 20 via a network.
[0025] The character string input unit 121 accepts input of a character string (that is, an instruction command) indicating the processing content to be applied to the image input by the device 10.
[0026] The operation method acquisition unit 122 acquires the operation method of the device 10 using a first model that receives as input a character string (i.e., an instruction command) indicating the processing content to be applied to an image input by the device 10 and outputs an operation method of the device 10 when inputting an image to which the processing content indicated by the character string is applied. Here, the "operation method" refers to information (a set of parameter values) that specifies the reading operation by the scanner 12, such as the reading resolution, number of gradations, reading γ, number of scans, and reading color (RGB color, monochrome, infrared light), when the image is input by reading the image with the scanner 12.
[0027] The first model is a machine learning model (e.g., a neural network, etc.) that is trained to reduce an error between data generated by the second model based on an image obtained by inputting an image of an input source by the device 10 based on a certain character string and an operation method output by the first model that inputs the certain character string, and data generated by the second model based on the certain character string and the certain image. In this embodiment, the operation method determination model m1 is an example of the first model, and the generation model m2 is an example of the second model. Note that the input source image is an image contained in an object to be imaged. When imaging is performed by scanning, when an image is input, the object to be imaged is a manuscript such as a paper document.
[0028] The operation method setting section 123 sets the operation method acquired by the operation method acquisition section 122 in the image input section 124 .
[0029] The image input unit 124 controls the scanner 12 in accordance with a set operation method, and captures an image of an object to be captured (a document set on the scanner 12) to generate input data from an image of an input source.
[0030] The request unit 125 issues a request to the information processing device 20, which is a computer having a second model (generative model m2), to generate data based on the character string (instruction command) accepted by the character string input unit 121 and an image input by the device 10 (the image input unit 124) based on the operation method acquired by the operation method acquisition unit 122. At this time, the request unit 125 transmits the generation request to the computer (information processing device 20) in response to a user instruction input via a screen obtained from the computer (information processing device 20).
[0031] The operation method determination unit 21 receives a character string as input and determines the operation method using a first model (operation method determination model m1) that outputs an operation method of the device 10 that inputs an image to which the processing content indicated by the character string is applied.
[0032] The movement method determination learning unit 24 learns the movement method determination model m1 so as to reduce an error between the generation data generated by the generation model m2 based on an instruction command related to a data generation request from the request unit 125 and input data related to the generation request, and the generation data generated by the generation model m2 based on the instruction command and an input source image of the input data. Learning a model means updating (optimizing) learnable parameters used by the model.
[0033] The learning data storage unit 25 stores learning data used for learning the generative model m1.
[0034] Fig. 8 is a diagram showing an example of the configuration of the learning data storage unit 25. As shown in Fig. 8, the learning data storage unit 25 stores in advance a plurality of learning data, each set being an instruction command and an input image, and a generated image corresponding to the instruction command and the input image. Although Fig. 8 shows three pieces of learning data as an example, it is preferable to actually prepare a huge number of pieces of learning data.
[0035] The learning data stored in the learning data storage unit 25 is also used for learning the movement method determination model m1.
[0036] The data generating unit 23 generates data (generated data) in response to a data generation request from the request unit 125, using the generative model m2.
[0037] The data generating and learning unit 22 performs learning of the generation model m2. The generation model m2 is a machine learning model that receives an image and an instruction command as input and outputs generation data according to the input.
[0038] 9 is a diagram for explaining the learning of the generative model m2. The data generating and learning unit 22 inputs an instruction command and an input image to the generative model m2 for each piece of learning data, and calculates the parameters (conversion coefficients) W of the generative model m2 so that the generated data (image) output from the generative model m2 approaches the generated image of the learning data. ij This W ij The values of can be calculated by a well-known machine learning method (for example, an extension of the CLIP method). ij The matrix in the figure is a schematic representation of the generative model m2 as a transformation matrix for transforming the instruction command and the input image into the generated image. In reality, the generative model m2 is composed of multiple layers rather than a single matrix.
[0039] The following describes the process steps executed in the information processing system: Fig. 10 is a flowchart illustrating an example of the process steps for learning the movement method determination model m1 in the first embodiment.
[0040] In step S110, the character string input unit 121 accepts an instruction command input by the user via the operation panel 15. The instruction command accepted by the character string input unit 121 is sent to the operation method acquisition unit 122. The instruction command input here (hereinafter referred to as a "target instruction command") is an instruction command included in any of the learning data (hereinafter referred to as "target learning data") stored in the learning data storage unit 25. Therefore, the character string input unit 121 may not accept a target instruction command from the user, but may acquire the target learning data from the learning data storage unit 25 and acquire the target instruction command from the target learning data. In addition, each learning data stored in the learning data storage unit 25 may also be stored in advance in the HDD 114 of the listener 10, etc.
[0041] Next, the operation method acquisition unit 122 acquires an operation method of the scanner 12 corresponding to the object instruction command from the operation method determination unit 21 of the information processing device 20 (S120). Specifically, the operation method acquisition unit 122 transmits the object instruction command to the operation method determination unit 21. The operation method determination unit 21 inputs the object instruction command to the operation method determination model m1 under learning, and determines the operation method output by the operation method determination model m1 as the operation method corresponding to the object instruction command. The operation method determination unit 21 transmits the operation method as a determination result to the operation method acquisition unit 122. The operation method acquired by the operation method acquisition unit 122 is hereinafter referred to as the "target operation method".
[0042] Next, the movement method setting unit 123 sets the target movement method in the image input unit 124 (S130).
[0043] Next, the image input unit 124 controls the scanner 12 according to the target operation method to scan (input) an image from a document set in the scanner 12 and generate input data representing the image (S140). The document is a high-quality image sample on which an input image included in the target learning data (i.e., corresponding to the target instruction command) is printed in advance.
[0044] Next, request unit 125 transmits a data generation request including the object specification command accepted by character string input unit 121 in step S110 and the input data scanned by image input unit 124 in step S140 to data generation unit 23 of information processing device 20 (S150).
[0045] When the data generation unit 23 receives the data generation request, the data generation unit 23 generates generated data using the trained generation model m2 (S160). Specifically, the data generation unit 23 inputs the target instruction command and the input image to the generation model m2 and obtains the generated data output by the generation model m2.
[0046] In data generation (image processing) by the generation model m2, a so-called diffusion model is not used, so there is no influence of noise components when using a diffusion model. Therefore, if the input data is equal to the input image included in the target learning data, an image close to the generated image of the target learning data should be generated based on the relationship between the instruction command, the input image, and the generated image shown in Figure 9. However, if the operation method of the scanner 12 of the device 10 (i.e., the judgment result by the operation method judgment model m1) is not appropriate, the operation method becomes a factor of variation, and a difference occurs between the generated data and the generated image.
[0047] Therefore, the movement method determination learning unit 24 calculates the error between the generated data generated by the data generating unit 23 and the generated image included in the target learning data (S170). The error may be calculated using the mean square error. However, taking into account the shift in the scanning position, the error calculation is performed using an image of 1 / 25 size averaged in 5×5 pixel units for both images. If the error is not small enough (for example, if it is not equal to or smaller than the threshold) (No in S180), the movement method determination learning unit 24 updates the parameters of the movement method determination model m1 by the error backpropagation method based on the error (S190).
[0048] By repeating steps S110 and after for each learning data, the learning of the movement method determination model m1 progresses, and when the error becomes sufficiently small (Yes in S180), a learned movement method determination model m1 can be obtained. The learned movement method determination model m1 is expected to extract the essence of the correspondence between the instruction command, the input image, and the generated image. For example, for an instruction command that requires a high-definition input image, it is expected to output an operation method that executes a high-resolution scan. Therefore, there is a high probability that an appropriate movement method can be determined even for instruction commands and input images other than the learning data.
[0049] Next, the inference process (data generation process using the trained motion method determination model m1) will be described. Fig. 11 is a diagram showing an example of the functional configuration of the information processing system at the inference process of the first embodiment. In Fig. 11, the same parts as those in Fig. 7 are given the same reference numerals, and the description thereof will be omitted as appropriate.
[0050] In FIG. 11, the device 10 further has a movement method correction unit 126, a correction receiving unit 127, and an output unit 128. Each of these units is realized by a process in which one or more programs installed in the device 10 are executed by the CPU 111. The device 10 also has a movement method determination model m1. The movement method determination model m1 is a movement method determination model m1 that has been learned by executing the processing procedure of FIG. 10. The device 10 further uses a movement method correction table 131. The movement method correction table 131 can be realized by using, for example, a storage unit such as the HDD 114 or a storage device connectable to the device 10 via a network.
[0051] When another operation method is stored in the memory unit (operation method modification table 131) in association with the operation method acquired by the operation method acquisition unit 122, the operation method modification unit 126 modifies the operation method of the device 10 to the other operation method.
[0052] The operation method modification table 131 is an example of a storage unit that stores an operation method acquired by the operation method acquisition unit 122 in association with an operation method resulting from a change made by the user to the acquired operation method.
[0053] The modification receiving unit 127 receives modifications (changes) to the operation method acquired by the operation method acquisition unit 122 by a user who has referred to the generated data generated by the information processing device 20. When the operation method is modified (changed), the modification receiving unit 127 registers the operation method before the change and the operation method after the change in the operation method modification table 131 in association with each other.
[0054] The output unit 128 executes an output process of the data generated by the information processing device 20.
[0055] On the other hand, the information processing device 20 does not necessarily have to include the movement method determination unit 21, the data generating learning unit 22, the movement method determination learning unit 24, and the movement method determination model m1.
[0056] Fig. 12 is a flowchart for explaining an example of a processing procedure of a data generation process using the trained operation method determination model m1. Also, Fig. 13 is a diagram for explaining an example of a screen transition in the data generation process using the trained operation method determination model m1.
[0057] In Fig. 12, the same steps as in Fig. 10 are given the same step numbers, and their explanations are omitted as appropriate. In Fig. 12, steps S120 and S160 are replaced with S120a and S160a. Also, step S125 is added between steps S120a and S130. Furthermore, steps S170 and after are replaced with steps S210 and after.
[0058] 10, for example, screen 510 in FIG. 13 is displayed on operation panel 15. Screen 510 is a screen for allowing the user to select a function (application) to be used. When the user selects command instruction button 511 on screen 510, character string input unit 121 displays screen 520 on operation panel 15, for example.
[0059] In step S110, character string input unit 121 accepts input of an instruction command from the user via screen 520. For example, the user inputs an instruction command using soft keyboard 521. The input instruction command is displayed in input command display area 522. Character string input unit 121 acquires the instruction command displayed in input command display area 522. Note that, among the steps in FIG. 12, the screens displayed in the steps having the same step numbers as those in FIG. 10 are also displayed when the processing procedure in FIG. 10 is executed.
[0060] In the present embodiment, an example in which an instruction command is input via screen 520 has been described, but an instruction command may be input by other methods. For example, character string input unit 121 may display a list of a good command collection prepared in advance and accept a selection of an instruction command from the list. Character string input unit 121 may also accept input of an instruction command by voice input from a microphone attached to operation panel 15. Character string input unit 121 may also input an instruction command embedded in a one-dimensional or two-dimensional code by photographing the code with a camera attached to operation panel 15. Character string input unit 121 may also receive an instruction command via a network or short-distance wireless communication.
[0061] In step S120a, the operation method acquisition unit 122 acquires an operation method (hereinafter referred to as a "target operation method") of the scanner 12 using the trained operation method determination model m1 to obtain an instruction command (S120). Specifically, the operation method acquisition unit 122 inputs the target instruction command to the trained operation method determination model m1, and acquires the operation method output by the operation method determination model m1 as an operation method corresponding to the target instruction command. Note that, similarly to the learning time, the device 10 may not have the operation method determination model m1, and the information processing device 20 may have the operation method determination unit 21 and the operation method determination model m1. In this case, in step S120a, a process similar to that of step S120 in FIG. 10 may be executed.
[0062] Next, the movement method correction unit 126 corrects the target movement method as necessary (S125). Correction means changing part or all of the movement method (parameter values). Therefore, the structure of the movement method is the same before and after correction. "As necessary" refers to a case where the condition that a correction method for the target movement method is registered (stored) in the movement method correction table 131 is satisfied. The registration of a correction method for a certain movement method is performed by the user in a step described later. Therefore, the correction of the target movement method will be described in detail later. When correction is performed, the corrected movement method is the target movement method in steps S130 and after.
[0063] Steps S130 to S150 are the same as those in FIG. 10, but screen transitions and the like that are omitted in FIG. 10 will be described below.
[0064] In step S140, when the image input unit 124 scans an image from the document set in the scanner 12, the image input unit 124 displays, for example, a screen 530 in FIG. 13 on the operation panel 15. That is, depending on the instruction command, an operation method for performing multiple scans may be set. For example, when a highly accurate photo image and a sharp character image are required for data generation, performing two scans with different gradations or resolutions may be set as the operation method. In such a case, it is preferable to issue an alert so that the user does not remove the document when the first scan is completed. A message is displayed on the screen 530 to notify the user that multiple scans are being performed.
[0065] In step S150, the request unit 125 displays a screen 540 (FIG. 13) and accepts the selection of a generation AI to be used by the user from among multiple generation AIs with which the device 10 can cooperate. Each generation AI uses a different generation model m1 trained using different learning data, for example. Therefore, even with the same input, if the generation AI is different, the generation data may be different. An information processing device 20 may be provided for each generation AI, or one information processing device 20 may have multiple generation AIs (generation models m1).
[0066] When the user selects a generated AI, the request unit 125 launches a browser window on the operation panel 15 and displays a screen including the WEB screen w1 of the generated AI, such as screen 550 (screen 550). More specifically, the request unit 125 transmits an object designation command and input data to the information processing device 20 corresponding to the selected generated AI. The information processing device 20 transmits display data (data in HTML format, etc.) for displaying the WEB screen w1 to the device 10. The request unit 125 displays the WEB screen w1 based on the display data.
[0067] The WEB screen w1 includes a text area 551, an input image area 552, a generation result display area 553, a RUN button 554, and the like. The text area 551 displays an object designation command. The input image area 552 displays the input data generated by the image input unit 124 in step S140. The generation result display area 553 is blank at the time of step S150.
[0068] There are many advantages to this method of displaying the web screen w1 prepared by the provider of the generation AI on the operation panel 15. First, because generation AI technology evolves quickly, websites may be updated frequently. Such updates can be easily accommodated. Also, depending on the generation AI, the amount of points that can be used in a certain period of time, as shown on screen 550, may be consumed each time data is generated. If the web screen w1 is displayed as is, such information can be confirmed on the operation panel 15 of the device 10.
[0069] When the RUN button 554 is pressed on the screen 550, the request unit 125 transmits a data generation request including the object designation command and the input data to the data generation unit 23 of the information processing device 20 (S150).
[0070] The data generation unit 23 generates generation data by inputting the target instruction command and input data into the generation model m2, and transmits to the request unit 125 a WEB screen w1 that has been synthesized with the generation data and is displayed in the generation result display area 553 of the WEB screen w1 that was displayed on the screen 550 (S160a).
[0071] When the request unit 125 receives the WEB screen w1, it redisplays the WEB screen w1 on the operation panel 15 (S210). At this time, the data generation unit 23 may reflect the consumption of points by the user on the WEB screen w1. As the WEB screen w1 is redisplayed, the screen displayed on the operation panel 15 transitions from screen 550 to screen 560.
[0072] Screen 560 includes three options, buttons 561 to 563. The user refers to the generated data and selects one of these three options.
[0073] If the generated data is not in the desired state, the user selects button 561 to instruct modification of the operation method. In this case (Yes in S220), the modification receiving unit 127 pops up a sub-window including screen 570 on screen 560 and receives a modification of the target operation method from the user (S230). That is, on screen 570, the values of parameters constituting the operation method, such as the number of gradations and resolution, can be changed. In the initial state of screen 570, the values of each parameter are those in the target operation method. Note that, in screen 570, in preparation for multiple scans, the operation method can be modified by switching tabs for each scan.
[0074] Following step S230, steps S140 and after are executed again with the modified movement method as the target movement method. Therefore, an image is scanned based on the modified movement method, and generated data is generated based on the input data related to the image.
[0075] On the other hand, when the desired generated data is obtained, the user selects button 562 on screen 560. In this case (No in S220 and Yes in S240), output unit 128 pops up a subwindow including screen 580 on screen 560 and accepts a selection of an output method for the generated data from the user (S250). Screen 580 shows an example in which the options are printing the generated data, saving the generated data in the cloud, and sending the generated data to oneself by email, but other output methods may be selectable. When any option is selected, output unit 128 outputs the generated data by a method corresponding to the option (S260).
[0076] Following step S260, the movement method correction unit 126 determines whether or not the movement method has been corrected (S230) (S270). If the movement method has been corrected (Yes in S270), the movement method correction unit 126 associates the movement method acquired by the movement method acquisition unit 122 with the corrected movement method and stores them in the movement method correction table 131 (S280). If multiple corrections have been made, the last corrected movement method is stored.
[0077] Fig. 14 is a diagram showing an example of the configuration of the operation method modification table 131. As shown in Fig. 14, each record (modification method) of the operation method modification table 131 has a data structure of two columns. The first column stores the operation method before modification (the operation method acquired by the operation method acquisition unit 122), and the second column stores the operation method after modification. If a record is already stored in the operation method modification table 131, a new record is added in step S280. Therefore, a new modification method is accumulated in the operation method modification table 131 every time step S280 is executed.
[0078] It should be noted that when button 563 is selected on screen 560 (No in S220 and No in S240), the generated data is not output and the process ends.
[0079] Next, a detailed description will be given of step S125. Fig. 15 is a flowchart for explaining an example of a processing procedure for correcting the operation method.
[0080] In step S301, the movement method correction unit 126 judges whether or not there is a record in the movement method correction table 131 that includes in the first column a movement method that matches the movement method (target movement method) acquired by the movement method acquisition unit 122. If there is a corresponding record (Yes in S301), the movement method correction unit 126 corrects (changes) the target movement method to the contents of the second column of the record (S302). If there is no corresponding record (No in S302), the movement method correction unit 126 does not correct (change) the target movement method.
[0081] If there are multiple applicable records, any one of the records may be selected based on a predetermined method.
[0082] In this way, the operation method is modified on a rule-based basis using the operation method modification table 131, rather than a trained machine learning model. There are several reasons why this method is desirable. First, the trained model is a general-purpose setting prepared with the average user in mind, whereas the operation method modified by each user for each device 10 is considered to be an individual setting due to the input image handled by the user of the device 10 and the sophistication of the generated data required for the business. When modifying a trained model on the information processing device 20, it is natural to provide the modified trained model to other users, but it is not desirable to provide individual settings to other users. It is preferable that the trained model can be obtained from the information processing device 20 as a general-purpose model, and that the individual model is handled in the local environment of the individual user.
[0083] In addition, since the correction of a trained model involves new learning, there is a disadvantage that the computational load is simply high. Furthermore, by making the processing performed on the device side rule-based, the processing load can be reduced. From the viewpoint of overall optimization of the processing load, it is preferable to execute the processing involving machine learning on the information processing device 20 side and the rule-based processing on the device side.
[0084] As described above, according to the first embodiment, the operation method for the device 10 to input the input data to the generation model m2 that generates data according to the instruction command is acquired using the operation method determination model m1, which is a machine learning model. Therefore, it is possible for the device 10 to input an image suitable for input to the generation model m2. As a result, it is possible to reduce the need for the user to determine the operation method, and the load on the user can be reduced.
[0085] Next, a second embodiment will be described. In the second embodiment, differences from the first embodiment will be described. Therefore, the points not specifically mentioned may be the same as the first embodiment.
[0086] Fig. 16 is a diagram showing an example of a functional configuration of an information processing system during learning according to the second embodiment. In Fig. 16, the same components as those in Fig. 7 are given the same reference numerals, and description thereof will be omitted.
[0087] 16, the device 10 has an input data prediction unit 129 instead of the image input unit 124. The input data prediction unit 129 is realized by a process that one or more programs installed in the device 10 cause the CPU 111 to execute.
[0088] The input data prediction unit 129 predicts (generates) an image to which image processing corresponding to the operation method acquired by the operation method acquisition unit 122 has been applied to the image of the input source. That is, the input data prediction unit 129 rewrites the image of the input source in a software manner to generate pseudo-input data corresponding to the operation method. For example, if the operating conditions regarding the resolution are 600 dpi and a certain image has a data amount equivalent to 1200 dpi, the input data prediction unit 129 averages the image by 4 pixels of 2×2 to generate an image of a quarter size, and uses this as a predicted image of the input data from the scanner 12. Similarly, if the operating method regarding the number of gradations is 4 gradations and a certain image has 256 gradations, the input data prediction unit 129 converts the number of gradations of the image to 4 gradations and uses this as a predicted image of the input data from the scanner 12.
[0089] Fig. 17 is a flowchart for explaining an example of a processing procedure of a learning process of the movement method determination model m1 in the second embodiment. In Fig. 17, the same steps as in Fig. 10 are given the same step numbers and their explanations are omitted. In Fig. 17, steps S130 and S140 in Fig. 10 are replaced with S130a and S140a.
[0090] In step S130a, the movement method setting section 123 sets the target movement method (the movement method acquired by the movement method acquisition section 122) in the input data prediction section 129.
[0091] Next, the input data prediction unit 129 generates a predicted image for the input image included in the target learning data (FIG. 8) based on the set target motion manner using the above-mentioned method (S140a).
[0092] Thereafter, the predicted image is processed as input data from the scanner 12 based on the target operation method.
[0093] Incidentally, the inference process may be the same as in the first embodiment.
[0094] As described above, according to the second embodiment, it is possible to reduce the time and effort required to print the input image of the learning data for learning the movement method determination model m1. Also, in the second embodiment, since all processing can be executed by software, if the prediction accuracy of the input data prediction unit 129 is high, the workload for learning the movement method determination model m1 can be reduced more than in the first embodiment.
[0095] The above-described embodiments may be applied to devices other than image forming devices as long as the devices are capable of inputting images. For example, various cameras (conference cameras, surveillance cameras, photo booths, cameras for ID photos, etc.) may be used as the device 10. In addition, the above-described embodiments have been described with respect to data generation using still images, but the above-described embodiments may also be applied to video.
[0096] Each function of the above-described embodiment can be realized by one or more processing circuits. In this specification, the term "processing circuit" includes a processor programmed to execute each function by software, such as a processor implemented by an electronic circuit, and a device such as an ASIC (Application Specific Integrated Circuit), a DSP (digital signal processor), an FPGA (field programmable gate array), or a conventional circuit module designed to execute each function described above.
[0097] Although the embodiment of the present invention has been described in detail above, the present invention is not limited to such specific embodiment, and various modifications and variations are possible within the scope of the gist of the present invention described in the claims.
[0098] For example, aspects of the present invention are as follows. <1> An apparatus comprising: a character string input unit that receives an input of a character string indicating a processing content to be applied to an image input by the device; an operation method acquisition unit that acquires an operation method of the device using a first model that receives the character string as an input and outputs an operation method of the device when an image to which the processing content indicated by the character string is to be applied is input; having the first model is trained so as to reduce an error between data generated by a second model based on a character string and the certain image, the data being generated by the second model based on an image obtained by inputting an image of an input source into the device based on the operation method output by the first model to which the character string is input, or an image obtained by applying image processing to the input source image in accordance with the operation method, and the data being generated by the second model based on the certain character string and the certain image; The device characterized by: <2> the input source image is an image of an object to be imaged, The device inputs the image of the input source by capturing an image of the object. The device described in <1. <3> a request unit that issues a request for generating data based on the character string received by the character string input unit and an image input by the device based on the operation method acquired by the operation method acquisition unit to a computer having the second model; characterized in that <1> or <2> The equipment listed. <4> The request unit transmits the generation request to the computer in response to a user instruction input via a screen obtained from the computer. Characterized by <3> The equipment listed. <5> an operation method correcting unit that corrects the operation method of the device to the other operation method when another operation method is stored in a storage unit in association with the operation method acquired by the operation method acquisition unit; having Characterized by <1> ~ <4> Any of the devices listed. <6> the storage unit stores the operation method acquired by the operation method acquisition unit in association with an operation method obtained by changing the operation method by a user; Characterized by <5> The equipment listed. <7> an operation method determination unit that uses a first model that receives a character string as an input and outputs an operation method of a device to which an image is input and to which a processing content indicated by the character string is to be applied, to determine the operation method; having the first model is trained so as to reduce an error between data generated by a second model based on a character string and an image obtained by inputting an image of an input source into the device based on the operation method output by the first model to which the character string is input, or an image obtained by applying image processing corresponding to the operation method to the input source image, and data generated by the second model based on the character string and the image; 23. An information processing apparatus comprising: <8> An information processing system including a device, a character string input unit that receives an input of a character string indicating a processing content to be applied to an image input by the device; an operation method acquisition unit that acquires an operation method of the device using a first model that receives the character string as an input and outputs an operation method of the device when an image to which a processing content indicated by the character string is to be applied is input; having the first model is trained so as to reduce an error between data generated by a second model based on a character string and the operation method output by the first model to which the character string is input, the image obtained by inputting an image of an input source to the device, or an image obtained by applying image processing corresponding to the operation method to the input source image, and data generated by the second model based on the character string and the image; An information processing system comprising: <9> The equipment, a character string input step of receiving an input of a character string indicating a processing content to be applied to an image input by the device; an operation method acquisition step of acquiring an operation method of the device using a first model in which the character string is input and an operation method of the device when an image to which the processing content indicated by the character string is applied is input is output; Run the first model is trained so as to reduce an error between data generated by a second model based on a character string and the certain image, the data being generated by the second model based on an image obtained by inputting an image of an input source into the device based on the operation method output by the first model to which the character string is input, or an image obtained by applying image processing to the input source image in accordance with the operation method, and the data being generated by the second model based on the certain character string and the certain image; 23. An information processing method comprising: <10> For equipment, a character string input step of receiving an input of a character string indicating a processing content to be applied to an image input by the device; an operation method acquisition step of acquiring an operation method of the device using a first model in which the character string is input and an operation method of the device when an image to which the processing content indicated by the character string is applied is input is output; Run the command, the first model is trained so as to reduce an error between data generated by a second model based on a character string and the certain image, the data being generated by the second model based on an image obtained by inputting an image of an input source into the device based on the operation method output by the first model to which the character string is input, or an image obtained by applying image processing to the input source image in accordance with the operation method, and the data being generated by the second model based on the certain character string and the certain image; A program characterized by: [Explanation of symbols]
[0099] 10 equipment 11 Controller 12 Scanner 13 Printers 14 Modem 15 Operation Panel 16 Network Interfaces 17 SD card slot 20 Information processing device 21 Operation method determination section 22 Data Generation and Learning Unit 23 Data Generation Division 24. Operation method determination learning unit 25 Learning data storage unit 80 SD card 111 CPU 112 RAM 113 ROM 114 HDD 115 NVRAM 121 String input section 122 Operation method acquisition section 123 Operation method setting section 124 Image input unit 125 Request part 126 Operation method modification section 127 Correction Reception Department 128 Output section 129 Input Data Prediction Unit 131 Operation Method Modification Table 200 Drive device 201 Recording media 202 Auxiliary storage device 203 Memory Device 204 Processor 205 Interface Device B Bus m1 Operation method judgment model m2 Generative Model [Prior art documents] [Patent documents]
[0100] [Patent Document 1] Japanese Patent Application Publication No. 4-355871
Claims
1. An apparatus comprising: a character string input unit that receives an input of a character string indicating a processing content to be applied to an image input by the device; an operation method acquisition unit that acquires an operation method of the device by using a first model that receives the character string as an input and outputs an operation method of the device when an image to which a processing content indicated by the character string is to be applied is input; having the first model is trained so as to reduce an error between data generated by a second model based on a character string and the operation method output by the first model to which the character string is input, the image being obtained by inputting an image of an input source to the device, or an image obtained by applying image processing corresponding to the operation method to the input source image, and data generated by the second model based on the character string and the image; An apparatus characterized by:
2. the input source image is an image of an object to be imaged, The device inputs the image of the input source by capturing an image of the object.
2. The device according to claim 1 .
3. a request unit that issues a request for generating data based on the character string received by the character string input unit and an image input by the device based on the operation method acquired by the operation method acquisition unit to a computer having the second model; 2. The device of claim 1, further comprising:
4. The request unit transmits the generation request to the computer in response to a user instruction input via a screen obtained from the computer.
4. The device according to claim 3.
5. an operation method correcting unit that corrects the operation method of the device to the other operation method when another operation method is stored in a storage unit in association with the operation method acquired by the operation method acquisition unit; having 2. The device according to claim 1 .
6. the storage unit stores the operation method acquired by the operation method acquisition unit in association with an operation method obtained by changing the operation method by a user; 6. The device according to claim 5.
7. an operation method determination unit that uses a first model that receives a character string as an input and outputs an operation method of a device to which an image is input and to which a processing content indicated by the character string is to be applied, to determine the operation method; having the first model is trained so as to reduce an error between data generated by a second model based on a character string and an image obtained by inputting an image of an input source into the device based on the operation method output by the first model to which the character string is input, or an image obtained by applying image processing corresponding to the operation method to the input source image, and data generated by the second model based on the character string and the image; 23. An information processing apparatus comprising:
8. An information processing system including a device, a character string input unit that receives an input of a character string indicating a processing content to be applied to an image input by the device; an operation method acquisition unit that acquires an operation method of the device using a first model that receives the character string as an input and outputs an operation method of the device when an image to which a processing content indicated by the character string is to be applied is input; having the first model is trained so as to reduce an error between data generated by a second model based on a character string and the operation method output by the first model to which the character string is input, the image being obtained by inputting an image of an input source to the device, or an image obtained by applying image processing corresponding to the operation method to the input source image, and data generated by the second model based on the character string and the image; An information processing system comprising:
9. The equipment, a character string input step of receiving an input of a character string indicating a processing content to be applied to an image input by the device; an operation method acquisition step of acquiring an operation method of the device using a first model in which the character string is input and an operation method of the device when an image to which the processing content indicated by the character string is applied is input is output; Run the first model is trained so as to reduce an error between data generated by a second model based on a character string and the operation method output by the first model to which the character string is input, the image being obtained by inputting an image of an input source to the device, or an image obtained by applying image processing corresponding to the operation method to the input source image, and data generated by the second model based on the character string and the image; 23. An information processing method comprising:
10. For equipment, a character string input step of receiving an input of a character string indicating a processing content to be applied to an image input by the device; an operation method acquisition step of acquiring an operation method of the device using a first model in which the character string is input and an operation method of the device when an image to which the processing content indicated by the character string is applied is input is output; Run the command, the first model is trained so as to reduce an error between data generated by a second model based on a character string and the operation method output by the first model to which the character string is input, the image being obtained by inputting an image of an input source to the device, or an image obtained by applying image processing corresponding to the operation method to the input source image, and data generated by the second model based on the character string and the image; A program characterized by:
Citation Information
Patent Citations
Device for inputting character and its supporting device
JP1992355871A