Information processing apparatus, program, and control method for information processing apparatus
The multifunction peripheral system addresses user navigation challenges by integrating a display and language model to receive and respond to natural language inputs, facilitating easy access to desired screens.
Patent Information
- Application Number
- JP2024113441
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2026-01-28
AI Technical Summary
Users of multifunction devices face difficulties in navigating complex screen transitions and operations to display desired screens due to small operation panels and lack of direct guidance from language models on UI operations.
A multifunction peripheral system that includes a display, reception, and transmission mechanism to receive natural language inputs, transmit screen information, and utilize a language model to provide operation instructions based on the current screen display, enabling easy access to desired screens.
Enables users to easily display desired screens on multifunction peripherals by providing intuitive operation guidance through natural language interactions.
Smart Images

Figure 2026013177000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a multifunction peripheral, a program for a multifunction peripheral, and a control method thereof. [Background technology]
[0002] Services such as CHATGPT (registered trademark) are known in which a user asks a question in natural language and a language model answers the user's question in natural language. Patent Document 1 discloses a technology in which a document from a specific domain is input into a language model, query data predicted from the document is created, and the query data and a search model are used to answer the user's question. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2023-76413 Summary of the Invention [Problem to be solved by the invention]
[0004] When a user operating software wants to display a certain screen, he or she may not know what operations to perform on the software's user interface (UI) to display that screen.
[0005] In particular, multifunction devices with functions such as printing and scanning may have small operation panels, requiring the user to perform many screen operations to display the desired screen. For example, the desired screen may be displayed by selecting button "a" on screen A to move to screen B, and then selecting button "b" on screen B to move to screen C.
[0006] The user can search for complex screen transitions and operation methods to display the desired screen by inputting a question in natural language into the language model. However, since the language model does not know the screen displayed on the operation panel when the user asks the question, it cannot answer how to operate from that screen to the desired screen.
[0007] The present invention has been made in view of the above-mentioned problems, and has as its object to enable a user to easily display a desired screen on a multifunction peripheral. [Means for solving the problem]
[0008] The multifunction peripheral of the present invention is a multifunction peripheral comprising: a display means for displaying a screen; a reception means for receiving natural language from a user; a transmission means for transmitting information about the screen displayed when the reception means receives the natural language and a prompt using the received natural language; and a reception means for receiving an output of a language model, the output being based on the transmitted screen information and the transmitted prompt; wherein the display means displays the output upon receiving the output, and the output is a method of operating the multifunction peripheral, the method of operating on the screen displayed by the display means when the reception means receives the natural language. [Effects of the Invention]
[0009] According to the present invention, a user can easily display a desired screen on a multifunction peripheral. [Brief explanation of the drawings]
[0010] [Figure 1] A diagram showing an example of the system configuration [Figure 2] FIG. 1 shows an example of the configuration of a printing device 100. [Figure 3] FIG. 2 shows an example of the configuration of an information processing device 200. [Figure 4] A diagram showing an example of the configuration of cloud 300. [Figure 5]FIG. 12 shows an example of screen transitions until an estimated ink level screen 1200 is displayed. [Figure 6] FIG. 22 shows an example of screen transitions until a maintenance screen 2200 is displayed. [Figure 7] FIG. 14 is a diagram showing an example of a document 1400 that describes the procedure for displaying the estimated ink level screen 1200. [Figure 8] FIG. 24 is a diagram showing an example of a document 2400 that describes the procedure for cleaning the ink head. [Figure 9] FIG. 1 is a diagram showing an example of a process for generating a knowledge base. [Figure 10] Conceptual diagram of the process of generating a knowledge base from a document 1400 [Figure 11] Conceptual diagram of the process of generating a knowledge base from document 2400 [Figure 12] A diagram showing an example of a knowledge base [Figure 13] FIG. 10 is a diagram showing an example of a prompt created by the CPU 101. [Figure 14] FIG. 10 is a diagram showing an example of processing executed by an application 150 of the printing device 100 according to the first embodiment. [Figure 15] A diagram showing an example of a message created in response to a question. [Figure 16] FIG. 10 is a diagram showing an example of processing executed by an application 150 of a printing device 100 according to a second embodiment. [Figure 17] A diagram showing an example of a message created in response to a question. [Figure 18] A diagram showing an example of the sequence from when a user asks a question to when the output of the language model 380 is displayed on the panel. DETAILED DESCRIPTION OF THE INVENTION
[0011] Preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the present invention. Although the embodiments describe multiple features, not all of these features are necessarily essential to the invention, and multiple features may be combined in any desired manner.
[0012] First Embodiment 1 is a diagram showing an example of the configuration of a system according to this embodiment. The system is composed of a printing device 100, an information processing device 200, and a cloud 300. The printing device 100 and the information processing device 200 are connected to the Internet 500 via a network 700. The Internet 500 is connected to the cloud 300 via the network 700. As described above, the printing device 100 and the information processing device 200 are connected to the cloud 300 and are able to communicate with each other.
[0013] It should be noted that the present invention may be configured in any combination of the printing device 100, the information processing device 200, and the cloud 300. For example, the information processing device 200 and the cloud 300 may not be included as components, and the printing device 100 may have these functions, and the system may be configured with only the printing device 100.
[0014] The printing device 100 is an image forming device or image processing device such as a multifunction peripheral, printer, copier, or scanner, and is a so-called office machine. The operations performed by the cloud 300 may be performed by the printing device 100 and / or the information processing device 200. In the following embodiment, the printing device 100 is used as an example, but the present invention is not limited to the printing device 100. Other devices such as cameras and home appliances may also be used. In addition, the printing device 100 may be replaced by an information processing device 200 such as a smartphone or a personal computer.
[0015] Alternatively, the system may be one in which the printing device 100 and the information processing device 200 operate in cooperation with each other. For example, the system may be configured such that the information processing device 200, such as a smartphone, receives a voice input from a user and acquires a response from a language model, the information processing device 200 transmits the result to the printing device 100, and the printing device 100 displays the response.
[0016] 2 shows an example of the configuration of the printing device 100. The printing device 100 includes a CPU 101, a ROM 102, a RAM 103, a storage 104, a panel 110, a communication I / F 120, a printing unit 130, a reading unit 131, a question receiving unit 140, a screen information acquisition unit 141, and an application 150. Note that CPU is an abbreviation for Central Processing Unit, ROM is an abbreviation for Read Only Memory, RAM is an abbreviation for Random Access Memory, and I / F is an abbreviation for interface. Note that not all of these components are required.
[0017] The CPU 101 controls the overall operation of the printing device 100. The CPU 101 reads a control program stored in the ROM 102 or storage 104, executes the program, and controls each unit and performs data calculations, thereby realizing the functions of the printing device 100, such as printing and reading.
[0018] ROM 102 stores control programs executable by CPU 101. ROM 102 is controlled by CPU 101, and the programs are executed by CPU 101 reading the control programs from ROM 102. RAM 103 is the main storage memory of CPU 101, and is used as a work area and a temporary storage area for expanding the various control programs stored in ROM 102 and storage 104. Storage 104 is a storage area that stores all kinds of information, such as print data, image data, various programs, and various setting information, in addition to the control programs executable by CPU 101.
[0019] The panel 110 is an operation panel that can display screens and accept inputs from the user. When the printing device 100 is turned on, icons and buttons for using the functions of the printing device 100, such as copy, scan, and print, are displayed on the panel 110. The user can use these functions by operating the panel 110.
[0020] The communication I / F 120 is an interface that connects to the network 700 and transmits and receives data to and from an external system on the network 700, such as the cloud 300.
[0021] The printing unit 130 and the printing unit 206 are for printing images based on image data stored in the RAM 203 or storage 204 onto a printing medium (recording paper) fed from a paper feed cassette (not shown).
[0022] The reading unit 131 reads an image of an original document, and the CPU 101 converts the image into image data such as binary data. Image data generated based on the image read by the reading unit 131 is sent to an external device or printed on recording paper. The reading unit may have a document table and may be configured to read the original document by transporting the original document set on the document table, for example, with a scanner. Alternatively, the reading unit may be configured to read the original document by photographing it with a camera. The reading unit 131 is not an essential component of the printing device 100.
[0023] The question receiving unit 140 receives questions from the user. The method for receiving questions may be voice input, or may be reception from a software key displayed on the panel 110, or a hardware key (not shown) connected to the printing device 100. The question receiving unit 140 may receive not only questions, but also instructions such as "Show me how to display the XX screen." In other words, the question receiving unit 140 receives input of natural language such as questions and instructions from the user.
[0024] The screen information acquisition unit 141 is used to acquire a screenshot of the screen displayed on the panel 110, an identifier of the displayed screen, etc. The application 150 is an application executed by the CPU 101 on the printing device 100, and is software for realizing the present invention.
[0025] 3 shows an example of the configuration of the information processing device 200. The information processing device 200 includes a CPU 201, a ROM 202, a RAM 203, a storage 204, a panel 210, a communication I / F 220, a question receiving unit 240, a screen information acquisition unit 241, and an application 250. Note that not all of these components are essential.
[0026] The CPU 201 controls the overall operation of the information processing device 200. The CPU 201 reads a control program stored in the ROM 202 or the storage 204, and executes the program to control each unit and perform data calculations, thereby realizing the functions of the information processing device 200. The ROM 202 stores a control program that can be executed by the CPU 201.
[0027] The ROM 202 is controlled by the CPU 201, and the CPU 201 executes the control program by reading it from the ROM 202. The RAM 203 is the main storage memory of the CPU 201, and is used as a work area and a temporary storage area for expanding the various control programs stored in the ROM 202 and the storage 204. The storage 204 is a storage area that stores all information such as print data, image data, various programs, and various setting information in addition to the control programs that the CPU 201 can execute.
[0028] The panel 210 is an operation panel that can display a screen and accept input from a user. The UI of the application 250 of the information processing device 200 and the like are displayed on the panel 210.
[0029] The communication I / F 220 is an interface that connects to the network 700 and transmits and receives data to and from an external system on the network 700, such as the cloud 300.
[0030] The question receiving unit 240 receives questions from the user. The question may be received by voice input, or may be received from a software key displayed on the panel 110 or a hardware key connected to the information processing device 200 (not shown).
[0031] The screen information acquisition unit 241 is used to acquire a screenshot of the screen displayed on the panel 210, an identification number of the displayed screen, etc. The application 250 is an application executed by the CPU 201 on the information processing device 200, and is software for realizing the present invention. The application 250 may be software in a form different from an application, such as a driver.
[0032] 4 shows an example of the configuration of the cloud 300. The cloud 300 includes a CPU 301, a ROM 302, a RAM 303, a storage 304, a database 305, a communication I / F 320, an application 350, a language model 380, and an embedded model 381. Note that not all of these components are essential.
[0033] The CPU 301 controls the overall operation of the cloud 300. The CPU 301 reads a control program stored in the ROM 302 or the storage 304, and executes the program to control each unit, calculate data, and so on, thereby realizing the functions of the cloud 300.
[0034] ROM 302 stores control programs executable by CPU 301. ROM 302 is controlled by CPU 301, and the programs are executed by CPU 301 reading the control programs from ROM 302. RAM 303 is the main memory of CPU 301, and is used as a work area and a temporary storage area for expanding the various control programs stored in ROM 302 and storage 304. Storage 304 is a storage area that stores all kinds of information, such as print data, image data, various programs, and various setting information, in addition to the control programs executable by CPU 301.
[0035] The database 305 is a database configured to operate in a cloud environment. The database 305 can store structured data, unstructured data, and semi-structured data. The communication I / F 320 is an interface that connects to the network 700 and transmits and receives data to and from external systems on the network 700, such as the printing device 100 and the information processing device 200.
[0036] The language model 380 is a model that has been trained on a large amount of data. When a question is input to the language model 380, an answer to the question is output in text. The input to the language model 380 is text or an image. The language model 380 may be a large language model (LLM) or a small language model (SLM), and the scale of the model does not matter. The language model 380 may be a multimodal language model that also handles image input.
[0037] The embedding model 381 is a model that converts input natural language into vectors. The language model 380 may be included in the printing device 100 or the information processing device 200.
[0038] FIG. 5 shows an example of the screen transitions that are displayed on the panel 110 of the printing device 100 until the estimated ink level screen 1200 is displayed.
[0039] 5(a) shows a home screen 1000 displayed on the panel 110 of the printing device 100. The home screen 1000 is the first screen displayed on the panel 110 when the printing device 100 is turned on. The home screen 1000 includes a copy icon 1001, a scan icon 1002, and a print icon 1003. The bottom of the home screen 1000 includes a Wi-Fi button 1011, a settings button 1012, and an information button 1013.
[0040] When the user selects copy icon 1001, the screen changes to a screen for executing copying. A description of the screen for executing copying will be omitted. When the user selects scan icon 1002, the screen changes to a screen for executing scanning. A description of the screen for executing scanning will be omitted. When the user selects print icon 1003, the screen changes to a screen for executing printing. A description of the screen for executing printing will be omitted.
[0041] When the user selects the Wi-Fi button 1011, the screen transitions to a screen for setting up Wi-Fi. A description of the screen for setting up Wi-Fi will be omitted. When the user selects the settings button 1012, the screen transitions to a screen for making various settings for the printing device 100 (hereinafter referred to as settings screen 2100). When the user selects the information button 1013, the screen transitions to a screen that displays various information about the printing device 100 (hereinafter referred to as information screen 1100).
[0042] 5(b) shows an information screen 1100 displayed on the panel 110 of the printing device 100. The information screen 1100 includes a quick guide button 1101, an estimated ink level button 1102, and a system information button 1103. On the left side of the information screen 1100, there are a home button 1111 and a back button 1112.
[0043] When the user selects the quick guide button 1101, the quick guide is displayed. A description of the quick guide will be omitted. When the user selects the estimated ink level button 1102, the estimated ink level screen 1200 shown in FIG. 5(c) is displayed. When the user selects the system information button 1103, a screen relating to system information is displayed. A description of the screen relating to system information will be omitted. When the user selects the home button 1111, the screen transitions to the home screen 1000. When the user selects the back button 1112, the screen transitions to the previous screen, in this case, the home screen 1000.
[0044] 5(c) shows an estimated ink level screen 1200 displayed on the panel 110 of the printing device 100. The estimated ink level screen 1200 displays the estimated ink levels (remaining amounts) of various inks as bar graphs. An estimated cyan ink level 1201, an estimated magenta ink level 1202, an estimated yellow ink level 1203, and an estimated black ink level 1204 are displayed, respectively.
[0045] By checking the estimated ink level screen 1200, the user can ascertain the remaining amount and consider, for example, when to order replacement ink. When the user selects the home button 1211, the screen transitions to the home screen 1000. When the user selects the back button 1212, the screen transitions to the previous screen, in this case, the information screen 1100.
[0046] FIG. 6 shows an example of the screen transitions that are displayed on the panel 110 of the printing device 100 until the maintenance screen 2200 is displayed.
[0047] 6(a) shows a home screen 1000 displayed on the panel 110 of the printing device 100. The home screen 1000 has been described above, so a detailed description will be omitted.
[0048] 6(b) shows a settings screen 2100 displayed on the panel 110 of the printing device 100. The settings screen 2100 includes a main body settings button 2101, a paper settings button 2102, and a maintenance button 2103. A home button 2111 and a back button 2112 are provided on the left side of the settings screen 2100.
[0049] When the user selects the device settings button 2101, a screen related to device settings is displayed. A description of the screen related to device settings will be omitted. When the user selects the paper settings button 2102, a screen related to paper settings is displayed. A description of the screen related to paper settings will be omitted. When the user selects the maintenance button 2103, the maintenance screen 2200 shown in FIG. 6(c) is displayed. When the user selects the home button 2111, the screen transitions to the home screen 1000. When the user selects the back button 2112, the screen transitions to the previous screen, in this case, the home screen 1000.
[0050] 6(c) shows a maintenance screen 2200 displayed on the panel 110 of the printing device 100. The maintenance screen 2200 includes a nozzle check pattern print button 2201, a cleaning button 2202, and a head adjustment button 2203. The left side of the maintenance screen 2200 includes a home button 2211 and a back button 2212. When the user selects the home button 2211, the screen transitions to the home screen 1000. When the user selects the back button 2212, the screen transitions to the previous screen, in this case, the settings screen 2100.
[0051] 7 is a document 1400 that describes the procedure for displaying the estimated ink level screen 1200. The document 1400 includes text 1500, 1501, 1502, and 1503, and images 1511, 1512, and 1513.
[0052] Text 1500 is the title of this document 1400, describing how to check the status of ink. Text 1501 is step 1 for displaying the estimated ink level screen 1200.
[0053] Text 1501 advises the user to select information button 1013 on home screen 1000. Image 1511 is an image showing home screen 1000.
[0054] Text 1502 is step 2 for displaying estimated ink level screen 1200. Text 1502 instructs the user to select estimated ink level button 1102 on home screen 1000. Image 1512 is an image showing information screen 1100.
[0055] Text 1503 is step 3 for displaying estimated ink level screen 1200. Text 1503 states that estimated ink level screen 1200 is displayed and the estimated ink level can be checked. Image 1513 is an image showing estimated ink level screen 1200.
[0056] 8 shows a document 2400 that describes the procedure for cleaning the ink head. That is, the document 2400 describes the procedure up to the display of the maintenance screen 2200. The document 2400 includes text 2500, 2501, 2502, and 2503, and images 2511, 2512, and 2513.
[0057] Text 2500 is the title of this document 2400, which describes cleaning the print head. Text 2501 is step 1 for displaying the maintenance screen 2200.
[0058] Text 2501 instructs the user to select the settings button 1012 on the home screen 1000. Image 2511 is an image showing the home screen 1000.
[0059] Text 2502 is Procedure 2 for displaying the maintenance screen 2200. The text 2502 states that the maintenance button 2103 should be selected on the home screen 1000. An image 2512 is an image showing the settings screen 2100.
[0060] Text 2503 is step 3 for displaying the maintenance screen 2200. The text 2503 states that the cleaning button 2202 should be selected. An image 2513 is an image showing the maintenance screen 2200.
[0061] FIG. 9 shows an example of a process for generating a knowledge base. A knowledge base, such as RAG (Retrieval-Augmented Generation), is a database that serves as an information source for adding information to a language model. In the following explanation, the symbol "S" stands for step. RAG searches the knowledge base for information similar to the content of a user's question, combines the information obtained by the search with the user's question, and inputs it into the language model 380 to obtain a highly accurate answer.
[0062] 9 is a flowchart showing the processing executed by application 350 of cloud 300. This processing is executed by CPU 301 of cloud 300 reading a program stored in ROM 302 into RAM 203 and running application 350. This processing is started when a knowledge base creation function (not shown) is started. The knowledge base may be created only once, or multiple times at any timing.
[0063] In S100, the CPU 301 stores a document in the storage 304 of the cloud 300. A document is an information source that describes information such as a file, and becomes an information source for the knowledge base when it is accumulated in the storage 304. Examples of documents stored in the storage 304 here include a document 1400 that describes how to check the estimated ink level and a document 2400 that describes how to clean the nozzles.
[0064] The document may be the user manual for the printing device 100, or a new document created when creating the system of the present invention. The document may be a single page or a multi-page file. For example, the document may be a PDF file containing multiple pages of instructions on how to use the printing device 100, such as how to check the estimated ink level and how to clean the nozzles. The file may be in any format.
[0065] In S101, CPU 301 divides the document into data types such as text, tables, images, etc. For example, if the document is a single PDF file consisting of multiple pages, the file often contains a mixture of various types of data such as text, tables, and images. Because it is often difficult for language model 380 to understand such files, the content of the document is divided into data types such as text, tables, and images before being input to language model 380.
[0066] Such division may be realized by a general existing application or library such as a PDF parser. If the divided text is long, it may be further divided into multiple texts. The division of text may be realized by using a general existing language processing application or library. When dividing text, it may be divided based on the number of characters in the text, the meaning of the text, or the like.
[0067] In S102, the CPU 301 combines the texts divided in S101 by combining texts that are similar in meaning. For example, if ten pieces of text are obtained in S101, those with similar meanings are combined to create a total of three pieces of text. This process of combining similar texts may be implemented using an existing language processing application or library. This combining process is not a required configuration for the present invention. For example, the text obtained in S101 may be used as is in steps S103 and onward. In that case, S102 is not performed.
[0068] In S103, the CPU 301 digitizes the text obtained in S101 or S102. The digitization of the text may be performed using the embedded model 381 for the divided text, and vector data is obtained by inputting the text into the embedded model 381. The present invention does not limit the method of conversion to vector data. For example, other methods such as Word2Vec may be used to convert the text to text. Furthermore, the type of data when the text is digitized is not limited to vector data, and other types of numerical data may also be used.
[0069] In S104, the CPU 301 converts the image segmented in S101 into text. One method for converting an image into text is to use, for example, a multimodal language model 380. For example, when a prompt such as "Please describe this image" and an image are input into the language model 380, text describing the content of the image is obtained.
[0070] A prompt is a string of characters, such as a word or a sentence, that instructs language model 380 what to generate. While an example of converting an image to text using a multimodal language model has been described, images may be converted to text using other methods. For example, an image may contain text information, and the text may be extracted by performing OCR (Optical Character Reader) processing on the image.
[0071] In S105, the CPU 301 digitizes the text obtained from the image in S104. This digitization may be performed using the same means as the method described in S103.
[0072] In S106, the CPU 301 associates the text obtained in S101 with the digitized data obtained in S103 and stores them in the database 305. The database 305 stores data as sets of keys and values. Here, the key is the digitized data obtained in S103, and the value is the text obtained in S101. Note that the text obtained in S101 may be converted in some way, such as by a summary, and then used as the value.
[0073] In S107, the CPU 301 stores the image obtained in S101 and the digitized data obtained in S105 in association with each other in the database 305.
[0074] 10 is a conceptual diagram of the process of generating a knowledge base from document 1400. Texts 1500 to 1503 and images 1511 to 1513 are obtained by dividing document 1400. Text 1510 is obtained by combining texts 1500 to 1503. V100 is obtained by converting text 1510 into a numerical value.
[0075] It should be noted that "V" means vector. Texts 1521 to 1523 are obtained by converting images 1511 to 1513 into text. These texts are texts with contents that explain the images. V111 to V113 are obtained by digitizing texts 1521 to 1523.
[0076] 11 is a conceptual diagram of the process of generating a knowledge base from a document 2400. Texts 2500 to 2503 and images 2511 to 2513 are obtained by dividing the document 2400. Text 2510 is obtained by combining the texts 2500 to 2503. V200 is obtained by digitizing the text 2510.
[0077] It should be noted that "V" means vector. Texts 2521 to 2523 are obtained by converting images 2511 to 2513 into text. These texts are texts with content that explains the images. V211 to V213 are obtained by digitizing texts 2521 to 2523.
[0078] Fig. 12 is a diagram showing an example of a knowledge base stored in database 305. Specifically, the information created in Figs. 10 and 11 is stored in database 305. Text 1510 and V100 obtained by digitizing the text are associated with each other and stored in database 305. Furthermore, texts 1521 to 1523 and V111 to V113 obtained by digitizing the images are associated with each other and stored in database 305.
[0079] Similarly, text 2510 and V200 obtained by digitizing the text are associated with each other and stored in database 305. Furthermore, texts 2521 to 2523 and V211 to V213 obtained by digitizing their images are associated with each other and stored in database 305. In this way, information relating to one or more documents is stored in database 305.
[0080] 14 is a diagram showing an example of processing executed by application 150 of printing device 100. This processing is executed by CPU 101 of printing device 100 reading a program stored in ROM 102 into RAM 203 and running application 150.
[0081] This process starts when the printing device 100 is ready to accept a user question, for example, by pressing a button. If the printing device 100 is ready to accept a user question at any time, this process may start when the printing device 100 is turned on. In explaining Figure 14, Figure 15, which illustrates the process of Figure 14, will also be used.
[0082] 14, the CPU 101 controls the question receiving unit 140 to receive a question from the user 10. That is, the CPU 101 receives natural language from the user. For example, the printing device 100 is provided with a voice input device such as a microphone, and the CPU 101 receives a question by voice input when the user 10 speaks into the microphone.
[0083] Instead of creating a question from voice input, an input device such as a software keyboard or hardware keyboard may be provided in the printing device 100, and the user may input (text input) using the input device, which allows the CPU 101 to accept the user's question. The present invention may also use input methods other than those described above, and is not limited to any means that can accept a question from the user.
[0084] 14, CPU 101 converts the question received in S200 into text. For example, if a question is received via voice input, the voice data is converted into text by transcribing it. In the following description, this converted text will be referred to as question text. Note that if a question is received as text from a user via keyboard input in S200, S201 does not need to be executed.
[0085] 14, the CPU 101 digitizes the question text obtained in S201. The present invention is not limited to the digitization method. For example, the question text may be input to the embedded model 381 and converted into numerical data such as vector data. If the embedded model 381 is in the cloud 300, the CPU 101 controls the communication I / F 120 to send the question text to the cloud 300, and the embedded model 381 in the cloud 300 digitizes the question text.
[0086] By receiving the digitized data, the printing device 100 can digitize the question text. The present invention is not limited to the method of converting to numeric data using the embedding model 381. For example, the question text may be converted to numeric data such as vector data using the probability of word occurrence, as in the Bag of Words method. Note that the present invention does not require digitizing the question text; the text may be retained as is without being digitized. In this case, step S202 is not performed.
[0087] In S203 of Fig. 14, CPU 101 obtains data similar to the question received from the user in S200 from a knowledge base. In this knowledge base, text extracted from documents and numerical data are stored in association with each other, as shown in Fig. 12. That is, eight pieces of text 1510, 1521-1523, 2510, and 2521-2523 extracted from documents are stored in association with eight pieces of numerical data V100, V110-V113, V200, and V211-V213.
[0088] The distance between the vector data (Vq) of the question text generated in S202 and V100, V111 to V103, V200, and V201 to V213 is calculated, and N items with close distances are found. In the present invention, any method for calculating the distance between two pieces of numerical data is acceptable. For example, when performing calculations using vector data, a method such as cosine similarity may be used. N values corresponding to N keys with close distances are considered similar data. For example, when N=2, if V100 and V111 are close in distance to Vq, the similar data are text 1510 and 1521 corresponding to V100 and V111.
[0089] In the above description, the question text is converted into a numerical value, and data with a close distance is regarded as similar data. However, the present invention is not limited to this method. For example, similar data may be acquired using a search method using text, such as a full-text search or a keyword search, without converting the text into a numerical value. Furthermore, in the above-described method for acquiring similar data, similar data is acquired from a knowledge base, but the present invention is not limited to this method. For example, similar data close to a question may be acquired by inputting the question into a learning model that has been additionally trained using specific documents, such as fine-tuning.
[0090] 14, CPU 101 controls screen information acquisition unit 241 to acquire information about the screen displayed on panel 110. Information about the screen includes, for example, the "name of the screen," "screenshot of the screen," "screenshot of the screen converted into text," and "screen identifier." Note that what CPU 101 acquires at this time is information about the screen displayed on panel 110 when the user's question was accepted in S200.
[0091] The first “screen name” is the name of a screen, such as a “home screen” or a “maintenance screen.” In this case, CPU 101 acquires the name of the screen displayed on panel 110 when the user asks a question.
[0092] The second "screenshot of the screen" is a screenshot of the screen being operated by the user. The CPU 101 acquires a screenshot of the screen displayed on the panel 110 when the user asks a question. In FIG. 15, the screenshot 30 is a screenshot of the home screen 1000. The file name of the screenshot 30 is screenshot.png and is saved in the RAM 103 or the storage 104. Note that the screenshot is an example of image data and is not limited to screenshots, and the extension may be .peg, .gif, or the like.
[0093] The third example, "a screenshot of a screen converted into text," is text obtained by, for example, CPU 101 inputting a prompt such as "Summarize the image in this screenshot" and screenshot 30 into multimodal language model 380. This text is text data that explains the information in screenshot 30.
[0094] Other methods for converting a screenshot into text may also be used. For example, OCR processing may be performed on screenshot 30 to extract characters from the screenshot and convert them. CPU 101 acquires a screenshot of the screen displayed on panel 110 when the user asks a question, and converts this screenshot into text.
[0095] The fourth item, "screen identifier," refers to a character string or number that is registered in storage 104 or the like in association with each screen. For example, an identifier is registered in association with each screen, such as Screen 1 for home screen 1000 and Screen 100 for estimated ink level screen 1200. When the user asks a question, CPU 101 acquires the identifier corresponding to the screen displayed on panel 110.
[0096] 14, CPU 101 creates a prompt to be sent to language model 380. The user's question (natural language) accepted in S200 is used for the prompt, and the accepted natural language may be used as is, or the natural language may be converted so that it is easier for the language model to understand.
[0097] Examples of information that can be included in a prompt include "information about the question," "information about similar data," and "information about the screen." By including this information in the prompt, an answer to the user's question can be obtained based on knowledge base information for the screen the user is operating. In other words, it is not necessary to include all of this information in the prompt.
[0098] FIG. 13 shows an example of the prompt created by the CPU 101 of the printing device 100 in S205.
[0099] 13(a) is an example in which the information about the screen is "the name associated with the screen." Prompt 600 states, "The user is operating the home screen. Please answer the question at the end of the sentence using the following context. Context: XXXXX Question: YYYYY."
[0100] The expression "home screen" included in prompt 600 is the above-mentioned "information about the screen." Information about similar data is entered in XXXXX of the context. Information about the question is entered in YYYYY of the question.
[0101] Figure 13(b) is an example where the information about the screen is a "screenshot of the screen." Prompt 601 states, "The user is operating the attached screen. Please answer the question at the end of the sentence using the following context. Context: XXXXX Question: YYYYY."
[0102] The expression "attached screen" included in prompt 601 means a screenshot of the screen. XXXXX in the context contains information about similar data. YYYYY in the question contains information about the question.
[0103] Figure 13(c) is an example of a case where the information about the screen is "a screenshot of the screen converted into text." Prompt 602 states, "The user is operating the home screen. The home screen has copy, scan, and print buttons. The bottom of the screen has Wi-Fi, settings, and information buttons. Please answer the question at the end of the sentence using the following context. Context: XXXXX Question: YYYYY."
[0104] The prompt 602 contains information about the home screen, such as buttons. This information is obtained by converting a screenshot of the screen into text. The context XXXXX contains information about similar data. The question YYYYY contains information about the question.
[0105] Figure 13(d) is an example where the information about the screen is an "identifier associated with the screen." Prompt 603 states, "The user is operating Screen1. Please answer the question at the end of the sentence using the following context. Context: XXXXX Question: YYYYY."
[0106] The expression "Screen1" included in prompt 603 is the "screen identifier" mentioned above. Information about similar data is entered in XXXXX of the context. Information about the question is entered in YYYYY of the question.
[0107] 15 is an example of a prompt 50 created by applying the processes of S200 to S205 to a question 20 of a user 10. As described in FIG. 13(b), the prompt 50 is an example of a prompt in which the information about the screen is a "screenshot of the screen." The prompt 50 is composed of a prompt opening 51, a context item 52, and a question item 53.
[0108] The first prompt 51 states that when the user is operating the screen corresponding to screenshot.png on the panel 110, he or she should use a context item 52 to answer a question item 53.
[0109] Context item 52 lists similar data to question 20. Here, text 1510, 1521, 1522, and 1523 are listed as similar data to question 20. In this example, the user has asked question 20 about estimated ink levels, so text 1510, 1521, 1522, and 1523 related to estimated ink levels are listed as similar data. Note that this similar data has been obtained from the knowledge base (S203).
[0110] The user's question 20 is entered in the question item 53. Although the content of the user's question 20 is entered as is in the question item 53 in FIG. 15 , the present invention is not limited to this. For example, under the control of the CPU 101, the language model 380 may convert the user's question into an easily understandable form and enter this in the question item 53. Furthermore, both the user's question as is and the converted question may be entered in the question item 53. For example, if the user's question is long, the content may be summarized in advance by the language model 380, and this summary may be entered in the question item 53.
[0111] 14, the CPU 101 controls the communication I / F 120 to input the prompt created in S205 to the language model 380. In the case of prompts 600, 602, and 603 shown in FIGS. 13(a), 13(c), and 13(d), the prompts are sent to the language model 380.
[0112] At this time, the corresponding images may also be input to the language model 380. On the other hand, in the case of the prompt 601 shown in FIG. 13(b), an image that is a screenshot of the screen is input to the language model 380 together with the prompt 601. In addition, the image input here is a screenshot of the screen that is displayed when the natural language is accepted in S200. In this case, the language model 380 is a multimodal language model that can also accept images as input.
[0113] When the language model 380 is included in the printing device 100, the CPU 101 transmits a prompt to the language model in the printing device 100. When the language model 380 is included in the cloud 300, the CPU 101 controls the communication I / F 120 to transmit a prompt to the cloud 300, and the received prompt is input to the language model 380 of the cloud 300. The language model 380 that receives the prompt in this manner executes processing in accordance with the received prompt.
[0114] 14, the CPU 101 receives the output of the language model 380 via the communication I / F 120. If the language model 380 is in the cloud 300, the CPU 101 receives the output of the language model 380 from the cloud 300 via the communication I / F 120. If the language model 380 is included in the printing device 100, the CPU 101 receives the output from the language model. At this time, the language model 380 performs output based on the prompt and image input by the CPU 101 in S206.
[0115] 14, the CPU 101 displays the output of the language model 380 received in S207 on the panel 110 and presents it to the user. Note that the output of the language model 380 is a method for operating the printing device 100, and is the method for operating on the screen displayed on the panel 110 when the natural language was accepted in S200.
[0116] 15, the output (answer) of language model 380 is displayed on panel 110 as message 80. Message 80 contains content that urges the user to select information button 1013 and then select estimated ink level button 1102 on the subsequently displayed information screen. This is an example of an answer that indicates the operations to take to ultimately arrive at estimated ink level screen 1200, which the user ultimately wants to reach.
[0117] The first embodiment is an example in which all the operations required to realize the function the user wants to execute are answered in one go. That is, an answer indicating the operations required to reach the screen the user ultimately wants to reach from the screen displayed on panel 110 at the time the user asks the question is displayed on the screen displayed on panel 110 at the time the user asks the question. By looking at this message 80, the user can know the operations required to reach estimated ink level screen 1200.
[0118] In this way, the answer of language model 380 is displayed on panel 110 of printing device 100. Note that the method of presenting the answer of language model 380 to the user is not limited to this. As an example of another presentation method, if printing device 100 is equipped with an audio output device such as a speaker, the answer of language model 380 may be converted into audio and communicated (transmitted) to the user by audio.
[0119] In S209 of FIG. 14, the CPU 10 detects whether the user's operation was intended. An intended operation is whether the user performed an operation in accordance with the content presented to the user in S208. For example, as shown in FIG. 15, when the user 10 asks the question "I want to check the remaining ink level. What should I do?" 20 on the home screen 1000, a message is displayed on the home screen 1000.
[0120] The message is "Press the [Information] button to display the information screen. Select [Estimated Ink Levels] on the information screen." 80. At this time, if the user 10 presses the information button 1013 on the home screen 1000, this is an intended operation. On the other hand, if the user 10 selects the settings button 1012 on the home screen 1000, this is an unintended operation. Note that the configuration of S209 is not essential to the present invention.
[0121] 14, the CPU 101 displays on the panel 110 a message indicating that an incorrect operation has been performed, and notifies the user of this. This is performed when the user performs an operation different from the content presented in S208 (NO in S209).
[0122] If the CPU 101 detects an unintended operation in S209, a message indicating that an unintended operation has been performed is displayed in S210. For example, a message such as "The [Settings] button was pressed instead of the [Information] button. Please use the [Back] button to return to the previous screen" or "An incorrect operation was performed. Please press the home button on the left side of the screen to return to the home screen" may be displayed. Note that the configuration of S210 is not essential to the present invention.
[0123] <Second embodiment> A second embodiment of the present invention will be described below. The contents shown in FIGS. 1 to 13 are the same as those in the first embodiment, and therefore will be omitted. In the first embodiment, an example was described in which operations up to the final screen display are displayed all at once in response to a user's question. In contrast, in the second embodiment, an example will be described in which, in response to a single user question, the CPU 101 queries the language model 380 and receives an answer each time the user transitions between screens, and the answer is presented to the user for each screen transition.
[0124] 16 is a flowchart illustrating an example of the second embodiment. This flowchart shows processing executed by application 150 of printing device 100. This processing is executed by CPU 101 of printing device 100 reading a program stored in ROM 102 into RAM 203 and running application 150. This processing is started when printing device 100 becomes ready to accept a question from the user, for example, by pressing a button.
[0125] If the printing device 100 is capable of accepting user questions at any time, this process may be started when the printing device 100 is turned on. FIG. 17 illustrates the process shown in FIG. 16. In FIG. 17, similar to FIG. 15, the user 10 asks a question 20 on the home screen 1000: "I want to check the remaining ink level. What should I do?"
[0126] S300 to S304 in Fig. 16 are the same as S200 to S204 described in Fig. 14 of the first embodiment, and therefore a description thereof will be omitted. As a result of S304 in Fig. 16, a screenshot 30 is acquired as information relating to the screen, as shown in Fig. 17. This screenshot 30 is a screenshot of the home screen 1000. The following description will be given assuming that the file name of the screenshot 30 is screenshot.png.
[0127] 16, the CPU 101 creates a prompt to be sent to the language model 380. The information included in the prompt is the same as that described in the first embodiment, and therefore will not be described here.
[0128] An example of a prompt is shown in Fig. 17. The prompt 60 shown in Fig. 17 is composed of a prompt opening 61, a context item 62, and a question item 63. The prompt opening 61 states that the user is operating the screen screenshot.png on the panel 110, and that the user should use the context item 62 to answer the question item 63. In addition to the contents of the prompt opening 51 described in the first embodiment, the prompt opening 61 also states, "Please answer only about what should be done on the screen the user is operating."
[0129] With this added description, the answer from language model 380 will be only the operation that the user should perform on this screen. In other words, unlike message 80 in embodiment 1, it does not explain the operations up to the final screen, but displays what should be done step by step. Context item 62 is similar to context item 52 in Fig. 15 and therefore will be omitted. Question item 63 is similar to question item 53 in Fig. 15 and therefore will be omitted.
[0130] 16, the CPU 101 controls the communication I / F 120 to input the screen information acquired in S304 and the prompt created in S305 to the language model 380. The explanation for this is omitted as it is the same as S206.
[0131] 16, the CPU 101 receives the output of the language model 380 via the communication I / F 120. The explanation for this is omitted as it is similar to S207.
[0132] 16, the CPU 101 displays the output of the language model 380 received in S307 on the panel 110 and presents it to the user. An example of the output (answer) of the language model 380 displayed on the panel 110 is shown in FIG.
[0133] When user 10 is displaying home screen 1000, message 81 is displayed on panel 110. Message 81 reads, "Please select the [Information] button at the bottom of the screen." The expression "bottom of the screen" is included in text 1521 obtained by converting image 1511. In this way, if position information of operation controls such as buttons is also included when converting from an image to text, a message can be displayed that makes it easier for the user to operate the device.
[0134] 16, the CPU 101 detects whether the intended operation has been performed. The explanation for this is omitted as it is the same as that for S209.
[0135] 16, the CPU 101 notifies the user that an incorrect operation has been performed. The explanation for this is omitted as it is the same as that for S210.
[0136] 16, CPU 101 detects whether or not a transition to the target screen has occurred. If CPU 101 detects that a transition to the target screen has occurred (YES in S311), it ends the process. On the other hand, if CPU 101 detects that application 150 has transitioned to a screen different from the target screen (NO in S311), the process returns to S304.
[0137] The target screen will be explained using FIG. 17. Here, the target screen is a screen that the user wants to display or a screen that executes a function that the user wants to execute. In FIG. 17, the question 20 from user 10 is "I want to check the remaining ink level. What should I do?", so the target screen is estimated ink level screen 1200.
[0138] When the user 10 receives the message 81 and selects the information button 1013, the screen of the panel 110 transitions to the information screen 1100. Since the information screen 1100 after the screen transition is different from the target screen, the estimated ink level screen 1200, in S311 the CPU 101 determines that the transition to the target screen has not occurred (NO in S311), and returns to the processing of S304.
[0139] 16, if the determination is NO, S304 is executed again. In S304, information about the screen is acquired, and the information about the screen here is the screen after the transition, that is, information about the information screen 1100. In this example, a screenshot 31 of the information screen 1100 is acquired as information about the information screen 1100. The file name of the screenshot 31 is assumed to be screenshot.png.
[0140] The processing of S305 to S307 from the second time onwards is the same as the processing of the first time, so it will be omitted. In this example, in S308 the second time, a message 82 saying "Please select estimated ink level" is displayed. This is because the screenshot.png specified at the beginning of the prompt 61 is a screenshot of the information screen 1100. When the user sees this message 82 and selects the estimated ink level button 1102, the screen of the panel 110 transitions to the estimated ink level screen 1200.
[0141] In S311 from the second time onwards, CPU 101 determines whether or not a transition to the target screen has occurred. In this example, the target screen is estimated ink level screen 1200, and the display on panel 110 also becomes estimated ink level screen 1200, so CPU 101 determines that the target screen has been reached (YES in S311). This ends the processing. Note that if CPU 101 determines that the transition to the target screen has not occurred even in the second S311 (NO in S311), the process returns to S304, and a third round of processing is started. CPU 101 repeats this process until it determines in S311 that the transition to the target screen has occurred.
[0142] 18 shows an example of the sequence from when a user asks a question to when the output of the language model 380 is displayed on a panel, in accordance with the first and second embodiments. The characters in the sequence diagram are the user 10, the printing device 100, and the cloud 300. The printing device 100 may also be the information processing device 200.
[0143] In F100, the user 10 asks a question to the printing device 100, which then accepts the question via the question accepting unit 140 (S200, S300).
[0144] In F101, the CPU 101 of the printing device 100 converts the question received in F100 into text (S201, S301).
[0145] In F102, the CPU 101 of the printing device 100 converts the question text converted in F101 into a numerical value (S202, S302).
[0146] In F103, the CPU 101 of the printing device 100 controls the communication I / F 120 to send to the cloud 300 a request to acquire data similar to the question received in F100, and the cloud 300 receives this.
[0147] In F104, the CPU 301 of the cloud 300 searches for data similar to the question accepted by the printing device 100 in F100 in response to the acquisition request received from the printing device 100 in F103.
[0148] In F105, the CPU 301 of the cloud 300 controls the communication I / F 320 to transmit the similar data searched for in F104 to the printing device 100 (S203, S303).
[0149] In F106, the CPU 101 of the printing device 100 acquires a screenshot of the screen displayed on the panel 110 (S204, S304).
[0150] In F107, the CPU 101 of the printing device 100 uses the question accepted in F100 and the similar data received in F105 to create a prompt to be sent to the language model 380 (S205, S305).
[0151] In F108, the CPU 101 of the printing apparatus 100 controls the communication I / F 120 to send the screenshot acquired in F106 and the prompt created in F107 to the cloud 300 (S206, S306).
[0152] In F109, the CPU 301 of the cloud 300 inputs the screenshot and the prompt into the language model 380 and processes it to obtain an answer.
[0153] In F110, the CPU 301 of the cloud 300 controls the communication I / F 320 to transmit the answer (output) of the language model 380 obtained in F109 to the printing device 100, which receives it (S207, S307).
[0154] In F111, the CPU 101 of the printing device 100 displays the answer (output) from the language model 380 received in F110 on the panel 110 (S208, S308).
[0155] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
Claims
1. A multifunction device, a display means for displaying a screen; Acceptance means for accepting natural language from a user; a transmitting means for transmitting information on the screen displayed when the receiving means receives the natural language and a prompt using the received natural language; receiving means for receiving a language model output, the output being based on the transmitted screen information and the transmitted prompt; the display means displays the output in response to the reception of the output by the reception means; The output is an operation method of the multifunction peripheral, and is the operation method displayed on the screen displayed by the display means when the accepting means accepts the natural language. A multifunction device characterized by the above.
2. 2. The multifunction peripheral according to claim 1, wherein the display means displays the received output on the screen that was displayed when the accepting means accepted the natural language upon receipt of the output by the receiving means.
3. 3. The multifunction peripheral according to claim 2, further comprising a notification unit that notifies the user of an operation after the output is displayed by the display unit.
4. 4. The multifunction peripheral according to claim 3, wherein the notification means notifies the user that an operation different from the operation method indicated by the displayed output has been performed when the user performs an operation different from the operation method indicated by the displayed output.
5. In response to a screen transition caused by an operation by the user after the output is displayed by the display means, the transmitting means transmits information about the screen after the transition and a prompt using the accepted natural language; the receiving means receives an output of the language model, the output being based on the transmitted information on the screen after the transition and the transmitted prompt; 2. The multifunction peripheral according to claim 1, wherein the display unit displays the output when the receiving unit receives the output.
6. Further comprising a search means, the search means searches for data similar to the natural language accepted by the acceptance means; The transmitting means transmits information on the screen displayed when the accepting means accepts the natural language and a prompt using the accepted natural language and the data searched by the searching means.
2. The multifunction device according to claim 1, wherein:
7. 2. The multifunction peripheral according to claim 1, wherein the transmitting means transmits a prompt based on the information on the screen displayed when the accepting means accepts the natural language and the accepted natural language.
8. 2. The multifunction peripheral according to claim 1, wherein the accepting unit accepts natural language input by voice.
9. 2. The multifunction peripheral according to claim 1, wherein the accepting unit accepts natural language input by text input.
10. 2. The multifunction peripheral according to claim 1, wherein the prompt using natural language is a prompt including natural language received from the user by the receiving means.
11. 2. The multifunction peripheral according to claim 1, wherein the prompt using natural language is a prompt converted based on the natural language accepted by the accepting unit from the user.
12. 2. The multifunction peripheral according to claim 1, wherein the information about the screen is a name associated with the screen.
13. 2. The multifunction peripheral according to claim 1, wherein the screen information is image data of the screen.
14. 14. The multifunction peripheral according to claim 13, wherein the image data is a screenshot of a screen.
15. 8. The multifunction peripheral according to claim 7, wherein the information on the screen is a screenshot of the screen converted into text.
16. 2. The multifunction peripheral according to claim 1, wherein the information about the screen is an identifier associated with the screen.
17. having an audio output means, 2. The multifunction peripheral according to claim 1, wherein the audio output unit outputs the output received by the receiving unit as audio.
18. 18. A program for causing a computer to execute each of the means of the multifunction peripheral according to claim 1.
19. A control method for a multifunction peripheral, comprising: a display step for displaying a screen; a receiving step of receiving natural language from a user; a sending step of sending information on the screen displayed when the natural language is received in the receiving step and a prompt using the received natural language; receiving a language model output based on the transmitted screen information and the transmitted prompt; the display step displays the output in response to the reception step of the output, The output is an operation method of the multifunction peripheral, and is the operation method on the screen displayed by the display step when the natural language is received in the receiving step. A control method for a multifunction peripheral.
Citation Information
Patent Citations
Method, computer device, and computer program for providing dialogue dedicated to domain by using language model
JP2023076413A