Image processing device, control method, and program

The image processing device automates the proofreading of paper manuscripts by connecting with a large-scale language model server, reducing user burden and enhancing accuracy through automated scanning and prompt generation.

JP2025136921APending Publication Date: 2025-09-19CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024035859
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-08
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing methods require manual pre-proofreading of paper manuscripts before scanning, which increases user burden and does not fully alleviate the proofreading task.

Method used

An image processing device that communicates with a large-scale language model server to automatically acquire, generate prompts for, and transmit image data for proofreading, receiving highly accurate results with simple operations.

Benefits of technology

Obtains highly accurate proofreading of paper documents with minimal user interaction by automating the scanning and proofreading process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025136921000001_ABST
    Figure 2025136921000001_ABST
Patent Text Reader

Abstract

To provide a mechanism that allows highly accurate proofreading results of paper-medium documents to be obtained with simple operations.SOLUTION: An image processing device 1 that is communicatively connected to a large-scale language model server 2 that performs inference using a trained model acquires image data read from a document by a scanner 220, generates a prompt for causing the image data to be proofread, transmits the image data and the prompt to the large-scale language model server 2, causes the large-scale language model server 2 to proofread the transmitted image data based on the transmitted prompt using the trained model and output the proofreading results, and outputs the proofreading results output from the large-scale language model server 2 to an output destination.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing apparatus, a control method, and a program, and more particularly to an image processing apparatus, a control method, and a program for performing proofreading using an external server. [Background technology]

[0002] In recent years, generative AI (Artificial Intelligence) such as ChatGPT has been attracting attention. In recent years, mainstream generative AI uses pre-trained models called LLMs (Large Language Models), which are characterized by generating sentences and images based on instructions in natural language text. Large-scale language models, such as GPT-4.0, are also emerging as multimodal models that can handle inputs such as image data and audio data in addition to natural language text.

[0003] Large-scale language models build databases by learning huge amounts of data, which requires large amounts of storage, so services are generally built on large servers.

[0004] Furthermore, the use of large-scale language models is expanding beyond generating text and images to include proofreading input text. In this type of use, users must prepare the data they want to proofread on their own device and then send the data from their device to the server on which the large-scale language model is built. When obtaining the data to be proofread as image data from a paper manuscript, users must take the extra step of scanning the manuscript with a scanner, importing the data into a device such as a PC, and then sending the data from their device to the server.

[0005] On the other hand, Patent Document 1 discloses a technique in which scanned data is sent to a server and the presence of proofreading information in the image data is detected. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] JP 2017-11490 A Summary of the Invention [Problem to be solved by the invention]

[0007] However, in Patent Document 1, the manuscript must be manually proofread before it is scanned, and there is a problem in that the burden of proofreading on the user is only partially alleviated.

[0008] Therefore, the present invention provides a mechanism that allows highly accurate proofreading results of paper-based manuscripts to be obtained with simple operations. [Means for solving the problem]

[0009] In order to solve the above problem, the image processing device of claim 1 of the present invention is an image processing device that is communicatively connected to a review server that performs inference using a learned model, and is characterized in that it has an acquisition means for acquiring image data read from a manuscript, a generation means for generating a prompt for reviewing the acquired image data, a review result request means for transmitting the acquired image data and the generated prompt to the review server, causing the review server to perform review of the acquired image data based on the generated prompt using the learned model and output the review results, and an output means for outputting the review results output from the review server to an output destination. [Effects of the Invention]

[0010] According to the present invention, highly accurate proofreading results for a paper document can be obtained with simple operations. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a diagram showing a basic configuration of a system including an image processing apparatus according to an embodiment of the present invention. [Figure 2] 1 is a block diagram showing a basic internal hardware configuration of an image processing apparatus. [Figure 3] FIG. 2 is a block diagram showing the basic internal hardware configuration of a large-scale language model server. [Figure 4] 10 is a flowchart of a proofreading scan process according to the present embodiment. [Figure 5] FIG. 2 is a diagram showing a main menu screen displayed on the display in FIG. 1. [Figure 6] FIG. 10 is a diagram showing a review scan screen displayed on a display. [Figure 7] 10 is a flowchart of a display control process of a preview screen. [Figure 8] FIG. 5 is a diagram showing a proofreading scan screen including a preview screen, which is displayed in step S411 of FIG. 4. [Figure 9] FIG. 10 is a diagram showing an example of image data output from the large-scale language model server as a proofreading result, in which the proofreading result is displayed on image data obtained by scanning a document on which a notice is printed. [Figure 10] 10 is a diagram showing image data converted from the list of text that has been converted from the proofreading results shown in FIG. 9 from the large-scale language model server. FIG. [Figure 11] FIG. 10 is a diagram showing an example in which image data is output from the large-scale language model server as a proofreading result, in which the proofreading result is displayed on image data obtained by scanning a document on which a form is printed. [Figure 12] 10 is a flowchart of a proof scan data transmission execution process performed by the CPU 201 while a proof scan data transmission screen is displayed. [Figure 13] FIG. 8 is a diagram showing a proofread scan data transmission screen displayed in step S714 of FIG. 7. [Figure 14] FIG. 10 is a diagram showing an AI destination setting screen. [Figure 15] FIG. 5 is a diagram showing a prompt generation table used in step S408 of FIG. DETAILED DESCRIPTION OF THE INVENTION

[0012] The best mode for carrying out the present invention will be described below with reference to the drawings. Note that the following embodiments do not limit the scope of the invention as claimed, and not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.

[0013] FIG. 1 is a diagram showing the basic configuration of a system including an image processing apparatus according to this embodiment.

[0014] In FIG. 1, an image processing device 1 and a large-scale language model server 2 are communicably connected via a network to form one system.

[0015] The large-scale language model server 2 generates output data internally based on instructions and data input from the image processing device 1 via the network, and returns the output data to the image processing device 1 via the network.

[0016] FIG. 2 is a block diagram showing the basic internal hardware configuration of the image processing device 1. As shown in FIG.

[0017] The image processing device 1 includes, as devices connected to a system bus 202, a CPU 201, a GPU 203, an eMMC 204, a RAM 205, a network controller 206, a USB host controller 208, and a USB device controller 210. The image processing device 1 also includes, as devices connected to the system bus 202, a display controller 212, an input unit controller 214, an RTC 216, a SATA I / F 217, a scanner I / F 219, a printer I / F 221, and a counter 223. These devices can communicate with each other via the system bus 202.

[0018] The CPU 201 is a central processing unit that runs software for operating the image processing device 1.

[0019] The GPU 203 is a graphics processing unit that performs image processing.

[0020] The eMMC 204 stores the software (programs) of the image processing device 1, as well as databases and temporary storage files required for the image processing device 1 to operate.

[0021] The RAM 205 is a random access memory in which the program of the image processing apparatus 1 is developed and which serves as a storage area for variables during program operation and data transferred from each unit by direct memory access (DMA).

[0022] The network controller 206 controls a network controller I / F 207 that communicates with the image processing device 1 and other devices on the network. The CPU 201 transmits image data and prompts (described later) to the large-scale language model server 2 via the network controller 206 and the network controller I / F 207. The CPU 201 also receives, via the network controller 206 and the network controller I / F 207, proofreading results inferred by a trained model (described later) based on the image data and prompts transmitted from the large-scale language model server 2.

[0023] A USB host controller 208 controls a USB host I / F 209 for connecting to and communicating with a USB device when the image processing apparatus 1 functions as a USB host.

[0024] The USB device controller 210 controls a USB device I / F 211 for connecting to and communicating with a USB host when the image processing apparatus 1 functions as a USB device.

[0025] The display controller 212 controls the display 213 that displays the operating status of the image processing device 1 so that the user can check it.

[0026] The input unit controller 214 controls the input unit 215, which receives instructions from a user to the image processing device 1. Specifically, the input unit 215 is an input system such as a keyboard, mouse, numeric keypad, cursor keys, touch panel, or operation keyboard. When the input unit 215 is a touch panel, it is physically mounted on the surface of the display 213.

[0027] The RTC216 is a real-time clock with clock, alarm, and timer functions.

[0028] The SATA I / F 217 is an interface for connecting an SSD 218 to the image processing apparatus 1. The SSD 218 stores large files, temporary application data, and the like.

[0029] The scanner I / F 219 is an interface for connecting a scanner 220 that reads an image from a document. The CPU 201 acquires image data of an image read from a document by the scanner 220 from the scanner I / F 219 (acquisition means).

[0030] The printer I / F 221 is an interface for connecting a printer 222 for printing image data.

[0031] Counter 223 records the number of pages in a print job.

[0032] FIG. 3 is a block diagram showing the basic internal hardware configuration of the large-scale language model server 2.

[0033] 3, the large-scale language model server 2 includes, as devices connected to a system bus 302, a CPU 301, a RAM 303, a GPU 304, a VRAM 305, a communication unit 306, and a storage unit 307. These devices can communicate with each other via the system bus 302.

[0034] The CPU 301 runs software that runs the large-scale language model server 2 .

[0035] The RAM 303 expands the programs executed by the CPU 301, and the storage unit 307 is used as a storage area for the programs and data.

[0036] The communication unit 306 can communicate with external devices including the image processing device 1 via a network, and inputs and outputs data to and from the external devices.

[0037] A model trained by machine learning (trained model) is stored in the storage unit 307. For example, the trained model may be a model using a neural network architecture called Transformer.

[0038] The GPU 304 is used for processing to perform inference using the trained model stored in the storage unit 307 and generate output data.

[0039] The VRAM 305 is used as a location for expanding data used when the GPU 304 performs processing. When the large-scale language model server 2 (proofreading server) receives image data and instructions called prompts input from an external device via the communication unit 306, the CPU 301 uses the GPU 304 to perform inference from the trained model stored in the storage unit 307. Data output as a result of the inference (e.g., image data showing the proofreading results described below) is returned to the external device via the communication unit 306. The large-scale language model server 2 can accept prompts in natural language, allowing the external device to provide flexible instructions and extract various outputs from the trained model. Note that while an example of performing inference using the GPU 304 has been described here, the CPU 301 may also perform the inference.

[0040] 4 is a flowchart of the proofreading scan process according to this embodiment. This process is executed by the CPU 201 reading out a program stored in the eMMC 204 and loading it into the RAM 205.

[0041] This process will be explained below using screens shown in Figures 5 and 6 that are displayed on the display 213 when the flowchart in Figure 4 is executed. Note that the icon arrangements and wording on the screens shown in Figures 5 and 6 are merely examples and do not limit the configuration of the present invention.

[0042] This process starts with the main menu screen of FIG.

[0043] First, in step S401, when the CPU 201 detects that the proof scan icon 501 in the main menu screen (FIG. 5) has been pressed, the process proceeds to step S402, where the proof scan screen (first display means) of FIG.

[0044] As shown in FIG. 6, the proofreading scan screen displays buttons 601 to 610 and 615 for receiving instructions from the user regarding the proofreading content, and a start button 611, which can be selected by the user.

[0045] 4, if CPU 201 detects that at least one of buttons 601 to 610, 615 has been pressed (step S404—Yes) before start button 611 is pressed (step S403—No), CPU 201 proceeds to step S405. On the other hand, if not (step S404—No), CPU 201 returns to step S403.

[0046] In step S405, CPU 201 changes the pressed state of the button for which pressing has been detected from "not pressed" to "pressed." Here, "pressed" refers to a state in which the user has selected the button, and "not pressed" refers to a state in which the user has not selected the button. The pressed state of each of buttons 601 to 610, 615 is saved in RAM 205. Note that, in step S404, if CPU 201 detects that a button whose pressed state has changed to "pressed" has been pressed again, CPU 201 changes the pressed state of the button for which re-pressing has been detected in step S405 to "not pressed." Furthermore, the "review only" button 601 and the "review + edit" button 602 are in an exclusive relationship, and the pressed state of only one of them can be changed to "pressed."

[0047] When the CPU 201 detects that the start button 611 has been pressed (step S403—Yes), the process proceeds to step S406, where it determines whether the pressed state of one or more of the buttons 603 to 610 (test items) is “pressed.” If none of the buttons is “pressed” (step S406—No), the CPU 201 displays a warning message (step S407), and then returns to step S402, where the proofreading scan screen is displayed again. On the other hand, if any button is “pressed” (step S406—Yes), the CPU 201 (generation means) generates a prompt corresponding to the “pressed” button (step S408). The generated prompt may be, for example, an instruction sentence in natural language, and creates a sentence instructing the large-scale language model server 2 to proofread the image data.

[0048] In this embodiment, the CPU 201 generates the prompt in step S408 by referring to a prompt generation table (FIG. 15) stored in advance in the eMMC 204.

[0049] As shown in FIG. 15, the prompt generation table includes items 1501 to 1512 and prompts 1513 to 1524 in natural language corresponding to the items 1501 to 1512, respectively.

[0050] Items 1502 to 1509 correspond to buttons 603 to 610, item 1510 corresponds to buttons 601 and 602, item 1511 corresponds to button 615, and item 1512 corresponds to button 602.

[0051] When generating the prompt in step S408, the CPU 201 first reads into the RAM 205 the prompt 1513 corresponding to the "basic instruction" item 1501.

[0052] Next, if any of the buttons 603 to 610 is in the “pressed” state, the CPU 201 reads into the RAM 205 the prompt corresponding to that button.

[0053] Furthermore, when the pressed state of either the "review only" button 601 or the "review and edit" button 602 is "currently being pressed," the CPU 201 reads out a prompt 1522 for "instructions to image the results" into the RAM 205. Furthermore, when the pressed state of the "review and edit" button 602 (instructions item for review and edit) is "currently being pressed," the CPU 201 reads out a prompt 1524 for "instructions to edit images" (correction instruction prompt) into the RAM 205. Furthermore, when the pressed state of the "text" button 615 (instructions item for text) is "currently being pressed," the CPU 201 reads out a prompt 1523 for "instructions to text the results" (text instruction prompt) into the RAM 205.

[0054] Finally, the CPU 201 concatenates all the prompts read into the RAM 205 to generate one prompt.

[0055] 4, after generating the prompt in step S408, CPU 201 scans the document (step S409). The document is scanned by CPU 201 issuing a read instruction to scanner 220 via scanner I / F 219. Then, CPU 201 (proofreading result request means) stores the scanned image data in eMMC 204 or RAM 205.

[0056] Next, the CPU 201 (proofreading result request means) transmits the image data and the prompt to the large-scale language model server 2 via the network controller 206 and the network I / F 207 (step S410). The data can be transmitted, for example, using a WebAPI that uses HTTP communication.

[0057] When the CPU 301 of the large-scale language model server 2 receives the image data and the prompt from the image processing device 1 in step S410, it reviews the image data based on the prompt and outputs the review results to the image processing device 1. The review results may be image data such as those shown in FIGS. 9 and 11 (described later), or text data showing a list of the review results as a text, as shown in FIG. 10. The CPU 301 transmits the output review results to the image processing device 1 via the communication unit 306, and the CPU 201 stores the received data in the eMMC 204 or RAM 205. At this time, if the "Review Only" button 601 is in the "Pressed" state, the CPU 201 receives only the reviewed image data. On the other hand, if the "Review + Correct" button 602 is in the "Pressed" state, the CPU 201 receives corrected image data in which errors have been corrected based on the review results, along with the reviewed image data. If the CPU 201 receives the corrected image data along with the reviewed image data, it adds the corrected image data to the reviewed image data as a separate page. If the "Text" button 615 is in the "Pressed" state, text data in which the proofreading results have been converted into text is received.

[0058] In step S411, the CPU 201 (output means) stores the received proofreading results in the eMMC 204 or RAM 205 (output destination), displays the received proofreading results on the preview screen 801 (FIG. 8) of the display 213 (output destination), and ends this process. In step S411, image data of the proofreading results, such as those shown in FIGS. 9 and 11, is displayed on the preview screen 801.

[0059] FIG. 8 is a diagram showing a proofreading scan screen including a preview screen 801, which is displayed in step S411 of FIG.

[0060] The proof scan screen includes a preview screen 801, a cancel button 802, an original data button 803, a proof result button 804, a corrected data button 805, and page switching buttons 806 and 807. The proof scan screen also includes a send button 808, a save button 809, and a print button 810. These buttons are configured as user-selectable icons.

[0061] Next, the flow of the display control process for the preview screen 801 displayed in step S411 will be described.

[0062] 7 is a flowchart of a preview screen display control process that is executed in response to a user operation after the preview screen is displayed in step S411. This process is executed by the CPU 201 reading a program stored in the eMMC 204 and loading it into the RAM 205.

[0063] 9 to 11, this process will be described using the proofreading results from the large-scale language model server 2, which are displayed on the preview screen 801 by the CPU 201. Note that the image data examples in FIGS. 9 to 11 are merely examples and do not limit the scope of the invention.

[0064] First, in step S701, the CPU 201 detects whether or not a user operation has been performed on the proofreading scan screen. If a user operation has been performed (step S701-Yes), the process proceeds to step S702. On the other hand, if a user operation has not been performed (step S701-No), the process waits until a user operation is performed.

[0065] When the CPU 201 detects that the cancel button 802 has been pressed (step S702-Yes), it switches the preview screen 801 of the proof scan screen in Fig. 8 to a blank display (step S703) and ends this process. On the other hand, if the cancel button 802 has not been pressed (step S702-No), the process proceeds to step S704 and detects whether the original data button 803 has been pressed.

[0066] When the CPU 201 detects that the Original Data button 803 has been pressed (step S704-Yes), it displays the unproofread manuscript data scanned in step S409 on the preview screen 801 (step S705). On the other hand, if the Original Data button 803 has not been pressed (step S704-No), the CPU 201 proceeds to step S706 and detects whether the Proofread Result button 804 has been pressed.

[0067] When the CPU 201 detects that the proofreading result button 804 has been pressed (step S706-Yes), it displays the proofreading result output from the large-scale language model server 2 on the preview screen 801. On the other hand, if the proofreading result button 804 has not been pressed (step S706-No), the CPU 201 proceeds to step S708, where it detects whether the corrected data button 805 has been pressed.

[0068] 9 is a diagram showing an example of image data 900 output from the large-scale language model server 2 as a proofreading result, in which the proofreading result is displayed on image data obtained by scanning a document on which a notice is printed. When the proofreading result button 804 (FIG. 8) is pressed, the image data 900 is displayed on the preview screen 801 by the CPU 201, and an overview of the proofreading result is displayed in association with the highlighted portions enclosed in frames, as in the proofreading items 901-904. Note that the highlighted portions may be highlighted in the image data 900 in a manner other than by enclosing them in a frame as in the example shown in FIG. 9. For example, the highlighted portions may be displayed in bold or red, or may blink.

[0069] Fig. 11 is a diagram showing an example of image data 1100 output as a proofreading result from the large-scale language model server 2, in which the proofreading result is displayed on image data obtained by scanning a document on which a form has been printed. Similar to the image data 900 in Fig. 9, the image data 1100 in Fig. 11 is displayed on the preview screen 801 when the proofreading result button 804 (Fig. 8) is pressed, and proofreading items corresponding to the pointed-out portions (in this case, numerical errors) surrounded by boxes are displayed, such as proofreading items 1101.

[0070] 7, when the CPU 201 detects that the corrected data button 805 has been pressed (step S708-Yes), it detects whether corrected image data exists in the eMMC 204 or RAM 205 (step S709). If corrected image data exists (step S709-Yes), the CPU 201 displays the corrected image data (step S710) and proceeds to step S711. On the other hand, if the corrected data button 805 has not been pressed or if corrected image data does not exist (step S708-No or step S709-No), the CPU 201 proceeds directly to step S711.

[0071] In step S711, the CPU 201 detects whether either of the page switching buttons 806 and 807 has been pressed.

[0072] The CPU 201 detects whether one of the page switching buttons 806, 807 has been pressed (user instruction) (step S711). If one of the page switching buttons 806, 807 has been pressed (step S711-Yes), the CPU 201 switches to display another page of the image data currently being displayed on the preview screen 801 (such as corrected image data or image data 1000 (FIG. 10)). On the other hand, if neither of the page switching buttons 806, 807 has been pressed (step S711-No), the CPU 201 proceeds to step S713 and detects whether the send button 808 has been pressed.

[0073] 10 is a diagram showing image data 1000 converted from the list of textualized proofreading results shown in FIG. 9 from the large-scale language model server 2. When the CPU 201 (conversion means) receives the list of textualized proofreading results shown in FIG. 9 from the large-scale language model server 2, it converts this list into image data 1000. Thereafter, the CPU 201 adds the image data 1000 as a separate page to the proofreading results (image data 900) output from the large-scale language model server 2. Furthermore, when either of the page forward buttons 806, 807 (FIG. 8) is pressed while the image data 900 in FIG. 9 is displayed on the preview screen 801, the CPU 201 switches the preview screen 801 to the image data 1000 in FIG. 10. The image data 1000 displays a list of pointed out items, such as pointed out items 1001 to 1004, each consisting of the pointed out location and content of the pointed out items in the image data 900.

[0074] Returning to FIG. 7, when the CPU 201 detects that the Send button 808 has been pressed (step S713—Yes), the process proceeds to step S714. In step S714, the CPU 201 displays a Send screen 1300 for proofread scan data (FIG. 13: second display means) that displays multiple destinations that can be set as destinations and allows the user to select one, and then terminates this process. The user can send the original data or image data 900 (and corrected data, if already acquired) to the desired destination by operating the Send screen 1300 for proofread scan data. On the other hand, if the Send button 808 has not been pressed (step S713—No), the process proceeds to step S715, where it is detected whether the Save button 809 has been pressed. The proofread scan data is one of the following: scanned, unproofread manuscript data; image data output from the large-scale language model server 2 as the proofreading result; and image data in which errors have been corrected based on the proofreading result output from the large-scale language model server 2.

[0075] When the CPU 201 detects that the Save button 809 has been pressed (step S715-Yes), it saves the scanned image data and the image data output by the large-scale language model server 2 in a predetermined area of ​​the eMMC 204 (step S716), and ends this processing. On the other hand, if the Save button 809 has not been pressed (step S715-No), the CPU 201 proceeds to step S717, where it detects whether the Print button 810 has been pressed.

[0076] If CPU 201 detects that print button 810 has been pressed (step S717-Yes), the process proceeds to step S718. On the other hand, if print button 810 has not been pressed (step S717-No), the process returns to step S702.

[0077] In step S718, the CPU 201 transmits the image data to the printer 222 via the printer I / F 221, executes printing, and then ends this process. Note that the image data to be printed may be specified by the user.

[0078] Next, the flow of the proof scan data transmission execution process that is executed when the proof scan data transmission screen 1300 is displayed in step S714 will be described.

[0079] 12 is a flowchart of the proof scan data transmission execution process performed by the CPU 201 while the proof scan data transmission screen 1300 is displayed. This process is executed by the CPU 201 reading out a program stored in the eMMC 204 and loading it into the RAM 205.

[0080] First, in step S1200, CPU 201 detects whether or not a user operation has been performed on proofread scan data transmission screen 1300. If a user operation has been performed (step S1200-Yes), the process proceeds to step S1201, where it is detected whether or not back button 1310 has been pressed. On the other hand, if a user operation has not been performed (step S1200-No), the process waits until a user operation is performed.

[0081] If the CPU 201 detects that the back button 1310 has been pressed (step S1201-Yes), it displays the preview screen 801 (step S1202) and ends this process. On the other hand, if the back button 1310 has not been pressed (step S1201-No), the process proceeds to step S1203, where it is detected whether the address book button 1302 has been pressed.

[0082] When CPU 201 detects that address book button 1302 has been pressed (step S1203-Yes), it accepts the destination setting from the address book, sets it as the transmission destination (step S1204), and then proceeds to step S1205. On the other hand, if address book button 1302 has not been pressed (step S1203-No), it proceeds directly to step S1205 and detects whether To Me button 1303 has been pressed.

[0083] When CPU 201 detects that To Me button 1303 has been pressed (step S1205-Yes), it sets the email address of the operating user as the destination (step S1206) and then proceeds to step S1207. On the other hand, if To Me button 1303 has not been pressed (step S1205-No), it proceeds directly to step S1207 and detects whether To Cloud button 1304 has been pressed.

[0084] When CPU 201 detects that To Cloud button 1304 has been pressed (step S1207-Yes), it accepts the cloud URL setting and sets it as the destination (step S1208), and then proceeds to step S1209. On the other hand, if To Cloud button 1304 has not been pressed (step S1207-No), it proceeds directly to step S1209 and detects whether Fax button 1305 has been pressed.

[0085] When CPU 201 detects that fax button 1305 has been pressed (step S1209-Yes), it accepts the fax destination setting, sets it as the destination (step S1210), and then proceeds to step S1211. On the other hand, if fax button 1305 has not been pressed (step S1209-No), it proceeds directly to step S1211 and detects whether original data button 1307 has been pressed.

[0086] When CPU 201 detects that Original Data button 1307 has been pressed (step S1211-Yes), it sets the scanned image data (original data) to be used as the transmission data (step S1212), and then proceeds to step S1213. On the other hand, if Original Data button 1307 has not been pressed (step S1211-No), it proceeds directly to step S1213 and detects whether or not Review Result button 1308 has been pressed.

[0087] When CPU 201 detects that proofreading result button 1308 has been pressed (step S1213-Yes), it sets the proofreading result data (here, the image data in FIGS. 9 and 10) to be used as transmission data (step S1214), and then proceeds to step S1215. On the other hand, if proofreading result button 1308 has not been pressed (step S1213-No), it proceeds directly to step S1215, where it detects whether corrected data button 1309 has been pressed. Here, if corrected data button 1309 has not been pressed (step S1215-No), it proceeds to step S1218, where it detects whether execute button 1306 has been pressed.

[0088] On the other hand, when CPU 201 detects that corrected data button 1309 has been pressed (step S1215-Yes), it detects whether corrected data exists (step S1216). If corrected data does not exist (step S1216-No), a warning message is displayed (step S1219), and the process proceeds to step S1218. On the other hand, if corrected data exists (step S1216-Yes), the process sets the corrected data to be used as transmission data (step S1217), and the process proceeds to step S1218.

[0089] When CPU 201 detects that execute button 1306 has been pressed (step S1218-Yes), it transmits the transmission data set as the transmission data to the transmission destination set as the destination (step S1220), and then ends this processing. On the other hand, if execute button 1306 has not been pressed (step S1218-No), it returns to step S1201.

[0090] 14 is a diagram showing the AI ​​destination setting screen. Here, the AI ​​destination setting screen (destination setting UI) is a screen on which access settings for the large-scale language model server 2 used for proofreading can be set. This screen is displayed when the AI ​​destination setting menu is selected on the menu screen (not shown) that is displayed when the setting button 1311 on the proofread scan data transmission screen (FIG. 13) is pressed.

[0091] In FIG. 14, the AI ​​destination setting screen includes an address setting section 1401 , an ID setting section 1402 , a password setting section 1403 , and a save button 1404 .

[0092] The address setting unit 1401 sets an IP address, URL, etc. The ID setting unit 1402 and password setting unit 1403 respectively set an ID and password required for authentication when using the services of the large-scale language model server 2. When the CPU 201 accesses the large-scale language model server 2 set in the address setting unit 1401, the CPU 201 performs authentication using the information input to the ID setting unit 1402 and password setting unit 1403.

[0093] As described above, in this embodiment, the proofreading scan screen is displayed in step S402, the settings of the inspection items are accepted in steps S404 and S405, and the scanned image is sent to the large-scale language model server 2 for proofreading in steps S408 to S411. This allows the user to obtain highly accurate proofreading results with simple operations.

[0094] Furthermore, by setting the destination of the large-scale language model server 2 and the authentication ID and password required for accessing it in advance on the AI ​​destination setting screen of FIG. 14, the processing of step S410 can be carried out smoothly.

[0095] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the present embodiment to a system or device via a network or a recording medium, and having one or more processors in the computer of the system or device read and run the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0096] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention.

[0097] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention.

[0098] The disclosure of this embodiment includes the following configuration, method, and program. (Configuration 1) An image processing device that is communicatively connected to a review server that performs inference using a trained model, comprising: an acquisition means for acquiring image data read from a manuscript; a generation means for generating a prompt for reviewing the acquired image data; a review result request means for transmitting the acquired image data and the generated prompt to the review server, causing the review server to perform review of the acquired image data based on the generated prompt using the trained model and output the review results; and an output means for outputting the review results output from the review server to an output destination. (Configuration 2) The image processing device according to Configuration 1, wherein the proofreading server is a large-scale language model server. (Configuration 3) The image processing device according to configuration 1 or 2, wherein the output means displays image data showing the proofreading results on a preview screen. (Configuration 4) The image processing device according to any one of configurations 1 to 3, wherein the output means transmits image data showing the proofreading results to a predetermined destination. (Configuration 5) The image processing device according to any one of configurations 1 to 4, wherein the generated prompt is a sentence in a natural language. (Configuration 6) An image processing device described in any one of configurations 1 to 5, further comprising a first display means for displaying a plurality of test items for generating the prompt in a user-selectable manner, wherein the generation means generates the prompt according to an test item selected by the user from among the plurality of test items. (Configuration 7) The image processing device described in Configuration 6 is characterized in that the first display means further displays user-selectable proofreading and correction instruction items for instructing proofreading and correction, and when the user selects the proofreading and correction instruction item, the generation means further generates a correction instruction prompt instructing the user to correct the acquired image data, and the proofreading result request means further transmits the correction instruction prompt to the proofreading server, causes the proofreading server to generate image data in which errors have been corrected based on the proofreading results, and outputs the corrected image data. (Configuration 8) The image processing device according to configuration 7, wherein the output means switches between displaying the corrected image data and image data showing the proofreading results on a preview screen in response to a user instruction. (Configuration 9) The image processing device described in any one of configurations 6 to 8, characterized in that the first display means further displays a text instruction item selectable by the user to instruct the acquisition of a text list of the proofreading results, and when the text instruction item is selected by the user, the generation means further generates a text instruction prompt instructing the acquisition of a text list of the proofreading results, and the proofreading result request means further transmits the text instruction prompt to the proofreading server, causes the proofreading server to generate a text list of the proofreading results, and outputs the text list. (Configuration 10) An image processing device as described in Configuration 9, further comprising a conversion means for converting the text list into image data, wherein the output means switches between displaying the converted image data and image data showing the proofreading results on a preview screen in response to a user instruction. (Configuration 11) An image processing device described in any one of configurations 1 to 10, further comprising a destination setting UI for setting access settings to the review server, and wherein the review result request means communicates with the review server using the access settings to the review server set in the destination setting UI. (Configuration 12) The image processing device according to configuration 3, characterized in that the image data showing the proofreading results highlights the points pointed out by the proofreading and displays the proofreading items linked to the points pointed out. (Configuration 13) The image processing device according to configuration 4, further comprising second display means for displaying a plurality of destinations that can be set as the predetermined destination so that the user can select one of them. (Method 1) A control method for an image processing device that is communicatively connected to a review server that performs inference using a trained model, the control method comprising: an acquisition step for acquiring image data read from a manuscript; a generation step for generating a prompt for reviewing the acquired image data; a review result request step for sending the acquired image data and the generated prompt to the review server, causing the review server to perform review of the acquired image data based on the generated prompt using the trained model and output the review results; and an output step for outputting the review results output from the review server to an output destination. (Program 1) A program for causing a computer to function as each means of the image processing device described in any one of configurations 1 to 13. [Explanation of symbols]

[0099] 1. Image processing device 2 Large-scale language model server 201 CPU 204 eMMC 205 RAM 207 Network I / F 213 Display 214 Input Controller 215 Input section 219 Scanner I / F 220 Scanner

Claims

1. An image processing device communicably connected to a review server that performs inference using a trained model, an acquisition means for acquiring image data read from a document; generating means for generating a prompt for reviewing the acquired image data; a proofreading result request means for transmitting the acquired image data and the generated prompt to the proofreading server, causing the proofreading server to proofread the acquired image data based on the generated prompt using the trained model, and outputting the proofreading result; An image processing apparatus comprising: an output unit for outputting the proofreading results output from the proofreading server to an output destination.

2. The image processing device according to claim 1 , wherein the proofreading server is a large-scale language model server.

3. 2. The image processing apparatus according to claim 1, wherein the output means displays the image data showing the proofreading result on a preview screen.

4. 2. The image processing apparatus according to claim 1, wherein the output means transmits image data showing the proofreading results to a predetermined destination.

5. The image processing device of claim 1 , wherein the generated prompt is a sentence in a natural language.

6. The apparatus further includes a first display means for displaying a plurality of test items for generating the prompt in a user-selectable manner, 2. The image processing apparatus according to claim 1, wherein the generating means generates the prompt in accordance with an examination item selected by a user from among the plurality of examination items.

7. the first display means further displays review and correction instruction items for instructing review and correction in a manner selectable by the user; The image processing device described in claim 6, characterized in that when the proofreading and correction instruction item is selected by the user, the generation means further generates a correction instruction prompt that instructs the user to correct the acquired image data, and the proofreading result request means further sends the correction instruction prompt to the proofreading server, causing the proofreading server to generate image data in which errors have been corrected based on the proofreading result, and output the corrected image data.

8. 8. The image processing apparatus according to claim 7, wherein the output means switches between displaying the corrected image data and displaying the image data showing the proofreading result on a preview screen in response to a user instruction.

9. the first display means further displays a text conversion instruction item selectable by the user for instructing acquisition of a list of the proofreading results converted into text; The image processing device described in claim 6, characterized in that when the text conversion instruction item is selected by the user, the generation means further generates a text conversion instruction prompt that instructs the user to obtain a list of the proofreading results converted into text, and the proofreading result request means further transmits the text conversion instruction prompt to the proofreading server, causing the proofreading server to generate a list of the proofreading results converted into text and output the text list.

10. further comprising a conversion means for converting the text list into image data; 10. The image processing apparatus according to claim 9, wherein the output means switches between displaying the converted image data and image data showing the proofreading result on a preview screen in response to a user instruction.

11. further comprising a destination setting UI for setting access settings to the review server; 2. The image processing apparatus according to claim 1, wherein the proofreading result requesting means communicates with the proofreading server using an access setting for the proofreading server set in the destination setting UI.

12. 4. The image processing device according to claim 3, wherein the image data showing the proofreading result highlights the portions pointed out by the proofreading and displays proofreading items linked to the portions pointed out.

13. 5. The image processing apparatus according to claim 4, further comprising a second display unit for displaying a plurality of destinations that can be set as the predetermined destination so that the user can select one of the destinations.

14. A control method for an image processing device communicably connected to a review server that performs inference using a trained model, comprising: an acquisition step of acquiring image data read from a document; generating a prompt for reviewing the captured image data; a proofreading result request step of transmitting the acquired image data and the generated prompt to the proofreading server, causing the proofreading server to proofread the acquired image data based on the generated prompt using the trained model, and outputting the proofreading result; and an output step of outputting the proofreading results output from the proofreading server to an output destination.

15. A program for causing a computer to function as each of the means of the image processing apparatus according to claim 1.

Citation Information

Patent Citations

  • Distribution server device, distribution method, and distribution system

    JP2017011490A