Image processing device, image processing system, image processing method and control program

JP2025021073A5Pending Publication Date: 2026-05-27PFU LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PFU LTD
Filing Date
2023-07-31
Publication Date
2026-05-27

AI Technical Summary

Benefits of technology

【0011】 本発明によれば、画像処理装置、画像処理システム、画像処理方法及び制御プログラムは、スキャン画像から認識した文字を、利用者の作業負担を軽減させつつ、効率良く管理することが可能となる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an image processing device, an image processing system, an image processing method and a control program capable of efficiently managing characters recognized from a scan image while reducing a workload on a user.SOLUTION: Ab image processing device has: a first acquisition unit that acquires characters recognized from a scan image and an objective for recognizing the characters from the scan image received from a user; a generation unit that generates a command syntax including the characters and objective acquired by the first acquisition unit; a second acquisition unit that inputs a command syntax into a large-scale language model, and acquires output information output from the large-scale language model; and an output control unit that outputs information related to the output information.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image processing device, an image processing system, an image processing method, and a control program. [Background technology]

[0002] In recent years, image processing devices have been developed that recognize characters from scanned images of media such as forms in order to automate data entry tasks. In order to properly manage the characters contained in the media, such image processing devices need to identify, from among the recognized characters, characters that indicate the items to be managed (such as company names and amounts).

[0003] Prompt engineering is used to develop and optimize prompts for efficient use of language models (see non-patent document 1). [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] "Prompt Engineering Guide," Internet<URL:https: / / www.promptingguide.ai / jp> Summary of the Invention [Problem to be solved by the invention]

[0005] It is desirable for image processing devices to efficiently manage characters recognized from scanned images while reducing the workload of users.

[0006] An object of the present invention is to provide an image processing device, an image processing system, an image processing method, and a control program that are capable of efficiently managing characters recognized from a scanned image while reducing the workload of a user. [Means for solving the problem]

[0007] An image processing device according to one aspect of the present invention has a first acquisition unit that acquires characters recognized from a scanned image and an objective for recognizing characters from the scanned image received from a user, a generation unit that generates a command syntax including the characters and objective acquired by the first acquisition unit, a second acquisition unit that inputs the command syntax into a large-scale language model and acquires output information output from the large-scale language model, and an output control unit that outputs information related to the output information.

[0008] An image processing system according to one aspect of the present invention is an image processing system having a plurality of devices, wherein any one of the plurality of devices has a first acquisition unit that acquires characters recognized from a scanned image and an objective of recognizing the characters from the scanned image received from a user, any one of the plurality of devices has a generation unit that generates a command syntax including the characters and objective acquired by the first acquisition unit, any one of the plurality of devices has an acquisition unit that inputs the command syntax into a large-scale language model and acquires output information output from the large-scale language model, and any one of the plurality of devices has an output control unit that outputs information related to the output information.

[0009] An image processing method according to one aspect of the present invention acquires characters recognized from a scanned image and a purpose for recognizing the characters from the scanned image received from a user, generates a command syntax including the acquired characters and purpose, inputs the command syntax into a large-scale language model, acquires output information output from the large-scale language model, and outputs information related to the output information.

[0010] A control program according to one aspect of the present invention is a computer control program that causes a computer to acquire characters recognized from a scanned image and a purpose for recognizing the characters from the scanned image that is received from a user, generate a command syntax including the acquired characters and purpose, input the command syntax into a large-scale language model, acquire output information output from the large-scale language model, and output information related to the output information. Effect of the Invention

[0011] According to the present invention, the image processing device, image processing system, image processing method, and control program are capable of efficiently managing characters recognized from a scanned image while reducing the workload of a user. [Brief description of the drawings]

[0012] [Figure 1] 1 is a diagram showing a schematic configuration of an image processing system 1 according to an embodiment. [Diagram 2] 3 is a diagram showing a schematic configuration of a third storage device 310 and a third processing circuit 320. FIG. [Diagram 3] 11 is a schematic diagram for explaining the data structure of a template table; FIG. [Figure 4] 11 is a sequence illustrating an example of an operation of a file generation process. [Diagram 5] FIG. 5 is a schematic diagram showing an example of a scanned image 500. [Figure 6] 6A is a schematic diagram showing an example of a template 600, and FIG. 6B is a schematic diagram showing an example of a first directive syntax 610. FIG. [Figure 7] 7 is a schematic diagram showing an example of first output information 700. FIG. [Figure 8] 13 is a flowchart showing an example of an operation of a feedback process. [Figure 9] 9A is a schematic diagram showing an example of a second instruction syntax 900, and FIG. 9B is a schematic diagram showing an example of second output information 910. FIG. [Figure 10] 1 is a schematic diagram showing first output information 1000. FIG. [Figure 11] 13 is a sequence showing an example of a part of the operation of another file generation process. [Figure 12] 13 is a block diagram showing a schematic configuration of another third processing circuit 520. FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Hereinafter, an image processing device, an image processing system, an image processing method, and a control program according to one aspect of the present invention will be described with reference to the drawings. However, it should be noted that the technical scope of the present invention is not limited to the embodiments, but extends to the inventions described in the claims and their equivalents.

[0014] FIG. 1 is a diagram showing a schematic configuration of an image processing system 1 according to an embodiment.

[0015] As shown in FIG. 1, the image processing system 1 includes an image reading device 100, a client device 200, a server device 300, and an LLM (Large Language Models) server device 400. The image reading device 100, the client device 200, the server device 300, and the LLM server device 400 are connected to each other for communication via a network N. The network N is the Internet, an intranet, or the like. The image reading device 100, the client device 200, and the server device 300 are examples of image processing devices. The LLM server device 400 is a server in which a large-scale language model is stored. The large-scale language model is, for example, GPT-3, GPT-3.5, GPT-4, or the like used in ChatGPT (Chat Generative Pre-trained Transformer). The LLM server device 400 may be a server provided in a cloud network.

[0016] The image reading device 100 is, for example, a scanner device. Scanner devices include an ADF (Auto Document Feeder) type scanner device that captures an image of a medium, which is an original, while conveying the medium, and a flatbed type scanner device that captures an image of a fixed medium. The medium includes various types of media, such as receipts, invoices, delivery notes, and other documents, or business cards, contracts, instruction manuals, and other documents. The image reading device 100 may be a mobile phone, a smartphone, a tablet computer, a notebook personal computer, or the like that captures an image of a medium.

[0017] The image reading device 100 includes a first communication device 101, a first input device 102, a first display device 103, an imaging device 104, a first storage device 110, and a first processing circuit 120.

[0018] The first communication device 101 has a wired communication interface circuit for transmitting and receiving signals through a wired communication line according to a predetermined communication protocol such as a wired LAN (Local Area Network). The first communication device 101 transmits and receives image data and various information to and from the client device 200 according to instructions from the first processing circuit 120. The first communication device 101 may have an antenna for transmitting and receiving wireless signals, and a wireless communication interface circuit for transmitting and receiving signals through a wireless communication line according to a predetermined communication protocol such as a wireless LAN.

[0019] The first input device 102 has an input device such as a button and an interface circuit that acquires a signal from the input device, accepts an operation by a user, and outputs to the first processing circuit 120 a signal according to the operation by the user.

[0020] The first display device 103 has a display such as a liquid crystal display, an organic EL (Electro-Luminescence) display, or the like, and an interface circuit that outputs image data to the display.

[0021] The imaging device 104 generates a scan image by capturing an image of a medium. The imaging device 104 has a reduction optical system type imaging sensor equipped with imaging elements based on CCDs (Charge Coupled Devices) arranged one-dimensionally or two-dimensionally. Furthermore, the imaging device 104 has a light source for irradiating light, a lens for forming an image on the imaging elements, and an A / D converter for amplifying and analog-to-digital (A / D) converting an electric signal output from the imaging elements. In the imaging device 104, the imaging sensor captures an image of the front and / or back surface of the medium, generates an analog image signal, and outputs it. The A / D converter A / D converts the analog image signal to generate a digital scan image, and outputs it to the first processing circuit 120. Note that a reduction optical system type imaging sensor equipped with a CMOS (Complementary Metal Oxide Semiconductor) imaging element may be used instead of a CCD. Also, a unit-magnification optical system type line sensor equipped with an imaging element based on a CMOS or CCD may be used.

[0022] The first storage device 110 is an example of a storage unit, and includes a memory device such as a random access memory (RAM) or a read only memory (ROM), a fixed disk device such as a hard disk, or a portable storage device such as a flexible disk or an optical disk. The first storage device 110 also stores computer programs, databases, tables, and the like used for various processes of the image reading device 100. The computer programs may be installed in the first storage device 110 from a computer-readable portable recording medium using a known setup program or the like. The portable recording medium is, for example, a compact disk read only memory (CD-ROM) or a digital versatile disk read only memory (DVD-ROM). The computer programs may be distributed from a server or the like and installed in the first storage device 110.

[0023] The first processing circuit 120 operates based on a program stored in advance in the first storage device 110. The first processing circuit 120 is, for example, a CPU (Control Processing Unit). Note that the first processing circuit 120 may be a DSP (digital signal processor), an LSI (large scale integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programming Gate Array), or the like.

[0024] The first processing circuit 120 is connected to the first communication device 101, the first input device 102, the first display device 103, the imaging device 104, the first storage device 110, etc., and controls each of these components. The first processing circuit 120 controls data transmission and reception with the client device 200 via the first communication device 101, controls input of the first input device 102, controls display of the first display device 103, controls imaging of the imaging device 104, etc.

[0025] The client device 200 is a personal computer, a notebook personal computer, a tablet computer, a smartphone, or the like.

[0026] The client device 200 includes a second communication device 201 , a second input device 202 , a second display device 203 , a second storage device 210 , and a second processing circuit 220 .

[0027] The second communication device 201 has a wired communication interface circuit for transmitting and receiving signals through a wired communication line according to a predetermined communication protocol such as a wired LAN. The second communication device 201 transmits and receives image data and various information to and from the image reading device 100 and the server device 300 according to instructions from the second processing circuit 220. The second communication device 201 may have an antenna for transmitting and receiving wireless signals, and a wireless communication interface circuit for transmitting and receiving signals through a wireless communication line according to a predetermined communication protocol such as a wireless LAN.

[0028] The second input device 202 has an input device such as a keyboard or a mouse, and an interface circuit that acquires a signal from the input device, and outputs to the second processing circuit 220 a signal according to an operation by the user.

[0029] The second display device 203 has a display such as a liquid crystal display, an organic EL display, or the like, and an interface circuit that outputs image data to the display, and displays various types of information on the display according to instructions from the second processing circuit 220.

[0030] The second storage device 210 is an example of a storage unit, and includes a memory device such as a RAM or a ROM, a fixed disk device such as a hard disk, or a portable storage device such as a flexible disk or an optical disk. The second storage device 210 stores computer programs, databases, tables, and the like used for various processes of the client device 200. The computer programs may be installed in the second storage device 210 from a computer-readable portable recording medium such as a CD-ROM or a DVD-ROM using a known setup program or the like. The computer programs may also be distributed from a server or the like and installed in the second storage device 210.

[0031] The second processing circuit 220 operates based on a program previously stored in the second storage device 210. The second processing circuit 220 is, for example, a CPU. Note that the second processing circuit 220 may be, for example, a DSP, an LSI, an ASIC, an FPGA, or the like.

[0032] The second processing circuit 220 is connected to the second communication device 201, the second input device 202, the second display device 203, the second storage device 210, etc., and controls each of these components. The second processing circuit 220 controls data transmission and reception with the image reading device 100 and the server device 300 via the second communication device 201, controls input of the second input device 202, controls display of the second display device 203, etc.

[0033] The server device 300 is a server used by an administrator of the image processing system 1. The server device 300 may be a server provided in a cloud network.

[0034] The server device 300 includes a third communication device 301 , a third input device 302 , a third display device 303 , a third storage device 310 , and a third processing circuit 320 .

[0035] The third communication device 301 has a wired communication interface circuit for transmitting and receiving signals through a wired communication line according to a predetermined communication protocol such as a wired LAN. The third communication device 301 transmits and receives image data and various information to and from the image reading device 100 and the server device 300 according to instructions from the third processing circuit 320. The third communication device 301 may have an antenna for transmitting and receiving wireless signals, and a wireless communication interface circuit for transmitting and receiving signals through a wireless communication line according to a predetermined communication protocol such as a wireless LAN.

[0036] The third input device 302 has an input device such as a keyboard or a mouse, and an interface circuit that acquires a signal from the input device, and outputs to the third processing circuit 320 a signal according to an operation by the user.

[0037] The third display device 303 has a display such as a liquid crystal display, an organic electroluminescence display, or the like, and an interface circuit that outputs image data to the display, and displays various types of information on the display according to instructions from the third processing circuit 320.

[0038] The third storage device 310 is an example of a storage unit, and includes a memory device such as a RAM or a ROM, a fixed disk device such as a hard disk, or a portable storage device such as a flexible disk or an optical disk. The third storage device 310 stores computer programs, databases, tables, and the like used for various processes of the server device 300. The computer programs may be installed in the third storage device 310 from a computer-readable portable recording medium such as a CD-ROM or a DVD-ROM using a known setup program or the like. The computer programs may also be distributed from a server or the like and installed in the third storage device 310.

[0039] A template table and the like are stored in advance as data in the third storage device 310. Details of the template table will be described later.

[0040] The third processing circuit 320 operates based on a program previously stored in the third storage device 310. The third processing circuit 320 is, for example, a CPU. Note that, as the third processing circuit 320, a DSP, an LSI, an ASIC, an FPGA, or the like may be used.

[0041] The third processing circuit 320 is connected to the third communication device 301, the third input device 302, the third display device 303, the third storage device 310, etc., and controls each of these components. The third processing circuit 320 controls data transmission and reception with the client device 200 and the LLM server device 400 via the third communication device 301, controls input of the third input device 302, controls display of the third display device 303, etc.

[0042] FIG. 2 is a diagram showing a schematic configuration of the third storage device 310 and the third processing circuit 320. As shown in FIG.

[0043] 2, the third storage device 310 stores a first acquisition program 311, a generation program 312, a second acquisition program 313, an output control program 314, an update program 315, etc. Each of these programs is a functional module implemented by software that runs on a processor. The third processing circuit 320 reads each program stored in the third storage device 310 and operates according to the read programs, thereby functioning as a first acquisition unit 321, a generation unit 322, a second acquisition unit 323, an output control unit 324, and an update unit 325.

[0044] FIG. 3 is a schematic diagram for explaining the data structure of the template table.

[0045] As shown in FIG. 3, the template table stores a template of the first instruction syntax for each type of medium. The types of media are, for example, receipts, invoices, delivery notes, business cards, contracts, instructions, etc. The first instruction syntax is an example of an instruction syntax. The first instruction syntax is an input syntax to be input to the large-scale language model, and is a text syntax (text sentence) for expressing an instruction or question to the large-scale language model, that is, a question or command sentence presented for the large-scale language model to answer. The first instruction syntax is, for example, a prompt of ChatGPT. The template of the first instruction syntax is a template of the first instruction syntax according to the type of medium.

[0046] FIG. 4 is a sequence showing an example of the operation of a file generation process.

[0047] An example of the operation of the file generation process of the image processing system 1 will be described below with reference to the sequence shown in Fig. 4. Note that the operation described below is executed mainly by each processing circuit of each device in cooperation with each element of each device, based on a program stored in advance in each storage device of each device included in the image processing system 1.

[0048] First, the second processing circuit 220 of the client device 200 acquires a scan image (step S101). The second processing circuit 220 transmits a request signal requesting acquisition of the scan image to the image reading device 100 via the second communication device 201. When the first processing circuit 120 of the image reading device 100 receives the request signal from the client device 200 via the first communication device 101, the first processing circuit 120 transmits the scan image generated by the imaging device 104 to the client device 200 via the first communication device 101. The second processing circuit 220 acquires the scan image by receiving it from the image reading device 100 via the second communication device 201. The scan image is not limited to an image generated by the image reading device 100, and may be any image as long as it is an image of an image of a medium. The second processing circuit 220 may acquire the scan image by receiving a scan image input by a user via an interface circuit (not shown). The second processing circuit 220 may also acquire the scan image by reading a scan image stored in advance from the second storage device 210.

[0049] FIG. 5 is a schematic diagram showing an example of a scanned image 500. As shown in FIG.

[0050] In the scanned image 500 shown in FIG. 5, a receipt is captured as a medium. The scanned image 500 includes store information 501, summary information 502, purchase information 503, and additional information 504. The store information 501 includes the name, address, telephone number, business hours, etc. of the store that issued the receipt. The summary information 502 includes the date and time the receipt was issued, the receipt identification number, and the payment method used by the purchaser. The purchase information 503 includes the product name, unit price, quantity, and subtotal for each purchased item, the total amount, consumption tax, the amount including tax, the amount received, and the change. The additional information 504 includes information such as the fact that point members are being recruited.

[0051] Next, the second processing circuit 220 receives from the user the purpose of recognizing characters from the scan image and the type of medium included in the scan image (step S102). The second processing circuit 220 receives the designation of the purpose of recognizing characters from the scan image and the type of medium included in the scan image input by the user using the second input device 202. The background from which data is acquired is designated as the purpose of recognizing characters from the scan image. The purpose of recognizing characters from the scan image may be any form, such as "acquiring data for expense application," "acquiring items required for data input in invoice processing," "creating an address book," "organizing deliverables," "managing business cards," "extracting important matters for concluding a contract," etc. The second processing circuit 220 may display a plurality of predetermined candidates as the purpose of recognizing characters from the scan image and / or the type of medium selectable on the second display device 203 and accept the selection by the user.

[0052] Next, the second processing circuit 220 transmits the scanned image acquired in step S101 and information indicating the purpose and type of medium received in step S102 to the server device 300 via the second communication device 201 (step S103).

[0053] Meanwhile, the first acquisition unit 321 of the server device 300 receives the scanned image, the purpose of recognizing characters from the scanned image, and information indicating the type of medium included in the scanned image from the client device 200 via the third communication device 301 (step S104). As a result, the first acquisition unit 321 acquires the scanned image, the purpose of recognizing characters from the scanned image, and the type of medium included in the scanned image, which are received from the user. The first acquisition unit 321 may acquire the scanned image by receiving the scanned image input from the user via an interface circuit (not shown). The first acquisition unit 321 may also acquire the scanned image by reading out the scanned image stored in advance from the third storage device 310. The first acquisition unit 321 may also accept the purpose of recognizing characters from the scanned image and the designation of the type of medium included in the scanned image, which are input by the user using the third input device 302.

[0054] Next, the first acquisition unit 321 executes an OCR (Optical Character Recognition) process on the scanned image to recognize characters from the scanned image (step S105). As a result, the first acquisition unit 321 acquires the characters recognized from the scanned image. The first acquisition unit 321 also rearranges each character (character string) recognized from the scanned image in a predetermined order (for example, scanning sequentially from the character in the upper left to the right, and scanning the next line when a line break is detected). The first acquisition unit 321 may delete, from the characters recognized from the scanned image, preregistered wording that is unnecessary for form recognition, or wording other than preregistered wording that is unnecessary for form recognition.

[0055] Next, the generating unit 322 reads out a template corresponding to the type of medium acquired in step S104 from the template table (step S106).

[0056] FIG. 6A is a schematic diagram showing an example of a template 600. As shown in FIG.

[0057] The template 600 shown in FIG. 6A is a template corresponding to a receipt. The template 600 includes a command statement 601, a type specification statement 602, a purpose specification statement 603, an output format specification statement 604, an OCR result specification statement 605, and the like. The command statement 601 is a command for a large-scale language model, and, for example, a command is written in the command statement 601 to extract appropriate items from the OCR results for a specified type of medium according to a specified purpose and output them in a specified output format. The type specification statement 602 is a statement that specifies the type of medium contained in the scanned image. The type specification statement 602 may be written in advance as the type of medium corresponding to the template. The purpose specification statement 603 is a statement that specifies the purpose of recognizing characters from the scanned image. The output format specification statement 604 is a statement that specifies the output format. The output format is JSON format, CSV format, XML format, or the like. A predetermined format may be written in the output format specification statement 604. The OCR result specification statement 605 is a statement that specifies the characters recognized from the scanned image.

[0058] Next, the generating unit 322 generates a first instruction syntax including the characters, purpose, and type of medium acquired by the first acquiring unit 321 (step S107). The generating unit 322 generates the first instruction syntax according to the template read in step S106.

[0059] FIG. 6B is a schematic diagram showing an example of a first directive syntax 610.

[0060] The first instruction syntax 610 shown in FIG. 6(B) is a first instruction syntax generated so as to include characters recognized from the scanned image 500 shown in FIG. 5 according to the template 600 shown in FIG. 6(A). As shown in FIG. 6(B), the command statement 601 is not changed from the template 600. The type designation statement 602 and the purpose designation statement 603 respectively describe the purpose and type of medium acquired by the first acquisition unit 321. The output format designation statement 604 describes a specific (desired) output format. If the type of medium is previously described in the type designation statement 602 and / or the output format designation statement 604 in the template 600, the type designation statement 602 and / or the output format designation statement 604 do not need to be changed. The OCR result designation statement 605 describes the characters acquired by the first acquisition unit 321.

[0061] Next, the second acquisition unit 323 transmits the first instruction syntax generated by the generation unit 322 and request information requesting that the first instruction syntax be input into the large-scale language model to the LLM server device 400 via the third communication device 301 (step S108).

[0062] Meanwhile, the LLM server device 400 receives the first instruction syntax and request information from the server device 300. The LLM server device 400 inputs the received first instruction syntax into the large-scale language model stored in the LLM server device 400, and obtains first output information output from the large-scale language model (step S109). The first output information is an example of output information.

[0063] FIG. 7 is a schematic diagram showing an example of the first output information 700. As shown in FIG.

[0064] The first output information 700 shown in FIG. 7 is information output from the large-scale language model when the first instruction syntax 610 shown in FIG. 6(B) is input to the large-scale language model. The first output information 700 includes store information 701, summary information 702, purchase information 703, and the like. The store information 701 includes only the store name and address of the store in the store information 501 of the scanned image 500, and does not include the store's phone number or business hours. The summary information 702 includes only the receipt identification number in the summary information 502 of the scanned image 500, and does not include the date and time the receipt was issued or the payment method by the purchaser. The purchase information 703 includes only the product name, unit price, quantity, and subtotal for each purchased product, the total amount, the consumption tax, and the amount including tax in the purchase information 503 of the scanned image 500, and does not include the deposit amount or change. The first output information 700 does not include information corresponding to the auxiliary information 504 of the scanned image 500.

[0065] In this way, the first output information includes structured data from the OCR results. The first output information appropriately includes only characters that correspond to the purpose of recognizing characters from the scanned image (extracting data for expense claims) among the characters recognized from the scanned image. In addition, item names such as "store name" and "address" that were not included as characters in the scanned image even though they were not set in advance for each item are estimated and described in the first output information. Furthermore, product information that was merely a list of characters is organized and classified in a table format in the first output information. In addition, the first output information shown in FIG. 7 is described in JSON format, and therefore has a format that is easy to process in subsequent processing.

[0066] Next, the LLM server device 400 transmits the acquired first output information to the server device 300 (step S110).

[0067] Meanwhile, the second acquisition unit 323 of the server device 300 acquires the first output information by receiving it from the LLM server device 400 via the third communication device 301 (step S111). As a result, the second acquisition unit 323 inputs the first instruction syntax generated by the generation unit 322 to the large-scale language model, and acquires the first output information output from the large-scale language model. Note that the second acquisition unit 323 may convert or edit the acquired first output information so as to be compatible with a system used by the user.

[0068] Next, the output control unit 324 outputs the first output information acquired by the second acquisition unit 323 or the first output information converted or edited by the second acquisition unit 323 by transmitting it to the client device 200 via the third communication device 301 (step S112). The first output information acquired by the second acquisition unit 323 or the first output information converted or edited by the second acquisition unit 323 is an example of information related to the output information. Note that the output control unit 324 may output the first output information acquired by the second acquisition unit 323 or the first output information converted or edited by the second acquisition unit 323 by displaying it on the third display device 303.

[0069] On the other hand, the second processing circuit 220 of the client device 200 receives the first output information from the server device 300 via the second communication device 201 (Step S113).

[0070] Next, the second processing circuit 220 displays the received first output information on the second display device 203 (step S114).

[0071] Next, the second processing circuit 220 receives a confirmation result for the first output information from the user (step S115). The second processing circuit 220 receives a designation of the confirmation result input by the user using the second input device 202. The confirmation result includes, for example, an OK instruction indicating that the first output information is appropriate, an NG instruction indicating that the first output information is inappropriate, and an end instruction instructing to complete the file generation process. If the confirmation result indicates an NG instruction, the second processing circuit 220 further receives, as the confirmation result, correction items to be corrected for the first output information from the user. For example, unnecessary items among the items included in the first output information and / or missing items that were not included in the first output information are designated as the correction items.

[0072] Next, the second processing circuit 220 transmits information indicating a confirmation result for the first output information received from the user to the server device 300 via the second communication device 201 (step S116).

[0073] Meanwhile, the first acquisition unit 321 of the server device 300 receives information indicating the confirmation result from the client device 200 via the third communication device 301 (step S117). As a result, the first acquisition unit 321 acquires the confirmation result for the first output information accepted from the user. Note that the first acquisition unit 321 may receive and acquire the confirmation result input by the user using the third input device 302.

[0074] Next, the third processing circuit 320 executes a feedback process based on the confirmation result acquired by the first acquisition unit 321 (step S118). With the above, the file generation process ends.

[0075] 8 is a flowchart showing an example of the operation of the feedback process. The feedback process is executed in step S118 of FIG.

[0076] First, the first acquisition unit 321 determines whether or not the acquired confirmation result indicates an NG instruction (Step S201).

[0077] If the check result indicates an NG instruction, the generating unit 322 corrects the first instruction syntax based on the correction item included in the check result acquired by the first acquiring unit 321 (step S202), and returns the process to step S108 in FIG. 4. The generating unit 322 corrects the first instruction syntax generated in step S107 in FIG. 4 so as to add to the command statement 601 of the first instruction syntax a statement that the unnecessary item included in the check result is not extracted and / or that the missing item included in the check result is extracted. For example, if the check result includes the unnecessary item "A", a sentence "A is an unnecessary item, so please do not include it in the output" is added to the command statement of the first instruction syntax. Also, if the check result does not include the missing item "B", a sentence "B is a necessary item, so please be sure to output it" is added to the command statement of the first instruction syntax. 4, the second acquisition unit 323 transmits the corrected first instruction syntax and request information requesting that the first instruction syntax be input to the large-scale language model to the LLM server device 400, and inputs the first instruction syntax to the large-scale language model. This enables the image processing system 1 to more accurately extract necessary characters from the scanned image based on the user's feedback.

[0078] On the other hand, if the confirmation result does not indicate an NG instruction, the first obtaining unit 321 determines whether the confirmation result indicates an OK instruction (step S203).

[0079] If the confirmation result indicates an OK instruction, the first acquisition unit 321 determines whether the generation unit 322 has corrected the first instruction syntax for the currently processed scanned image based on the correction items included in the confirmation result (step S204). If the generation unit 322 has not corrected the first instruction syntax based on the correction items included in the confirmation result, the first acquisition unit 321 proceeds to step S206 without performing any particular process.

[0080] On the other hand, if the generating unit 322 corrects the first instruction syntax based on the correction items included in the confirmation result, the updating unit 325 updates the template of the first instruction syntax stored in the third storage device 310 based on the corrected first instruction syntax (step S205). The updating unit 325 updates the template of the first instruction syntax by replacing the command statement 601 of the corresponding template, i.e., the used template, with the command statement corrected (added) by the generating unit 322. In this way, the updating unit 325 updates the template of the first instruction syntax based on the confirmation result acquired by the first acquiring unit 321. As a result, the first instruction syntax is generated according to the updated template from the next time onwards, so that the image processing system 1 can extract appropriate characters from the scanned image before receiving feedback from the user. Therefore, the image processing system 1 can extract necessary characters from the scanned image with higher accuracy while saving the user's effort.

[0081] Next, the generation unit 322 generates a second command syntax requesting the large-scale language model to suggest a file name (step S206), and returns the process to step S108 in Fig. 4. The generation unit 322 generates the second command syntax based on the first output information acquired by the second acquisition unit 323. As in the case where the first command syntax is generated, the server device 300 may store a template of the second command syntax in the third storage device 310 in advance, and the generation unit 322 may generate the second command syntax according to the template stored in the third storage device 310.

[0082] FIG. 9A is a schematic diagram showing an example of a second directive syntax 900.

[0083] The second instruction syntax 900 shown in FIG. 9(A) is a second instruction syntax generated based on the first output information 700 shown in FIG. 7. As shown in FIG. 9(A), the second instruction syntax 900 includes a command statement 901, a file name example specification statement 902, an output result specification statement 903, and the like. The command statement 901 is a command for the large-scale language model, and as the command statement 901, for example, a command to determine and output an appropriate file name from the first output information and a file name example is described. In the file name example specification statement 902, an example of a file name (format example) is described. The command statement 901 and the file name example specification statement 902 may be described in a template of the second instruction syntax. In the output result specification statement 903, the first output information acquired by the second acquisition unit 323 is described.

[0084] In this case, in step S108 of Fig. 4, the second acquisition unit 323 transmits the newly generated second instruction syntax and request information requesting that the second instruction syntax be input into the large-scale language model to the LLM server device 400, and inputs them into the large-scale language model. In step S109 of Fig. 4, the LLM server device 400 receives the second instruction syntax and the request information from the server device 300. The LLM server device 400 inputs the received second instruction syntax into the large-scale language model stored in the LLM server device 400, and acquires second output information output from the large-scale language model.

[0085] FIG. 9B is a schematic diagram showing an example of the second output information 910.

[0086] The second output information 910 shown in Fig. 9(B) is the second output information output from the large-scale language model when the second instruction syntax 900 shown in Fig. 9(A) is input to the large-scale language model. The second output information 910 includes a file name proposed by the large-scale language model. The file name proposed by the large-scale language model is proposed according to the file name specified in the second instruction syntax.

[0087] In step S110 of FIG. 4, the LLM server device 400 transmits the acquired second output information to the server device 300. In step S111 of FIG. 4, the second acquisition unit 323 of the server device 300 acquires the second output information by receiving it from the LLM server device 400 via the third communication device 301. In step S112 of FIG. 4, the output control unit 324 outputs the second output information acquired by the second acquisition unit 323 or the second output information converted or edited by the second acquisition unit 323 by transmitting it to the client device 200 via the third communication device 301. The second output information acquired by the second acquisition unit 323 or the second output information converted or edited by the second acquisition unit 323 is an example of information related to the second output information. In this way, the second acquisition unit 323 inputs the second instruction syntax to the large-scale language model, acquires the second output information output from the large-scale language model, and the output control unit 324 outputs information related to the second output information. This allows the image processing system 1 to suggest to the user a file name for appropriately managing the first output information, thereby improving user convenience.

[0088] In step S113 of Fig. 4, the second processing circuit 220 of the client device 200 receives second output information from the server device 300 via the second communication device 201. In step S114 of Fig. 4, the second processing circuit 220 displays the received second output information on the second display device 203. In step S114 of Fig. 4, the second processing circuit 220 accepts an end instruction from the user as a confirmation result for the second output information, the end instruction instructing to complete the file generation process.

[0089] Returning to FIG. 8, if the confirmation result in step S203 does not indicate an OK instruction, that is, if the confirmation result is neither an NG instruction nor an OK instruction, the first acquisition unit 321 determines that the confirmation result is an end instruction. In this case, the generation unit 322 generates a file including the first output information for which the confirmation result by the user was an OK instruction, and stores the file in the third storage device 310 with a file name indicated in the second output information. Then, the generation unit 322 outputs the generated file by transmitting it to the client device 200 via the third communication device 301 (step S207), and ends a series of steps. Note that the second processing circuit 220 of the client device 200 may receive a correction instruction for the first output information and / or the second output information from the user using the second input device 202. In this case, the generation unit 322 generates a file including the first output information corrected according to the correction instruction by the user, and / or stores the file in the third storage device 310 with a file name indicated in the second output information corrected according to the correction instruction by the user.

[0090] In step S104 of FIG. 4, the first acquisition unit 321 may not acquire the type of medium included in the scanned image received from the user. In that case, for example, the first acquisition unit 321 identifies the type of medium included in the scanned image based on the characters acquired by recognizing the scanned image in step S105. For example, the server device 300 stores one or more keywords corresponding to each type of medium in the third storage device 310 in advance. For example, if the type of medium is an invoice, "invoice", "fee", etc. are set as keywords. If the type of medium is a delivery note, "delivery note", "goods", etc. are set as keywords. If the type of medium is a receipt, "receipt", "product name", etc. are set as keywords. The first acquisition unit 321 refers to the keywords stored in the third storage device 310 and identifies the type of medium corresponding to the keyword that matches the acquired characters as the type of medium included in the scanned image.

[0091] Alternatively, the first acquisition unit 321 extracts ruled lines from the scanned image using a known image processing technique. The server device 300 stores images in which ruled lines corresponding to each type of medium are arranged in the third storage device 310 in advance for each type of medium. The first acquisition unit 321 calculates the similarity between an image in which only ruled lines extracted from the scanned image are arranged and each image stored in the third storage device 310, and specifies the type corresponding to the image with the highest similarity as the type of image included in the scanned image. The similarity is, for example, a normalized cross-correlation value.

[0092] Alternatively, the first acquisition unit 321 may identify the type of medium by a learning model that has been pre-trained to output the type of medium when a character (character string) or an image is input. The learning model is pre-trained using a set of images of characters included in various types of media or various types of media captured and the types of each medium by supervised learning such as a neural network or a support vector machine, and is stored in advance in the third storage device 310. The first acquisition unit 321 inputs the characters recognized from the scanned image or the scanned image into the learning model, and identifies the type of medium output from the learning model as the type of medium included in the scanned image.

[0093] As a result, the user does not need to specify the type of medium included in the scanned image, and the image processing system 1 can reduce the user's effort and improve user convenience.

[0094] Also, the first acquisition unit 321 may not acquire the type of medium. In that case, the generation unit 322 generates a first instruction syntax that does not include the type of medium by using a template that is determined in advance regardless of the type of medium. When using a first instruction syntax that does not include the type of medium, the accuracy of extracting the necessary items by the large-scale language model may be lower than when using a first instruction syntax that includes the type of medium. However, even in this case, the large-scale language model can extract the necessary items with high accuracy because the first instruction syntax includes the purpose of recognizing characters from the scanned image. That is, when the type of medium is not included in the first instruction syntax, the image processing system 1 can extract the necessary items from the scanned image while saving the user's effort, and when the type of medium is included in the first instruction syntax, the image processing system 1 can extract the necessary items from the scanned image with high accuracy.

[0095] Also, if the scanned image is born digital, the process of step S105 may be omitted. In this case, the first acquisition unit 321 acquires the characters included in the scanned image, the purpose of using the characters included in the scanned image received from the user, and the type of medium included in the scanned image. The generation unit 322 generates a first instruction syntax including the characters, purpose, and type of medium acquired by the first acquisition unit 321.

[0096] Furthermore, the process of step S106 in Fig. 4 may be omitted, and the generation unit 322 may generate the first instruction syntax without using a template. The processes of steps S115 to S118 in Fig. 4 may be omitted, and the first acquisition unit 321 may not accept a confirmation result from the user. The processes of steps S201 to S202, steps S204 to S205, and / or step S206 in Fig. 8 may be omitted. When the processes of steps S204 to S206 are omitted, the process of step S203 is also omitted.

[0097] The technical significance of including the purpose of recognizing characters from a scanned image in the first instruction syntax will be explained below.

[0098] FIG. 10 is a schematic diagram showing first output information 1000 from a large-scale language model when a first instruction syntax that does not include an objective of recognizing characters from a scanned image is input to the large-scale language model.

[0099] The first output information 1000 shown in FIG. 10 is information output from the large-scale language model when the first imperative syntax 610 shown in FIG. 6B that does not include the object specification statement 603 is input to the large-scale language model. The first output information 1000 includes store information 1001, summary information 1002, purchase information 1003, and additional information 1004. The store information 1001 includes all of the store information 501 in the scanned image 500 (store name, address, telephone number, and business hours). The summary information 1002 does not include the date and time of receipt issuance among the summary information 502 in the scanned image 500, but includes the receipt identification number and the payment method by the purchaser. The purchase information 1003 includes all of the purchase information 503 in the scanned image 500 (product name, unit price, quantity, and subtotal for each purchased product, the total amount, consumption tax, amount including tax, deposit amount, and change). In this way, compared to the first output information 700 when a purpose is specified, the first output information 1000 when a purpose is not specified contains many items that are not necessary for an expense claim. Therefore, the user needs to specify many unnecessary items and obtain the first output information from the large-scale language model again, or needs to modify the first output information to delete many unnecessary items from the first output information.

[0100] The image processing system 1 generates a first instruction syntax so as to include a purpose of recognizing characters from a scanned image, and inputs the first instruction syntax to the large-scale language model. This allows the image processing system 1 to suppress the inclusion of items that do not match the purpose in the first output information output from the large-scale language model, thereby saving the user time and reducing the processing load and processing time required for file generation processing.

[0101] As described above in detail, the image processing system 1 inputs the first instruction syntax including the purpose of recognizing the characters from the scanned image, in addition to the characters recognized from the scanned image of the medium, to the large-scale language model to obtain the items to be extracted. This allows the image processing system 1 to extract items matching the purpose from the characters recognized from the scanned image with high accuracy and provide them to the user. Therefore, the image processing system 1 can efficiently manage the characters recognized from the scanned image while reducing the workload of the user.

[0102] In addition, by using a large-scale language model, the image processing system 1 does not need to set the items to be extracted in advance. Therefore, the image processing system 1 does not need to set the position where the items to be extracted are located in the medium, and the convenience of the user can be improved. In addition, the image processing system 1 can appropriately extract the items to be extracted even from a medium in which the position where the items to be extracted are located is not fixed. In general, the purpose of recognizing characters from a scanned image differs depending on the business content, and the items to be extracted differ depending on the purpose of recognizing characters from a scanned image. Therefore, in data entry work, it is necessary to set the items to be extracted in advance for each business content or each purpose of recognizing characters from a scanned image. The image processing system 1 can accurately extract the items to be extracted without setting them in advance by inputting the purpose of recognizing characters from a scanned image into the large-scale language model.

[0103] FIG. 11 is a sequence showing an example of a part of the operation of a file generation process according to another embodiment.

[0104] The sequence shown in Fig. 11 is executed in place of steps S102 to S105 of the sequence shown in Fig. 4. In the sequence shown in Fig. 11, the image processing system 1 acquires the purpose of recognizing characters from a scanned image and the type of medium included in the scanned image by using a large-scale language model.

[0105] First, the second processing circuit 220 of the client device 200 transmits the scanned image acquired in step S101 to the server device 300 via the second communication device 201 (step S301).

[0106] On the other hand, the first acquisition unit 321 of the server device 300 receives and acquires the scanned image from the client device 200 via the third communication device 301 (Step S302).

[0107] Next, the first obtaining unit 321 performs OCR processing on the scanned image to recognize characters from the scanned image, and obtains the recognized characters from the scanned image (Step S303).

[0108] Next, the generating unit 322 generates a third command syntax requesting the large-scale language model to identify the type of medium included in the scanned image (step S304). The generating unit 322 generates the third command syntax including the characters acquired by the first acquiring unit 321 in the same manner as the first command syntax. However, the command statement in the third command syntax describes a command to identify and output the type of medium from the specified OCR result. In addition, the third command syntax does not include a type specification statement and a purpose specification statement. As in the case where the first command syntax is generated, the server device 300 may store a template of the third command syntax in the third storage device 310 in advance, and the generating unit 322 may generate the third command syntax according to the template stored in the third storage device 310.

[0109] Next, the second acquisition unit 323 transmits the third instruction syntax generated by the generation unit 322 and request information requesting that the third instruction syntax be input into the large-scale language model to the LLM server device 400 via the third communication device 301 (step S305).

[0110] Meanwhile, the LLM server device 400 receives the third instruction syntax and the request information from the server device 300. The LLM server device 400 inputs the received third instruction syntax into the large-scale language model stored in the LLM server device 400, and obtains third output information output from the large-scale language model (step S306). This third output information includes the type of medium identified by the large-scale language model.

[0111] Next, the LLM server device 400 transmits the acquired third output information to the server device 300 (step S307).

[0112] Meanwhile, the second acquisition unit 323 of the server device 300 acquires the type of medium by receiving the third output information from the LLM server device 400 via the third communication device 301 (step S308). As a result, the second acquisition unit 323 inputs the third instruction syntax generated by the generation unit 322 to the large-scale language model, and acquires the third output information output from the large-scale language model.

[0113] Next, the generating unit 322 generates a fourth command syntax requesting the large-scale language model to identify one or more candidates for the purpose of recognizing characters from a scanned image (step S309). The generating unit 322 generates the fourth command syntax including the characters acquired by the first acquiring unit 321 in the same manner as the first command syntax. However, the command statement in the fourth command syntax describes a command to identify and output candidates for the purpose of recognizing characters from a scanned image from the OCR result for the specified type of medium. In addition, the fourth command syntax includes a type specification statement that specifies the type of medium acquired by the second acquiring unit 323. On the other hand, the fourth command syntax does not include a purpose specification statement. As in the case where the first command syntax is generated, the server device 300 may store a template of the fourth command syntax in the third storage device 310 in advance, and the generating unit 322 may generate the fourth command syntax according to the template stored in the third storage device 310.

[0114] Next, the second acquisition unit 323 transmits the fourth instruction syntax generated by the generation unit 322 and request information requesting that the fourth instruction syntax be input into the large-scale language model to the LLM server device 400 via the third communication device 301 (step S310).

[0115] Meanwhile, the LLM server device 400 receives the fourth instruction syntax and the request information from the server device 300. The LLM server device 400 inputs the received fourth instruction syntax into the large-scale language model stored in the LLM server device 400, and obtains fourth output information output from the large-scale language model (step S311). This fourth output information includes one or more candidates for the purpose of recognizing characters from a scanned image, which are identified by the large-scale language model.

[0116] Next, the LLM server device 400 transmits the acquired fourth output information to the server device 300 (step S312).

[0117] Meanwhile, the second acquisition unit 323 of the server device 300 receives the fourth output information from the LLM server device 400 via the third communication device 301, and extracts candidates for the purpose of recognizing characters from the scanned image from the received fourth output information (step S313). As a result, the second acquisition unit 323 inputs the fourth instruction syntax generated by the generation unit 322 to the large-scale language model, and acquires the fourth output information output from the large-scale language model.

[0118] Next, the output control unit 324 transmits information indicating candidates for the purpose of recognizing characters from the scanned image, extracted by the second acquisition unit 323, to the client device 200 via the third communication device 301 (step S314).

[0119] Meanwhile, the second processing circuit 220 of the client device 200 receives information indicating candidates for the purpose of recognizing characters from the scanned image from the server device 300 via the second communication device 201 (step S315).

[0120] Next, the second processing circuit 220 displays, on the second display device 203, selectable candidates for recognizing characters from the scanned image (step S316).

[0121] Next, the second processing circuit 220 accepts a selection by the user of an objective of recognizing characters from a scanned image from among the candidate objectives displayed in a selectable manner (step S317). The second processing circuit 220 accepts the selection of an objective selected by the user using the second input device 202.

[0122] Next, the second processing circuit 220 transmits information indicating the purpose selected by the user to the server device 300 via the second communication device 201 (step S318).

[0123] Meanwhile, the first acquisition unit 321 of the server device 300 receives information indicating the purpose selected by the user from the client device 200 via the third communication device 301 (step S319). As a result, the first acquisition unit 321 acquires the purpose of recognizing characters from a scanned image accepted from the user. Thereafter, the processes of steps S106 to S118 in FIG. 4 are executed. In step S106, the generation unit 322 reads out a template corresponding to the type of medium acquired in step S308 from the template table. In step S107, the generation unit 322 generates a first instruction syntax including the characters acquired in step S303, the purpose of recognizing characters from a scanned image acquired in step S319, and the type of medium acquired in step S303.

[0124] In addition, when a file generation process is performed for the same purpose and the same type of medium, in the second or subsequent file generation processes, steps S304 to S319 may be omitted, and the instruction syntax may be generated using the purpose and type of medium obtained or modified the first time.

[0125] As described above in detail, the image processing system 1 is capable of efficiently managing characters recognized from scanned images while reducing the workload of the user, even when the purpose of character recognition and type of medium are obtained using a large-scale language model.

[0126] In particular, image processing system 1 allows a user, even if he or she is unfamiliar with data entry work, to appropriately set the purpose of character recognition and the type of medium, thereby making it possible to accurately extract necessary items from the characters recognized from the scanned image.

[0127] FIG. 12 is a block diagram showing a schematic configuration of a third processing circuit 520 in a server device according to another embodiment.

[0128] The third processing circuit 520 is used in place of the third processing circuit 320, and executes file generation processing, etc. The third processing circuit 320 has a first acquisition circuit 521, a generation circuit 522, a second acquisition circuit 523, an output control circuit 524, an update circuit 525, etc.

[0129] The first acquisition circuit 521 is an example of a first acquisition unit, and has the same function as the first acquisition unit 321. The first acquisition circuit 521 receives a scanned image from the third communication device 301, as well as information indicating the purpose of recognizing characters from the scanned image and the type of medium included in the scanned image, and stores the information in the third storage device 310. The first acquisition circuit 521 also recognizes characters from the scanned image and stores the characters in the third storage device 310. The first acquisition circuit 521 also receives a confirmation result for the first output information accepted from the user from the third communication device 301, and stores the confirmation result in the third storage device 310.

[0130] The generation circuit 522 is an example of a generation unit, and has the same function as the generation unit 322. The generation circuit 522 reads out from the third storage device 310 the characters recognized from the scanned image, the purpose of recognizing the characters from the scanned image, the type of medium included in the scanned image, the first output information, and / or the confirmation result for the first output information. The generation circuit 522 generates a first instruction syntax and / or a second instruction syntax based on each of the read information, and stores it in the third storage device 310.

[0131] The second acquisition circuit 523 is an example of a second acquisition unit, and has the same function as the second acquisition unit 323. The second acquisition circuit 523 reads the first instruction syntax and / or the second instruction syntax from the third storage device 310, and transmits it to the LLM server device 400 via the third communication device 301. The second acquisition circuit 523 receives the first output information and / or the second output information, etc. from the LLM server device 400 via the third communication device 301, and stores it in the third storage device 310.

[0132] The output control circuit 524 is an example of an output control unit, and has the same function as the output control unit 324. The output control circuit 524 reads out the first output information and / or the second output information, etc. from the third storage device 310, and outputs them to the third communication device 301.

[0133] The update circuit 525 is an example of an update unit, and has the same function as the update unit 325. The update circuit 525 reads, from the third storage device 310, a confirmation result for the first output information accepted from the user and a template of the first instruction syntax, updates the template based on the read confirmation result, and stores it in the third storage device 310.

[0134] As described above in detail, even when the server device uses the third processing circuit 520, it is possible to efficiently manage characters recognized from a scanned image while reducing the workload of the user.

[0135] Although the preferred embodiments have been described above, the embodiments are not limited to these. For example, the large-scale language model may be stored in advance in the third storage device 310 of the server device 300, instead of being stored in the LLM server device 400. In this case, the processes of steps S108 to S111 in Fig. 4 are omitted, and the second acquisition unit 323 inputs the first instruction syntax into the large-scale language model stored in the third storage device 310, and acquires the first output information output from the large-scale language model.

[0136] Furthermore, the processes of steps S105 to S108, S111, and S118 may be executed by the client device 200, not by the server device 300. In this case, the second storage device 210 of the client device 200 stores the programs, databases, tables, information, and the like stored in the third storage device 310. The second processing circuit 220 reads the programs stored in the second storage device 210 and operates according to the read programs, thereby functioning as a first acquisition unit, a generation unit, a second acquisition unit, an output control unit, and an update unit. The first acquisition unit, the generation unit, the second acquisition unit, the output control unit, and the update unit of the second processing circuit 220 have the same functions as the first acquisition unit 321, the generation unit 322, the second acquisition unit 323, the output control unit 324, and the update unit 325 of the third processing circuit 320.

[0137] The processing of steps S103 to S104 is omitted, and the first acquisition unit of the client device 200 acquires the purpose of recognizing characters from the scanned image and the type of medium included in the scanned image, which are accepted from the user in step S102. In step S108, the second acquisition unit of the client device 200 transmits each command syntax and request information to the LLM server device 400 via the second communication device 201. In step S111, the second acquisition unit of the client device 200 receives each output information from the LLM server device 400 via the second communication device 201. The processing of steps S112 to S113 is omitted. In step S114, the output control unit of the client device 200 outputs each output information by displaying it on the second display device 203. The processing of steps S116 to S117 is omitted, and in step S207 of FIG. 6, the generation unit of the client device 200 outputs the generated file by displaying it on the second display device 203.

[0138] Similarly, in the file generation process shown in FIG. 11, each process executed by the server device 300 may be executed by the client device 200.

[0139] In these cases, the large-scale language model may be stored in advance in the second storage device 210 of the client device 200, rather than being stored in the LLM server device 400. In this case, the second acquisition unit of the client device 200 inputs each instruction syntax into the large-scale language model stored in the second storage device 210, and acquires each piece of output information output from the large-scale language model.

[0140] Furthermore, the processes of steps S101 to S102, S105 to S108, S111, S114 to S115, and S118 may be executed by the image reading device 100, instead of the server device 300 and the client device 200. In this case, the first storage device 110 of the image reading device 100 stores the programs, databases, tables, information, and the like stored in the third storage device 310. The first processing circuit 120 reads the programs stored in the first storage device 110 and operates according to the read programs, thereby functioning as a first acquisition unit, a generation unit, a second acquisition unit, an output control unit, and an update unit. The first acquisition unit, the generation unit, the second acquisition unit, the output control unit, and the update unit of the second processing circuit 220 have the same functions as the first acquisition unit 321, the generation unit 322, the second acquisition unit 323, the output control unit 324, and the update unit 325 of the third processing circuit 320.

[0141] In step S101, the first acquisition unit of the image reading device 100 acquires the scanned image generated by the imaging device 104. In step S102, the first acquisition unit of the image reading device 100 acquires the purpose of recognizing characters from the scanned image and the type of medium included in the scanned image, which are accepted from the user using the first input device 102. The processes of steps S103 to S104 are omitted. In step S108, the second acquisition unit of the image reading device 100 transmits each command syntax and request information to the LLM server device 400 via the first communication device 101. In step S111, the second acquisition unit of the image reading device 100 receives each output information from the LLM server device 400 via the first communication device 101. The processes of steps S112 to S113 are omitted. In step S114, the output control unit of the image reading device 100 outputs each output information by displaying it on the first display device 103. The processes in steps S116 to S117 are omitted, and in step S207 in FIG. 6, the generating unit of the image reading device 100 outputs the generated file by displaying it on the first display device 103.

[0142] Similarly, in the file generation process shown in FIG. 11, each process executed by the client device 200 and the server device 300 may be executed by the image reading device 100.

[0143] In these cases, the large-scale language model may be stored in advance in the first storage device 110 of the image reading device 100, instead of being stored in the LLM server device 400. In this case, the second acquisition unit of the image reading device 100 inputs each instruction syntax into the large-scale language model stored in the first storage device 110, and acquires each piece of output information output from the large-scale language model.

[0144] 4 or 11 can be arbitrarily set as a combination of devices that execute each process of the file generation process shown in Fig. 4 or 11. That is, it can be arbitrarily set as to whether the first acquisition unit, the generation unit, the second acquisition unit, the output control unit, and the update unit are arranged in the server device 300, the client device 200, or the image reading device 100. [Explanation of symbols]

[0145] 1 Image processing system, 100 Image reading device, 200 Client device, 300 Server device, 310 Third storage device, 321 First acquisition unit, 322 Generation unit, 323 Second acquisition unit, 324 Output control unit, 325 Update unit

Claims

1. A computer control program, The objective is to recognize characters from scanned images. Generate an instructional syntax that includes the acquired objective, The aforementioned directive syntax is input into a large-scale language model, and the output information output from the large-scale language model is obtained. A control program characterized by causing the computer to perform the following action.

2. The control program according to claim 1, further causing the computer to output information relating to the output information.

3. The control program according to claim 1 or 2, wherein the generation of the instruction syntax further includes characters recognized from the scanned image.

4. Further cause the computer to obtain the type of medium contained in the scanned image, The control program according to claim 1 or 2, wherein the generation of the instruction syntax further includes the type of medium acquired.

5. Obtain the confirmation result for the output information, Based on the verification results obtained, the instruction syntax is modified. The control program according to claim 1 or 2, further causing the computer to input the modified instruction syntax into a large-scale language model.

6. The computer is further instructed to store a template of instruction syntax, In the above generation, the instruction syntax is generated according to the template, The verification result for the output information is obtained, The control program according to claim 1 or 2, further causing the computer to update the template based on the acquired verification results.

7. It stores template instruction sentences for each type of medium, The control program according to claim 1 or 2, further causing the computer to obtain the type of medium contained in the scanned image based on the characters in the scanned image.

8. Based on the output information, a second instruction syntax is generated that requests a file name suggestion, The control program according to claim 1 or 2, further causing the computer to input the second instruction syntax into a large-scale language model and to obtain the second output information output from the large-scale language model.

9. An image processing system having multiple devices, Any of the aforementioned multiple devices has a first acquisition unit that acquires the purpose of recognizing characters from a scanned image, Any of the above-mentioned devices has a generation unit that generates an instruction syntax including the purpose acquired by the first acquisition unit, Any of the above-mentioned devices has a second acquisition unit that inputs the instruction syntax into a large-scale language model and acquires output information output from the large-scale language model. An image processing system characterized by the following:

10. The objective is to recognize characters from a scanned image, Generate an instructional syntax that includes the acquired objective, The aforementioned directive syntax is input into a large-scale language model, and the output information output from the large-scale language model is obtained. An image processing method characterized by the following:

11. A first acquisition unit for acquiring the purpose of recognizing characters from a scanned image, A generation unit that generates an instruction syntax including the acquired objective, A second acquisition unit inputs the aforementioned instruction syntax into a large-scale language model and acquires output information output from the large-scale language model, An image processing apparatus characterized by having