Image forming system, image forming method, and image forming program
The image forming system uses a large-scale language model to automatically correct printed materials without user memorization, addressing the limitations of existing systems by enabling flexible and efficient image data modification.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- KONICA MINOLTA INC
- Filing Date
- 2024-12-10
- Publication Date
- 2026-06-22
AI Technical Summary
Existing image correction systems require users to memorize correction rules and are limited in the types of modifications they can perform, making them cumbersome and inflexible.
An image forming system that includes an image data acquisition unit, recognition unit, correction instruction acquisition unit, and a large-scale language model to automatically generate corrected image data without requiring users to remember rules, allowing for various modifications.
Enables flexible and efficient correction of printed materials by automatically generating corrected image data using natural language processing, reducing the need for user memorization and accommodating diverse modifications.
Smart Images

Figure 2026100877000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image forming system, an image forming method, and an image forming program.
Background Art
[0002] Generally, information is input using a personal computer and printed to create paper-based materials.
[0003] When correcting mistakes in paper-based materials, the correction process can be cumbersome. For example, when a mistake is found in a printed material, it may be necessary to go back to the personal computer, correct the mistake in the input data, and print it again. Also, for example, when it is desired to correct a paper-based material received from the creator, it may be necessary to ask the creator to make corrections or to receive the original data of the material, make corrections, and print it.
[0004] Furthermore, mistakes in materials can occur in various patterns such as typos, numerical errors, and incorrect font usage, and corrections may be required for various mistake patterns.
[0005] The following prior art is disclosed in Patent Document 1 below. A corrected manuscript in which first specific information indicating editing processing to be performed on a printed matter formed on a sheet based on image data by an image forming apparatus and second specific information indicating an editing target portion are handwritten by a user is read by an image reading unit. Then, based on the corrected image data obtained by reading the corrected manuscript by the image reading unit, the first specific information and the second specific information are detected, and the editing processing indicated by the first specific information is executed on the editing target portion indicated by the second specific information.
[0006] Patent Document 2 discloses the following prior art: Text data is generated based on the portion of the first image data corresponding to characters in the first image data generated by reading a document with an image reading unit, and first attribute information indicating the position and size of the text data in the first image data is detected. The detected first attribute information is stored in a storage unit in association with the string information indicated by the text data. If it is determined that the string contains a first string with a spelling mistake, the information of the first string in the stored string information is replaced with the information of a second string that does not contain a spelling mistake, and the first attribute information is changed according to the difference between the first string and the second string. Then, second image data is generated by inserting the text data indicating the stored string information according to the position and size indicated by the changed first attribute information. [Prior art documents] [Patent Documents]
[0007] [Patent Document 1] Japanese Patent Publication No. 2019-29823 [Patent Document 2] Japanese Patent Publication No. 2020-91650 [Overview of the Initiative] [Problems that the invention aims to solve]
[0008] However, the prior art disclosed in Patent Document 1 has the problem that it is burdensome for the user to have to remember the rules, as it requires them to make correction instructions according to a predetermined rule of writing predetermined codes, such as first specific information and second specific information, on the printed material.
[0009] Furthermore, the prior art disclosed in Patent Documents 1 and 2, respectively, has the problem that the possible modifications are limited to modifications based on pre-set rules, such as changing fonts, deleting characters, and correcting spelling mistakes, making it difficult to accommodate a wide range of modifications.
[0010] This invention was made to solve these problems. Specifically, it aims to provide an image forming system, an image forming method, and an image forming program that do not require the user to memorize rules for modification and that can handle a variety of modifications. [Means for solving the problem]
[0011] The above-mentioned problems of the present invention are solved by the following means.
[0012] (1) An image forming system comprising: an image data acquisition unit that scans a printed document with an image formed on a recording medium to acquire image data of the printed document; a recognition unit that recognizes information in the acquired image data; a correction instruction acquisition unit that acquires correction instructions in natural language for the recognized information; a corrected information acquisition unit that inputs the recognized information and the acquired correction instructions into a large-scale language model and acquires corrected information from the large-scale language model in which the information has been corrected based on the correction instructions; and a corrected image data generation unit that generates corrected image data for forming an image on a recording medium based on the corrected information.
[0013] (2) The image forming system according to (1) above, wherein the correction instruction acquisition unit adds a predetermined prompt to the correction instruction, and the corrected information acquisition unit inputs the correction instruction with the predetermined prompt added and the information to the large-scale language model, and the large-scale language model acquires the corrected information from the large-scale language model in which the information has been corrected based on the predetermined prompt and the correction instruction.
[0014] (3) The image forming system according to (2) above, wherein the predetermined prompt is a prompt that gives instructions for further necessary modifications in relation to the modifications made by the modifications acquired by the modification instruction acquisition unit.
[0015] (4) The image forming system according to (2) above, wherein the predetermined prompt is a prompt that instructs the recognition unit to correct an error in the information recognized.
[0016] (5) The image forming system according to (1) above, further comprising an image forming unit that forms an image on the recording medium based on the corrected image data, and further comprising a display unit that previews the image of the corrected image data or previews the image of the corrected image data with the corrected parts highlighted, after the corrected image data is generated by the corrected image data generation unit and before the image is formed on the recording medium by the image forming unit.
[0017] (6) The image forming system according to (1) above, further comprising a transmission unit that transmits the modified image data to a user, or transmits the modified image data to a user with the modified parts clearly indicated.
[0018] (7) An image forming method performed by an image forming system, comprising: (a) scanning a printed material on a recording medium to acquire image data of the printed material; (b) recognizing information of the image data acquired in step (a); (c) acquiring natural language correction instructions for the information recognized in step (b); (d) inputting the information recognized in step (b) and the correction instructions acquired in step (c) into a large-scale language model to acquire corrected information from the large-scale language model in which the information has been corrected based on the correction instructions; and (e) generating corrected image data for forming an image on a recording medium based on the corrected information acquired in step (d).
[0019] (8) Further comprising a step (f) of adding a predetermined prompt to the correction instruction, wherein in step (d), the correction instruction with the predetermined prompt added and the information are input into the large language model, and the corrected information obtained from the large language model based on the predetermined prompt and the correction instruction is obtained from the large language model, the image forming method according to (7) above.
[0020] (9) The predetermined prompt is a prompt for instructing further necessary corrections in relation to the correction by the correction instruction obtained in step (c), the image forming method according to (8) above.
[0021] (10) The predetermined prompt is a prompt for instructing correction of errors with respect to the information recognized in step (b), the image forming method according to (8) above.
[0022] (11) Further comprising a step (g) of forming an image on the recording medium based on the corrected image data, and after the corrected image data is generated in step (e) and before the image is formed on the recording medium in step (g), further comprising a step (h) of previewing the image of the corrected image data or previewing the image of the corrected image data with the correction locations highlighted, the image forming method according to (7) above.
[0023] (12) Further comprising a step (i) of transmitting the corrected image data to the user or transmitting the corrected image data to the user with the correction locations clearly indicated, the image forming method according to (7) above.
[0024] (13) An image forming program for causing a computer to execute the image forming method according to any one of (7) to (12) above.
Advantages of the Invention
[0025] Recognize the information of the image data obtained by scanning the printed matter with the formed image, obtain a natural language correction instruction for the recognized information, input the information and the correction instruction into a large language model, and obtain the corrected information after correction. Then, based on the information after correction, generate the image data after correction used for image formation. Thereby, the user does not need to remember the rules for correction, and can handle various corrections.
Brief Description of the Drawings
[0026] The advantages and features provided by one or more embodiments of the present invention will be more fully understood from the following detailed description and the accompanying drawings. However, these are for illustrative purposes only and are not intended to limit the present invention. [Figure 1] It is a diagram showing a schematic configuration of an image forming system. [Figure 2] It is a block diagram showing the hardware configuration of an image forming apparatus. [Figure 3] It is a diagram showing a schematic configuration of an image forming apparatus. [Figure 4] It is a functional block diagram of a control unit. [Figure 5] It is a diagram showing manuscript image data information. [Figure 6] It is a diagram showing a correction instruction. [Figure 7] It is a diagram showing manuscript image data information after correction. [Figure 8] It is a diagram showing a prompt for related correction. [Figure 9] It is a block diagram showing the hardware configuration of a server. [Figure 10] It is a functional block diagram of a control unit of a server. [Figure 11] It is a diagram showing the time-series mutual relationship of the operations of a user, an image forming apparatus, and a server.
Modes for Carrying Out the Invention
[0027] Hereinafter, an image forming system, an image forming method, and an image forming program according to embodiments of the present invention will be described with reference to the attached drawings. However, the scope of the present invention is not limited to the disclosed embodiments. In the description of the drawings, the same elements are denoted by the same reference numerals, and redundant descriptions are omitted. Also, the dimensional ratios in the drawings are exaggerated for illustrative purposes and may differ from the actual ratios.
[0028] Figure 1 shows a schematic configuration of the image forming system 1.
[0029] The image forming system 1 includes an image forming apparatus 100 and a server 200. The image forming system 1 may also consist only of the image forming apparatus 100, which performs the functions of the server 200 described later. The server 200 is, for example, a cloud server. The server 200 may also be an on-premises server.
[0030] The image forming apparatus 100 and the server 200 can be connected to each other in a way that allows them to communicate with one another.
[0031] Figure 2 is a block diagram showing the hardware configuration of the image forming apparatus 100. Figure 3 is a diagram showing the schematic configuration of the image forming apparatus 100.
[0032] The image forming apparatus 100 comprises a control unit 110, a storage unit 120, a communication unit 130, an operation display unit 140, an image reading unit 150, an audio acquisition unit 160, an image control unit 170, and an image forming unit 180. These components are connected to each other via a bus 190 so as to be able to communicate with one another. The image forming apparatus 100 may be configured as an MFP (MultiFunction Peripheral). For the sake of simplicity, the following explanation will be based on the example where the recording medium on which the image forming apparatus 100 forms an image is paper 900. In addition to paper 900, the recording medium may include resin film, etc. The image reading unit 150 constitutes the image data acquisition unit. The control unit 110, storage unit 120, communication unit 130, and operation display unit 140 of the image forming apparatus 100 constitute a computer.
[0033] The control unit 110 is equipped with a CPU (Central Processing Unit) and various types of memory, and controls the above-mentioned parts and performs various calculations according to the program. Details of the functions of the control unit 110 will be described later.
[0034] The storage unit 120 is composed of an SDD (Solid State Drive) or HDD (Hard Disk Drive), etc., and stores various programs and various data.
[0035] The communication unit 130 is an interface for communication between the image forming apparatus 100 and external devices. Network interfaces such as Ethernet (registered trademark), SATA, and IEEE 1394 can be used as the communication unit 130. Alternatively, various local connection interfaces such as Bluetooth (registered trademark) and IEEE 802.11 wireless communication interfaces can also be used as the communication unit 130.
[0036] The operation display unit 140 is equipped with a touch panel, a numeric keypad, a start button, a stop button, etc., and is used for displaying various information and inputting various instructions.
[0037] The image reading unit 150 scans the original document and acquires image data of the document. The image reading unit 150 has a light source such as a fluorescent lamp and an image sensor such as a CCD (Charge Coupled Device) image sensor. The image reading unit 150 shines light from the light source onto the original document set at a predetermined reading position, converts the reflected light into electrical signals using the image sensor, and generates image data from the resulting electrical signals. The original document includes printed materials on which an image has been formed on paper 900. Printed materials include, for example, forms.
[0038] The audio acquisition unit 160 detects sound and converts the detected sound into audio data by sampling and quantizing it. This allows the sound to be acquired as audio data. The audio acquisition unit 160 is composed of, for example, a microphone and an AD converter.
[0039] The image control unit 170 performs layout processing and rasterization processing of print data included in print jobs, etc., received by the communication unit 130, and generates image data in bitmap format.
[0040] A print job is a general term for print commands to the image forming apparatus 100, and includes print data and print settings. Print data is the data of the document to be printed, and may include various types of data such as image data, vector data, and text data. Specifically, print data may be PDL (Page Description Language) data, PDF (Portable Document Format) data, or TIFF (Tagged Image File Format) data. Print settings are settings related to image formation on paper 900, and include various settings such as the number of pages, number of copies, paper type, color or monochrome selection, double-sided printing, and page layout.
[0041] The image forming unit 180 includes an image forming unit 40, a fixing unit 50, a paper feeding unit 60, and a paper transport unit 70. The paper transport unit 70 has a transport path for transporting the paper 900 using a plurality of transport rollers 72.
[0042] The image-forming unit 40 has image-forming units 41Y, 41M, 41C, and 41K corresponding to toners of the following colors: Y (yellow), M (magenta), C (cyan), and K (black). Based on the image data, each image-forming unit 41Y, 41M, 41C, and 41K forms a toner image on the photoreceptor drum 42 through the processes of charging, exposure, and development. Exposure is performed by scanning the photoreceptor drum 42 with laser light. The toner image formed on the photoreceptor drum 42 is sequentially superimposed onto the intermediate transfer belt 43 by electrostatic force from a constantly controlled transfer voltage applied to the primary transfer roller 44, and is transferred in the primary. This holds the color toner image on the intermediate transfer belt 43. The color toner image on the intermediate transfer belt 43 is then transferred to the paper 900 by the secondary transfer roller 45.
[0043] The fixing unit 50 includes a fixing roller 51a and a pressure roller 52. The fixing roller 51a and the pressure roller 52 are pressed against each other, forming a nip between them. The fixing unit 50 heats and pressurizes the paper 900 that has been transported to the nip, and rotates the fixing roller 51a and the pressure roller 52 to heat and fix the toner image on the paper 900 to the surface of the paper 900.
[0044] The paper 900, on which the toner image has been heated and fixed, is discharged as a printed material into the output tray 90 by the transport roller 72.
[0045] If the print job's print settings are set to double-sided printing, the paper transport unit 70 transports the paper 900, on which the toner image has been heated and fixed to the surface, to the ADU (Auto Duplex Unit) transport path 80. After being transported to the ADU transport path 80, the paper 900 is flipped over via a switchback path, then rejoins the transport path 71, where the image forming unit 180 forms an image on the back side of the paper 900 again.
[0046] The functions of the control unit 110 will now be described.
[0047] Figure 4 is a functional block diagram of the control unit 110.
[0048] The control unit 110 functions as a document data analysis unit 111, an audio data analysis unit 112, a communication control unit 113, an image processing unit 114, and a display control unit 115 by executing a program. The document data analysis unit 111 constitutes the recognition unit. The audio data analysis unit 112, together with the audio acquisition unit 160, constitutes the correction instruction acquisition unit. The communication control unit 113, together with the communication unit 130, constitutes the transmission unit. The image processing unit 114, together with the image control unit 170, constitutes the corrected image data generation unit. The display control unit 115 constitutes the display unit.
[0049] The document data analysis unit 111 performs document analysis processing to recognize image data information of the document by analyzing the image data of the document acquired by the image reading unit 150. Specifically, the document data analysis unit 111 analyzes characters (text), lines, and their positions (coordinates) on the document as image data information. The image data information can be recognized using known techniques such as OCR, and line segment analysis processing including edge detection and binarization.
[0050] The manuscript data analysis unit 111 can output the recognition results of the image data information of the manuscript in text format. Hereinafter, the text format information of the image data of the manuscript output by the manuscript data analysis unit 111 will also be referred to as "manuscript image data information".
[0051] Figure 5 shows the original image data information.
[0052] In the example shown in Figure 5, the grid lines in the original image data are included as text data indicating the origin coordinates, width, and height of 10 rectangles, Rect0 to Rect9. Additionally, the text information from the original image data—"Price List," "Price," "A," "400," "B," "200," "C," "300," "Total," and "900"—is included in the original image data, associated with the rectangle (Rect0 to Rect9) in which each character is written. According to the text information in the original image data, for example, as underlined in Figure 5, the price for "A" in the form is "400." Furthermore, the value of "Total" in the form is "900."
[0053] The audio data analysis unit 112 performs audio data analysis processing to convert the audio data acquired by the audio acquisition unit 160 into text. As a result, the audio data analysis unit 112 acquires and outputs the text-converted audio data. The audio data can be converted into text using known speech recognition technology. In particular, if audio data of natural language correction instructions for a document is acquired by the audio acquisition unit 160, the audio data analysis unit 112 converts this audio data into text. As a result, the audio data analysis unit 112 acquires and outputs the natural language correction instructions in text format. The correction instructions may be in the voice of a user who has viewed the document. Hereinafter, the text-format, natural language correction instructions acquired and output by the audio data analysis unit 112 will also be simply referred to as "correction instructions."
[0054] Figure 6 shows the correction instructions.
[0055] As shown in Figure 6, possible correction instructions include, for example, "Please change the price of A to 100."
[0056] The communication control unit 113 transmits the document image data information output by the document data analysis unit 111 and the correction instructions output by the audio data analysis unit 112 to the server 200 via the communication unit 130. As will be described later, the server 200 uses the large-scale language model 270 to acquire the document image data information corrected according to the correction instructions, based on the document image data information and the correction instructions. Hereinafter, the document image data information corrected according to the correction instructions will also be referred to as "corrected document image data information".
[0057] The large-scale language model 270 is a natural language processing model trained using deep learning with a relatively large amount of text data. It can recognize input natural language and generate and output natural language.
[0058] Figure 7 shows the image data information of the corrected original document.
[0059] As shown by the solid underline in Figure 7, in the corrected original document image data information, the fee for "A" has been corrected to "100" in accordance with the correction instructions shown in Figure 6.
[0060] Furthermore, as shown by the dashed underline in Figure 7, in the revised original image data information, the "Total" value is corrected to "600" in response to the correction of the "A" fee to "100". This correction is an additional correction necessary in relation to the correction instructions. Hereafter, additional corrections necessary in relation to correction instructions will also be referred to as "related corrections".
[0061] As described later, related modifications are input to the large-scale language model 270 along with the original image data, with a predetermined prompt for instructing related modifications added to the modification instruction at the server 200. As a result, the large-scale language model 270 outputs the modified original image data information, which includes the modifications made by the modification instruction along with the related modifications.
[0062] Figure 8 shows the prompts for the related corrections.
[0063] As shown in Figure 8, a prompt for related corrections might look like this: "Please correct the OCR analysis results according to my instructions. If there are any items that need to be corrected in relation to the change instructions, please correct those as well. Please output the results in JSON format only once." The communication control unit 113 receives the revised manuscript data information from the server 200 via the communication unit 130.
[0064] The image processing unit 114, in cooperation with the image control unit 170, generates corrected image data for image formation on the paper 900 by performing an image processing that converts the corrected original document data information in text format into an image. The control unit 110 then uses the image forming unit 180 to form an image on the paper 900 based on the corrected image data.
[0065] The display control unit 115 may preview the corrected image data on the operation display unit 140 after the corrected image data is generated by the image processing unit 114 and before the image is formed on the paper 900 by the image forming unit 180 based on the corrected image data. The display control unit 115 may also preview the corrected image data on the operation display unit 140, highlighting the corrected areas. In this case, the display control unit 115 can identify the corrected areas by comparing the corrected image data information with the image data information stored in the storage unit 120, which corresponds to the corrected image data information. The image data information can be associated with each other, for example, by assigning a common number. Highlighting the corrected areas includes, for example, underlining the corrected areas, making the text of the corrected areas bold, or adding a marker to the corrected areas.
[0066] The communication control unit 113 may send the corrected image data to the user. Specifically, for example, the communication control unit 113 may send the corrected image data to the email address of a user who printed the original document or scanned the original document with the image reading unit 150, and who has been registered in advance. The communication control unit 113 may also send the corrected image data to the user with the corrected parts clearly indicated or highlighted. The communication control unit 113 may also upload the corrected image data to a database accessible to the user.
[0067] Figure 9 is a block diagram showing the hardware configuration of server 200. Figure 10 is a functional block diagram of the control unit 210 of server 200. For the sake of simplicity, a large-scale language model 270 is also shown in Figure 10.
[0068] As shown in Figure 9, the server 200 comprises a control unit 210, a storage unit 220, a communication unit 230, and an operation display unit 240. The basic functions of these components are the same as those of the corresponding components of the image forming apparatus 100, so their explanation is omitted.
[0069] As shown in Figure 10, the control unit 210 functions as a communication control unit 211 and a natural language processing unit 212 by executing a program. The natural language processing unit 212 constitutes a modified information acquisition unit and a modification instruction acquisition unit.
[0070] The communication control unit 211 receives original image data information and correction instructions from the image forming apparatus 100 via the communication unit 230.
[0071] The natural language processing unit 212 inputs the original image data information and the correction instructions to the large-scale language model 270. The natural language processing unit 212 then retrieves the corrected original image data information, which has been modified by the large-scale language model 270 based on the correction instructions. If a predetermined prompt for instructing related corrections is pre-registered, the natural language processing unit 212 adds this prompt to the correction instructions. The natural language processing unit 212 inputs the correction instructions with the predetermined prompts and the original image data information to the large-scale language model 270. The natural language processing unit 212 retrieves the corrected image data information, which has been modified by the large-scale language model 270 based on the corrections made by the correction instructions and related corrections based on the predetermined prompts, from the large-scale language model 270. The predetermined prompts are prompts that instruct further related corrections necessary in relation to the corrections made by the correction instructions, and can be pre-registered by storing them in the storage unit 220. The large-scale language model 270 can be stored in the storage unit 220.
[0072] The predetermined prompt may include, in response to cases where the correction instructions include instructions to correct errors in specific proper nouns, etc., the predetermined prompt may also include, as related corrections, instructions to correct all instances of the same proper noun in the original image data information. In this case, the predetermined prompt might be something like, "If there are multiple names to be corrected, please correct them all according to the correction instructions." The large-scale language model 270, in accordance with the input predetermined prompt, outputs corrected image data information, in which the errors in the original image data information have been corrected, as related corrections.
[0073] Furthermore, the specified prompts may include prompts that instruct the correction of errors in the original image data information. For example, the specified prompts may include a message such as, "Please correct any errors."
[0074] The large-scale language model 270 can recognize instructions for various string manipulations such as substitution, deletion, insertion, and calculation, and can therefore correct various patterns of errors contained in the document.
[0075] Figure 11 is a diagram showing the interrelationship between the operations of the user, the image forming apparatus 100, and the server 200 in chronological order. The numbers shown in Figure 11 indicate the order in which the operations are performed in chronological order.
[0076] (1) The user places the original document at a predetermined reading position on the image reading unit 150 of the image forming apparatus 100.
[0077] (2) The control unit 110 of the image forming apparatus 100 scans the document using the image reading unit 150 to acquire image data of the document. The control unit 110 may also start scanning the document in response to instructions from the user entered on the operation display unit 140.
[0078] (3) The control unit 110 of the image forming apparatus 100 performs a document analysis process to recognize the image data information of the document by analyzing the image data of the document acquired by the image reading unit 150. As a result, the control unit 110 obtains document image data information in text format.
[0079] (4) The user inputs correction instructions by voice. The control unit 110 of the image forming apparatus 100 receives the correction instructions. The control unit 110 may start accepting voice correction instructions when it receives instructions from the user input on the operation display unit 140.
[0080] (5) The control unit 110 of the image forming apparatus 100 performs an audio data analysis process to convert the audio data into text and obtains a text-format correction instruction.
[0081] (6) The control unit 110 of the image forming apparatus 100 sends the original image data information and correction instructions to the server 200, thereby giving instructions for correcting the original image data information using the large-scale language model 270.
[0082] (7) The control unit 210 of the server 200 inputs the original image data information and correction instructions to the large-scale language model 270, causing the large-scale language model 270 to perform correction (inference) of the original image data information.
[0083] (8) The control unit 210 of the server 200 transmits the corrected document image data information in text format, which is output from the large-scale language model 270 as the result of correction (inference result) of the document image data information, to the image forming apparatus 100 as a response to the correction instruction.
[0084] (9) The control unit 110 of the image forming apparatus 100 performs image processing of the corrected original image data information to generate corrected image data in image format.
[0085] (10) The control unit 110 of the image forming apparatus 100 forms an image on the paper 900 based on the corrected image data. This provides the user with the corrected original.
[0086] Although the image forming apparatus 100 performs operations (2) and (3) between operations (1) and (4) above, operations (2), (3), (5), and (6) above may be performed after operations (1) and (4) have been performed consecutively.
[0087] In the operation described in (7) above, the control unit 210 of the server 200 may add the predetermined prompts described above to the modification instructions and input them to the large-scale language model 270.
[0088] Once the operation in (9) above is completed, the image forming apparatus 100 may display a preview of the corrected image data. Subsequently, if the user who has confirmed the preview inputs an instruction to form an image on the paper 900 into the operation display unit 140, the operation in (10) above may be performed. If the user who has confirmed the preview inputs an intention to perform further corrections into the operation display unit 140, the operation returns to the operation in (4) above, and the user may give another correction instruction.
[0089] Before or after the operation described in (10) above, the modified image data may be sent to the user via email or other means, or uploaded to a database accessible to the user.
[0090] The embodiment provides the following effects:
[0091] The system recognizes information in image data obtained by scanning printed materials with image formations, acquires correction instructions in natural language for the recognized information, inputs this information and correction instructions into a large-scale language model, and obtains corrected information. Then, based on the corrected information, it generates corrected image data to be used for image formation. This eliminates the need for users to memorize correction rules and allows for diverse corrections.
[0092] Furthermore, a predetermined prompt is added to the correction instruction, and the correction instruction with the predetermined prompt added, along with the relevant information, is input into a large-scale language model. The corrected information, with the information modified based on the predetermined prompt and correction instruction, is then obtained from the large-scale language model. This suppresses the need for additional correction instructions to the large-scale language model and allows for obtaining a reasonable corrected manuscript in a short amount of time.
[0093] Furthermore, predetermined prompts are used to provide instructions for further necessary modifications related to the modifications made by the correction instructions. This helps to suppress additional correction instructions for large-scale language models and allows for obtaining more appropriate corrected manuscripts in a shorter time.
[0094] Furthermore, a predetermined prompt is set to instruct the system to correct errors in the information in question. This reduces the need for additional correction instructions for large-scale language models and allows for the rapid acquisition of a more accurate corrected manuscript with fewer errors.
[0095] Furthermore, after the corrected image data is generated, and before the image is formed on the recording medium, the corrected image data is previewed, or the corrected areas are highlighted in the preview. This gives the user the opportunity to check the corrected information on the original document before printing, and helps to suppress the increase in costs due to reprinting.
[0096] Furthermore, the corrected image data will be sent to the user, or the corrected image data will be sent to the user with the corrected parts clearly indicated. This will allow the user to review the corrected information in the original document more flexibly.
[0097] The present invention is not limited to the embodiments described above.
[0098] For example, some functions performed by the image forming apparatus 100 may be included in the functions of the server 200. Some functions performed by the server 200 may be performed by the image forming apparatus 100.
[0099] Furthermore, in the embodiment, some or all of the processing performed by the program may be replaced with hardware such as circuits.
[0100] While embodiments of the present invention have been described and illustrated in detail, the disclosed embodiments are for illustrative purposes only and are not limiting. The scope of the present invention should be interpreted in accordance with the language of the appended claims. [Explanation of Symbols]
[0101] 1. Image forming system, 100 Image forming apparatus, 110 Control unit, 111 Manuscript Data Analysis Department, 112 Voice Data Analysis Unit, 113 Communication Control Unit, 114 Image processing unit, 115 Display control unit, 120 storage section, 130 Communications Department, 140 Operation display section, 150 Image acquisition unit, 160 Voice acquisition unit, 170 Image control unit, 180 Image forming unit, 200 servers, 210 Control unit, 211 Communication Control Unit, 212 Natural Language Processing Unit, 220 storage section, 230 Communications Department, 240 Operation display section, 270 Large-scale language models.
Claims
1. An image data acquisition unit scans a printed document with an image formed on a recording medium to acquire image data of the printed document, A recognition unit that recognizes the information of the acquired image data, A correction instruction acquisition unit that acquires correction instructions in natural language for the recognized information, A modified information acquisition unit inputs the recognized information and the acquired modification instructions into a large-scale language model, and acquires modified information from the large-scale language model in which the information has been modified based on the modification instructions. A modified image data generation unit generates modified image data for forming an image on a recording medium based on the modified information, An image forming system having the following features.
2. The correction instruction acquisition unit adds a predetermined prompt to the correction instruction, The image forming system according to claim 1, wherein the modified information acquisition unit inputs the modification instruction with the predetermined prompt added and the information to the large-scale language model, and the large-scale language model acquires the modified information from the large-scale language model, in which the information has been modified based on the predetermined prompt and the modification instruction.
3. The image forming system according to claim 2, wherein the predetermined prompt is a prompt for giving instructions for further necessary modifications in relation to the modifications made by the modification instructions acquired by the modification instruction acquisition unit.
4. The image forming system according to claim 2, wherein the predetermined prompt is a prompt that instructs the recognition unit to correct an error in the information it has recognized.
5. The system further includes an image forming unit that forms an image on the recording medium based on the modified image data, The image forming system according to claim 1, further comprising a display unit that previews the image of the modified image data or previews the image of the modified image data with the modified portion highlighted, after the modified image data is generated by the modified image data generation unit and before the image is formed on the recording medium by the image forming unit.
6. The image forming system according to claim 1, further comprising a transmission unit that transmits the modified image data to a user, or transmits the modified image data to a user with the modified parts clearly indicated.
7. A method performed by an image forming system, Step (a) scans a printed document on a recording medium to obtain image data of the printed document, Step (b) involves recognizing the information of the image data acquired in step (a), Step (c) is to obtain natural language correction instructions for the information recognized in step (b), Step (d) involves inputting the information recognized in step (b) and the correction instructions obtained in step (c) into a large-scale language model, and obtaining corrected information from the large-scale language model in which the information has been corrected based on the correction instructions by the large-scale language model. Step (e) generates corrected image data for forming an image on a recording medium based on the corrected information obtained in step (d), An image forming method having the following characteristics.
8. The method further includes step (f) of adding a predetermined prompt to the correction instruction, The image forming method according to claim 7, wherein in step (d), the correction instruction with the predetermined prompt added and the information are input to the large-scale language model, and the corrected information obtained from the large-scale language model, in which the information has been corrected by the large-scale language model based on the predetermined prompt and the correction instruction.
9. The image forming method according to claim 8, wherein the predetermined prompt is a prompt for giving instructions for further necessary modifications in relation to the modifications made by the modification instructions obtained in step (c).
10. The image forming method according to claim 8, wherein the predetermined prompt is a prompt that instructs the correction of an error in the information recognized in step (b).
11. The process further includes a step (g) of forming an image on the recording medium based on the modified image data, The image forming method according to claim 7, further comprising the step (h) of previewing the image of the modified image data, or previewing the image of the modified image data with the modified portion highlighted, after the modified image data is generated in step (e) and before the image is formed on the recording medium in step (g).
12. The image forming method according to claim 7, further comprising step (i) of sending the modified image data to a user, or sending the modified image data to a user with the modified parts clearly indicated.
13. An image forming program for causing a computer to execute the image forming method according to any one of claims 7 to 12.
Citation Information
Patent Citations
JP2019029823A
JP2020091650A