Image reader and image forming apparatus

The image reading device addresses the limitation of requiring text input by using an AI model to generate images from characters in original documents, facilitating image formation on a recording medium.

JP2025128698APending Publication Date: 2025-09-03KYOCERA DOCUMENT SOLUTIONS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024025511
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-22
Publication Date
2025-09-03

AI Technical Summary

Technical Problem

Existing methods require text data input for character conversion, failing to generate images from characters present in original images.

Method used

An image reading device with a reading unit, image processing unit, and display unit that detects and generates images from characters in input images using an AI model, specifically a generative adversarial network (GAN), enabling image formation on a recording medium.

Benefits of technology

Enables image generation based on characters in original documents, allowing for versatile image creation and display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025128698000001_ABST
    Figure 2025128698000001_ABST
Patent Text Reader

Abstract

To provide an image reader and an image forming apparatus capable of generating an image based on characters included in an image formed on a document.SOLUTION: An image reader 20 includes a reading unit 22, an image processing unit 30, an image generation unit 40, and a display unit 50. The reading unit 22 reads an image of a document S with the image including characters formed thereon, to obtain an input image. The image processing unit 30 performs image processing on the input image. The image generation unit 40 generates an output image based on an input character string. The display unit 50 can display the output image. The image processing unit 30 detects characters included in the input image. The image generation unit 40 receives, as an input character string, a character string composed of the characters detected by the image processing unit 30.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image reading device and an image forming device. [Background technology]

[0002] Patent Document 1 discloses a method for converting characters into images, in which an image is generated from input characters based on cross-modal correspondence between the characters and the image. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Chinese Patent No. 109543159 Summary of the Invention [Problem to be solved by the invention]

[0004] The method of Patent Document 1 requires that the characters to be input be prepared as text data from the beginning, and therefore the method of Patent Document 1 cannot generate an image based on the characters contained in the image formed on the original.

[0005] The present invention has been made in view of the above-mentioned problems, and aims to provide an image reading device and an image forming device that can generate an image based on characters included in an image formed on a document. [Means for solving the problem]

[0006] An image reading device according to one aspect of the present invention includes a reading unit, an image processing unit, an image generation unit, and a display unit. The reading unit reads an image of a document on which an image including characters is formed to obtain an input image. The image processing unit performs image processing on the input image. The image generation unit generates an output image based on an input character string. The display unit is capable of displaying the output image. The image processing unit detects the characters included in the input image. The image generation unit accepts a character string consisting of the characters detected by the image processing unit as the input character string.

[0007] An image forming apparatus according to one aspect of the present invention includes the image reading apparatus according to one aspect of the present invention and an image forming unit, the image forming unit forming an image on a recording medium, and the image forming unit being capable of forming the output image generated by the image generating unit on the recording medium. [Effects of the Invention]

[0008] According to an image reading apparatus and an image forming apparatus according to one aspect of the present invention, an image can be generated based on characters included in an image formed on a document. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a diagram schematically illustrating the overall shape of an image forming apparatus including an image reading apparatus according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram schematically illustrating the configuration of an image reading device and an image forming device. [Figure 3] FIG. 2 is a block diagram illustrating the functions of a control unit. [Figure 4] 10 is a flowchart of the operation of the image reading device when image generation is executed. [Figure 5] FIG. 10 is a diagram showing an example of a number designation screen. [Figure 6] FIG. 10 is a diagram illustrating an example of a reception screen. [Figure 7] FIG. 1 is a diagram illustrating a learning method when an AI model using a GAN is used for image generation. [Figure 8] This is a diagram showing a schematic diagram of image generation by an AI model using StackGAN. [Figure 9] FIG. 10 is a diagram illustrating an example of a transmission image selection screen. [Figure 10] FIG. 10 is a diagram illustrating an example of an evaluation screen. [Figure 11] FIG. 10 is a diagram showing an example of a display on the display unit when a single output image is displayed. DETAILED DESCRIPTION OF THE INVENTION

[0010] An image reading device 20 and an image forming device 10 according to an embodiment of the present invention will be described below with reference to the drawings. In the drawings, the same or corresponding parts are designated by the same reference numerals, and description thereof will not be repeated.

[0011] First, an image forming apparatus 10 including an image reading device 20 according to an embodiment of the present invention will be described with reference to Fig. 1. Fig. 1 is a diagram schematically illustrating the overall shape of the image forming apparatus 10 including the image reading device 20.

[0012] The image forming apparatus 10 includes an image reading device 20 according to an example embodiment of the present invention, and an image forming unit 11. The image reading device 20 reads an image formed on a document S to obtain an input image.

[0013] The image forming unit 11 forms an image on a recording medium, such as paper, based on an input image read by the image reading device 20 or image data transmitted from an external device. The image forming device 10 is a device equipped with an image reading function using the image reading device 20. The image forming device 10 may be, for example, a multifunction device further equipped with a printer function, a copy function, a facsimile function, etc. The image forming device 10 forms an image on a recording medium using, for example, an electrophotographic method or an inkjet method.

[0014] The image reading device 20 includes a control unit 12 and a display unit 50. The control unit 12 controls the operation of the image reading device 20. The control unit 12 is, for example, a unit including a processor and a memory. The display unit 50 is capable of displaying various images. The display unit 50 is, for example, a liquid crystal display. The display unit 50 may also be a touch panel display in which a liquid crystal display and a touch panel are integrated.

[0015] The control unit 12 includes a reading unit 22, an image processing unit 30, and an image generation unit 40. The reading unit 22 reads an image of an original S on which an image including characters is formed, and acquires an input image. The image processing unit 30 performs image processing on the input image. The image generation unit 40 generates an output image based on an input character string that has been input. The display unit 50 can display the output image.

[0016] The image reading device 20 in FIG. 1 further includes a supply tray 19, a document transport device 13, and a reading table 14. The supply tray 19 supports one or more documents S. The documents S are recording media on which an image is formed. In this embodiment, an image including text is formed on the documents S. For example, a sheet of paper on which text is printed or a sheet on which text is handwritten can be used as the documents S on which an image including text is formed.

[0017] The document transport device 13 is a device that transports the document S supported on the supply tray 19 one by one to the reading table 14. The document transport device 13 is, for example, an ADF (Automatic Document Feeder). The reading table 14 is a member having a flat surface. The reading table 14 is, for example, a glass plate. The document S is transported to the flat surface of the reading table 14. Then, the image of the document S transported to the reading table 14 is read by the reading unit 22.

[0018] The image forming apparatus 10 in FIG. 1 further includes a paper storage unit 15 and a discharge tray 16. The paper storage unit 15 stores paper sheets or the like as recording media. The image forming unit 11 forms an image on the paper sheets stored in the paper storage unit 15. The image forming unit 11 discharges the paper sheets on which the image has been formed to the discharge tray 16.

[0019] Next, the configurations of the image reading device 20 and the image forming device 10 will be described with reference to Fig. 2. Fig. 2 is a block diagram that schematically shows the configurations of the image reading device 20 and the image forming device 10. The image forming device 10 in Fig. 2 further includes a paper transport unit 25 in addition to the image reading device 20 and the image forming unit 11. The paper transport unit 25 transports paper stored in the paper storage unit 15 to the image forming unit 11.

[0020] The image reading device 20 includes a control unit 12, a document transport device 13, an input / output unit 18, and a scanning mechanism unit 23. The control unit 12 includes an execution unit 26 and a storage unit 27. The input / output unit 18 includes a display unit 50, an operation unit 51, and a communication unit 52.

[0021] The execution unit 26 of the control unit 12 is a unit that executes various functions in accordance with given data and instructions. The execution unit 26 includes a processor such as a central processing unit (CPU). The memory unit 27 is a unit that stores various data. The memory unit 27 includes semiconductor memory such as a random access memory (RAM) and a read only memory (ROM). The execution unit 26 executes programs stored in the memory unit 27, thereby executing various functions of the image reading device 20. For example, the reading unit 22, image processing unit 30, and image generation unit 40 of the control unit 12 may be program data stored in the memory unit 27 and executed by the execution unit 26.

[0022] The input / output unit 18 receives input to the image reading device 20 and outputs various types of information from the image reading device 20. The display unit 50 of the input / output unit 18 realizes output from the image reading device 20 by displaying various images. The operation unit 51 receives operations for the image reading device 20. The operation unit 51 is, for example, a group of switches provided near the display unit 50. Furthermore, if the display unit 50 is a touch panel display, the touch panel portion of the touch panel display may function as the operation unit 51. The communication unit 52 of the input / output unit 18 is a communication interface that communicates data and various signals with the image reading device 20. The communication unit 52 is, for example, a communication interface that communicates via a LAN (Local Area Network).

[0023] The scanning mechanism 23 is a mechanism that actually operates when the reading unit 22 reads the image of the original S. The scanning mechanism 23 includes, for example, a reading head that moves along the reading table 14 below the reading table 14. The reading head irradiates light onto the original S from below the reading table 14 and detects how the light is reflected by the original S, thereby reading one line of data from the image of the original S along the length direction of the reading head. The reading head moves in a direction intersecting the length direction of the reading head while reading one line of data, thereby reading the entire image of the original S.

[0024] Next, various functions executed by the control unit 12 will be described with reference to FIG. 3. FIG. 3 is a block diagram schematically illustrating the functions of the control unit 12. For the sake of explanation, the execution unit 26 and the storage unit 27 are omitted from FIG. 3. The control unit 12 in FIG. 3 further includes an operation management unit 21 and a reward storage unit 29 in addition to the reading unit 22, image processing unit 30, and image generation unit 40. The image processing unit 30 includes a character detection unit 31 and a character string combination unit 32. The image generation unit 40 includes an AI model 41. The operation management unit 21, the reading unit 22, the image processing unit 30, and the image generation unit 40 are, for example, program data stored in the storage unit 27 and executed by the execution unit 26. The reward storage unit 29 is, for example, part of the storage area of ​​the storage unit 27.

[0025] The operation management unit 21 manages the overall operation of the image reading device 20. For example, the operation management unit 21 manages the operations of the document transport device 13 and the scanning mechanism unit 23 when the reading unit 22 reads an image of the document S. The operation management unit 21 also manages the progress of various processes such as image processing performed by the image processing unit 30 and image generation performed by the image generation unit 40.

[0026] The reading unit 22 reads an image of the original S on which an image including characters is formed to obtain an input image, and then inputs the input image to the image processing unit 30. The image processing unit 30 performs image processing on the input image. For example, the image processing unit 30 detects characters included in the input image. Specifically, the character detection unit 31 of the image processing unit 30 detects characters included in the input image. Then, the character string combining unit 32 of the image processing unit 30 combines one or more characters detected by the character detection unit 31 to form one or more character strings.

[0027] The image generation unit 40 receives, as an input string, a string of characters detected by the image processing unit 30. The image generation unit 40 then generates an output image based on the input string using an AI model 41. The AI ​​model 41 is a learned model of artificial intelligence (AI) that has been trained to input a string of characters and output an image. In this embodiment, the output image is generated by the AI ​​model 41 that uses a generative adversarial network (GAN).

[0028] The image reading device 20 obtains an input image by reading an image formed on the original S using the reading unit 22. Then, characters included in the input image are detected by the image processing unit 30. The image generating unit 40 receives a character string consisting of the characters detected by the image processing unit 30 as an input character string and generates an output image based on the input character string. Therefore, the image reading device 20 can generate an image based on the characters included in the image formed on the original S.

[0029] The reward storage unit 29 stores the reward for the AI ​​model 41. The reward is data created based on the classification result or evaluation of the output image generated by the image generation unit 40. The reward is used, for example, for training the AI ​​model 41. Training of the AI ​​model 41 proceeds in a direction aimed at generating output images that will generate a large amount of reward. When a reward is stored in the reward storage unit 29, data related to the input string and output image that resulted in the reward may also be stored in the reward storage unit 29 together with the reward.

[0030] Next, the operation of image reading device 20 will be described with reference to Figures 1, 2, 3, and 4. Figure 4 is a flowchart of the operation of image reading device 20 when image generation is executed. In image reading device 20 having an image generation function, first, in step S10, the image generation function is enabled, thereby starting the execution of image generation. For example, the image generation function is enabled by the user of image reading device 20 operating operation unit 51.

[0031] Next, in step S11, image reading device 20 accepts a designation of the number of output images to be generated by image generation unit 40. For example, a number designation screen for designating the number of output images may be displayed on display unit 50. The user designates the number using operation unit 51 while checking the number designation screen. If the number designation screen can be displayed on display unit 50, image reading device 20 can generate multiple output images from the same input character string.

[0032] Next, in step S12, the image formed on the document S is read by the reading unit 22 (the document S is scanned). The reading unit 22 acquires the image formed on the document S as an input image.

[0033] Next, in step S13, the image processing unit 30 detects characters included in the input image of the document S. To detect characters included in the input image, for example, an OCR (Optical Character Recognition) function provided in the image reading device 20 is used. Specifically, the character detection unit 31 of the image processing unit 30 detects characters included in the input image. Then, the character string combining unit 32 of the image processing unit 30 combines one or more characters detected by the character detection unit 31 to form one or more character strings.

[0034] Next, in step S14, an input character string to be input to the image generating unit 40 is accepted. If the image formed on the document S includes multiple character strings, multiple character strings will also appear combined by the character string combining unit 32. In step S14, a process is performed to accept one of the character strings combined by the character string combining unit 32 as an input character string for the image generating unit 40. For example, a reception screen for accepting one of the character strings combined by the character string combining unit 32 as an input character string for the image generating unit 40 may be displayed on the display unit 50. The user specifies an input character string using the operation unit 51 while checking the reception screen. By accepting the character string combined by the character string combining unit 32 as an input character string for the image generating unit 40, the image reading device 20 can generate an output image corresponding to a specific character string included in the document S, rather than the entire image formed on the document S. Furthermore, if the reception screen can be displayed on the display unit 50, an appropriate character string can be specified as the input character string even if the document S includes multiple character strings.

[0035] Next, in step S15, image generation unit 40 generates an output image based on the input character string. If the number of images specified in step S11 is multiple (two or more), image generation unit 40 generates multiple output images. The generated output images are displayed on display unit 50 in step S16. If multiple output images are generated, the multiple output images are displayed on display unit 50.

[0036] Next, in step S17, one of the multiple output images is selected as a transmission image to be transmitted to the outside. For example, it is preferable that a transmission image selection screen for selecting one of the output images as a transmission image to be transmitted to the outside can be displayed on the display unit 50. The user selects the transmission image using the operation unit 51 while checking the transmission image selection screen. Note that if there is only one output image, step S17 may be skipped. If a transmission image selection screen for selecting the transmission image can be displayed on the display unit 50, an appropriate output image can be selected as the transmission image even if multiple output images are generated.

[0037] Next, in step S18, an evaluation of the output image is accepted. For example, it is preferable that an evaluation screen for accepting evaluations of the output image can be displayed on the display unit 50. The user inputs an evaluation of the output image using the operation unit 51 while checking the evaluation screen. The evaluation is determined, for example, based on how accurately the output image expresses the input character string, how realistic the output image is, etc. Note that evaluation may be performed based on the user's preference without any particular evaluation criteria being set. If an evaluation screen for accepting evaluations of the output image can be displayed on the display unit 50, the image reading device 20 alone can accept evaluations of the output image.

[0038] The evaluation on the evaluation screen is converted into a reward that can be used for training the AI ​​model 41 of the image generation unit 40, and the reward based on the evaluation is stored in the reward storage unit 29. Since the user's aesthetic sense is reflected in the reward based on the evaluation, when the AI ​​model 41 is trained using the reward stored in the reward storage unit 29, the AI ​​model 41 can output an output image that reflects the user's aesthetic sense.

[0039] Next, in step S19, the transmission image selected in step S17 is transmitted to an external device of the image reading device 20. For example, when the transmission image is transmitted to the image forming unit 11, the image forming unit 11 may be capable of forming the output image generated by the image generating unit 40 on a recording medium. If the image forming unit 11 is capable of forming the output image on a recording medium, the image forming device 10 alone completes the conversion from the document S on which characters are formed to a recording medium on which an image is formed. The transmission image may also be transmitted to an external recording device and recorded in the recording device. Alternatively, the transmission image may not be transmitted to an external device and may be stored in the memory unit 27 of the image reading device 20.

[0040] After step S19 is completed, image reading device 20 ends image generation and waits for the next operation from the user (END). Output images other than the transmitted image are either deleted or stored in storage unit 27 of image reading device 20 together with information (label) that says "not selected." Note that all output images, including the transmitted image, may be deleted without being stored.

[0041] Next, an example of the content displayed on the display unit 50 when enabling the image generation function and specifying the number of output images will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of a number specification screen 62. In the following description, it is assumed that the display unit 50 is a touch panel display that also serves as the operation unit 51.

[0042] 5, the display content of display unit 50 is divided into multiple sections, and each section displays a screen with a different function. In Fig. 5, the display content of display unit 50 includes a number of copies specification screen 62 as well as an image generation function activation screen 61. Furthermore, the display content of display unit 50 includes an operation panel (the upper section in Fig. 5) for executing functions (such as a copy function) of image forming apparatus 10 and image reading apparatus 20, but a description of functions that do not affect the image generation function will be omitted.

[0043] The user of the image reading device 20 first selects whether to disable (OFF) or enable (ON) the image generation function using the image generation function enable screen 61. In Fig. 5, enabling the image generation function has been selected, and the "ON" button is highlighted. Note that if "ON" is not selected on the image generation function enable screen 61, the copy count specification screen 62 is preferably inoperable (for example, grayed out).

[0044] After selecting the "ON" button to enable the image generation function, the user then uses the number of output images specification screen 62 to specify the number of output images. The number of output images specification screen 62 in FIG. 5 includes a button for specifying a single output image (Single Output) and a button for specifying multiple output images (Multiple Output). In FIG. 5, the "Multiple Output" button for specifying multiple output images is selected. When the "Multiple Output" button is selected, a number of output specification field 62a (Number of output) that accepts a specific number of images becomes active. The user inputs the desired number of output images ("3" in FIG. 5) into the number of output specification field 62a. It should be noted that it may be possible to specify a single output image by inputting "1" into the number of output specification field 62a without providing the "Single Output" button.

[0045] Next, a reception screen 63 for receiving an input character string 81 will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of the reception screen 63. After the number of output sheets is specified, an image of the original S is read to obtain an input image. Then, characters included in the input image are detected by the character detection unit 31 of the image processing unit 30, and one or more detected characters are combined by the character string combining unit 32 to obtain one or more character strings. When a plurality of character strings are combined by the character string combining unit 32, as shown in Fig. 6, a reception screen 63 for receiving any one character string from the plurality of character strings as an input character string 81 is displayed on the display unit 50.

[0046] Five character strings (character string 59a, character string 59b, character string 59c, character string 59d, and character string 59e) are displayed on reception screen 63 in FIG. 6. Character detection unit 31 of image processing unit 30 first detects each character included in the input image. Then, when two or more characters detected by character detection unit 31 are adjacent to each other in the input image, character string combining unit 32 combines the two or more adjacent characters into a single character string. For example, character string combining unit 32 may consider a series of characters from a predetermined delimiter (e.g., the beginning of a line, a comma, a colon, a semicolon, a space, a line break, etc.) to the next delimiter as a single character string.

[0047] 6, character strings 59a, 59b, 59c, 59d, and 59e are ready to accept a user selection. In FIG. 6, character string 59e ("Ice Mountain") selected by the user is accepted as input character string 81.

[0048] An input character string 81 is input to the AI ​​model 41 of the image generation unit 40. The AI ​​model 41 generates an output image based on the input character string 81. The AI ​​model 41 in this embodiment is an AI model 41 that uses GAN.

[0049] Next, image generation by an AI model 41 using a GAN (Generative Adversarial Network) and a StackGAN will be described with reference to Fig. 7 and Fig. 8. Fig. 7 is a diagram schematically showing a learning method when an AI model 41 using a GAN is used for image generation. Fig. 8 is a diagram schematically showing image generation by an AI model 41 using a StackGAN.

[0050] 7, an AI model 41 using a GAN includes a generator 72 and a discriminator 74. The generator 72 outputs a generated image 73 based on an input 71 (input character string 81 in this embodiment). The discriminator 74 receives the generated image 73 as input and outputs a discrimination result 75 indicating whether the generated image 73 is an appropriate image to output for the input.

[0051] For example, the classifier 74 compares the generated image 73 with a known image (an image considered to be the "correct answer") and changes the value of the classification result 75 depending on how close the generated image 73 is to the known image. For example, the classification result 75 is expressed as a number between "0" and "1", and if the generated image 73 completely matches the known image, the classification result 75 is "1", and if the generated image 73 is completely different from the known image, the classification result 75 is "0".

[0052] The generator 72 receives a reward based on the classification result 75. The closer the classification result 75 is to "1", the better the reward. The generator 72 learns so that it can obtain a better reward, that is, so that the classification result 75 approaches "1". After repeated learning, the generator 72 can obtain a better reward and can generate images that are close to known images (images considered to be "correct").

[0053] Furthermore, in this embodiment, the AI ​​model 41 may be one that uses StackGAN, which is a type of GAN. As shown in FIG. 8 , the AI ​​model 41 that uses StackGAN includes a first stage 83a and a second stage 83b. The first stage 83a outputs a first-resolution image 84a based on an input character string 81. The first stage 83a converts the first-resolution image 84a into a second-resolution image 84b that has a higher resolution than the first-resolution image 84a.

[0054] First, an input character string 81 is converted into an embedding 82, which is a vector representation for embedding the meaning of the input character string 81 in the real number space of the neural network. The embedding 82 is input to a first stage 83a of the AI ​​model 41, which then outputs a first-resolution image 84a. The first stage 83a is trained to generate an output image based on the input character string 81 (its embedded representation 82), but the resolution of the output first-resolution image 84a is relatively low (for example, a 64 × 64 dot image).

[0055] Next, the first-resolution image 84a is input to the second stage 83b together with the embedded representation 82. The second stage 83b outputs the second-resolution image 84b. The second stage 83b is trained to generate an output image based on the input string 81 (the embedded representation 82) and the first-resolution image 84a. By using the first-resolution image 84a as a base, the second stage 83b can generate a second-resolution image 84b (e.g., an image of 256 × 256 dots) with a higher resolution than the first-resolution image 84a. In this embodiment, a high-resolution output image is obtained by generating an output image using an AI model 41 that uses StackGAN.

[0056] Note that the GAN-based AI model 41 generally uses noise as an initial value and generates images by converting the noise based on the learning results. Therefore, the AI ​​model 41 can generate multiple different output images for the same input string 81 by changing the initial noise value.

[0057] Next, with reference to Fig. 9, the transmission image selection screen 64 when multiple output images are generated will be described. Fig. 9 is a diagram showing an example of the transmission image selection screen 64. The transmission image selection screen 64 displays multiple generated output images. The output image in Fig. 9 is an image generated based on the input character string 81, "Ice Mountain." While checking the transmission image selection screen 64, the user determines which output image is appropriate as the output image for the input character string 81 (which output image accurately represents "Ice Mountain" in Fig. 9), and selects one of the output images. The selected output image becomes the transmission image that is transmitted outside the image reading device 20.

[0058] Next, an evaluation screen 65 for receiving evaluations of output images will be described with reference to Fig. 10. Fig. 10 is a diagram showing an example of the evaluation screen 65. When one output image is selected on the transmission image selection screen 64, the transmission image selection screen 64 transitions to the evaluation screen 65. On the evaluation screen 65 in Fig. 10, five stars are displayed below the selected transmission image.

[0059] On the evaluation screen 65, the user gives an evaluation indicating how suitable the selected transmission image is as an output image for the input character string 81. For example, as shown in FIG. 10, when the star symbol on the far right is selected, all five stars are highlighted (e.g., yellow), and the image being transmitted this time receives "five stars," i.e., a high evaluation. When the star symbol on the far left is selected, only the star symbol on the far left is highlighted, and the remaining stars remain a light color (e.g., white), and the image being transmitted this time receives "one star," i.e., a low evaluation. It may also be possible to give an evaluation to output images that were not selected as transmission images.

[0060] The user's rating received on the rating screen 65 is converted into a reward to be used for training the AI ​​model 41, and the reward is stored in the reward storage unit 29 of the image reading device 20. For example, the reward based on the rating is considered to be better the more stars the user gives to the output image. When the reward based on the rating is used for training the AI ​​model 41, the AI ​​model 41 trains so that it can generate output images that will earn a rating of "five stars." Therefore, when the reward based on the rating is used for training the AI ​​model 41, it is expected that the AI ​​model 41 will be able to generate images that will earn high ratings from users.

[0061] Next, the display on the display unit 50 when there is a single output image will be described with reference to Fig. 11. Fig. 11 is a diagram showing an example of the display on the display unit 50 when there is a single output image. When there is a single output image, that is, when the number of outputs is specified as one on the number-of-images specification screen 62, the transmission image selection screen 64 is not displayed. Therefore, as soon as the input character string 81 is selected on the reception screen 63, the single output image and the evaluation screen 65 for receiving an evaluation of the single output image are displayed on the display unit 50.

[0062] After the evaluation of the output image (transmitted image) is received on the evaluation screen 65, it is preferable that the user's choice be left to decide how to handle the output image. For example, the user may transmit the output image selected as the transmitted image to the outside of the image reading device 20. For example, by pressing a button labeled "Send" displayed on the display unit 50, a dialog box for selecting the destination of the transmitted image may be displayed on the display unit 50. For example, when the image forming device 10 is selected as the transmission destination, the transmitted image may be formed on a recording medium (e.g., paper) by the image forming unit 11.

[0063] The embodiments of the present disclosure have been described above with reference to the drawings. However, the present disclosure is not limited to the above embodiments and can be implemented in various forms without departing from the spirit and scope of the present disclosure. The drawings mainly show each component in a schematic manner for ease of understanding, and the thickness, length, number, spacing, etc. of each illustrated component may differ from the actual components due to the convenience of creating the drawings. Furthermore, the materials, shapes, dimensions, etc. of each component shown in the above embodiments are merely examples and are not particularly limited, and various modifications are possible within a scope that does not substantially deviate from the configuration of the present disclosure. [Industrial Applicability]

[0064] The present invention provides an image reading device and an image forming device, and has industrial applicability. [Explanation of symbols]

[0065] 10 Image forming device 11 Image forming unit 12 Control Unit 20 Image reader 22 Reading unit 29 Reward Storage Department 30 Image processing section 31 Character detection unit 32 String Joiner 40 Image generation unit 41 AI models 50 Display 83a 1st stage 83b 2nd stage 84a First resolution image 84b Second resolution image 62 Number of copies selection screen 63 Reception screen 64 Send image selection screen 65 Evaluation screen 81 Input string S manuscript

Claims

1. a reading unit that reads an image including characters from a document and acquires an input image; an image processing unit that performs image processing on the input image; an image generation unit that generates an output image based on an input character string; a display unit capable of displaying the output image; Equipped with the image processing unit detects the characters included in the input image; The image generating unit receives, as the input character string, a character string made up of the characters detected by the image processing unit.

2. The image processing unit a character detection unit that detects the characters included in the input image; a character string combining unit that combines one or more of the characters detected by the character detection unit to form one or more character strings; and The image reading device according to claim 1 , wherein the image generating unit receives, as the input character string, any one of the character strings combined by the character string combining unit.

3. The image reading device according to claim 2 , wherein the display unit is capable of displaying a reception screen for receiving one of the character strings combined by the character string combining unit as the input character string for the image generating unit.

4. The image generation unit generates the output image using an AI model that uses a generative adversarial network; The AI ​​model is a first stage for outputting a first resolution image based on the input string; a second stage for converting the first resolution image into a second resolution image having a higher resolution than the first resolution image; The image reading device according to claim 1 , further comprising:

5. The image reading device according to claim 1 , wherein the display unit is capable of displaying an evaluation screen for receiving an evaluation of the output image.

6. The image reading device according to claim 5 , further comprising a reward storage unit that stores a reward based on the evaluation on the evaluation screen.

7. The image reading device according to claim 1 , wherein the display unit is capable of displaying a number designation screen for designating the number of the output images to be generated by the image generation unit.

8. 8. The image reading device according to claim 7, wherein the display unit is capable of displaying a transmission image selection screen for selecting one of the output images as a transmission image to be transmitted to an external device when the image generation unit generates a plurality of the output images.

9. The image reading device according to any one of claims 1 to 3, an image forming unit that forms an image on a recording medium; Equipped with The image forming apparatus, wherein the image forming unit is capable of forming the output image generated by the image generating unit on the recording medium.

Citation Information

Patent Citations

  • A text generation image method and device

    CN109543159A