Information processing device, information processing system, and program

The information processing device efficiently digitizes bound documents by capturing moving images and extracting still image data based on document characteristics, addressing the inefficiencies of traditional automatic paper feed systems.

JP7823388B2Active Publication Date: 2026-03-04FUJIFILM BUSINESS INNOVATION CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing automatic paper feed functions in scanners and multifunction printers are inefficient when dealing with multiple sheets of paper bound together, requiring users to manually remove staples and re-bind or read sheets individually, which is time-consuming.

Method used

An information processing device and system that captures a moving image of bound documents, extracts still image data based on characteristics such as character orientation and binding position, and manages the data as a group, allowing for efficient digitization without manual disassembly.

Benefits of technology

Enables faster digitization of images on multiple bound sheets by automatically extracting and managing still image data, reducing user effort and time compared to traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007823388000001
    Figure 0007823388000001
  • Figure 0007823388000002
    Figure 0007823388000002
  • Figure 0007823388000003
    Figure 0007823388000003
Patent Text Reader

Abstract

To facilitate, in comparison with conventional technologies, an operation to computerize an image formed in each of plural pieces of documents put in together for a user.SOLUTION: A user terminal 10 that is an information processing device includes a control unit 11. The control unit 11 obtains data on motion images obtained by imaging plural pieces of documents put in together, extracts, from the data on the motion images, data on still images on the basis of the feature of each of the plural pieces of documents originating from a state put in together, and manages the group of such pieces of data on the still images.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing system, and a program. [Background technology]

[0002] When a user scans an image formed on a paper document using a scanner or multifunction printer, the time required for the process can be reduced by using an automatic paper feed function, and related technologies exist (e.g., Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-004403 Summary of the Invention [Problem to be solved by the invention]

[0004] However, if multiple sheets of paper are bundled together and bound with a stapler or the like, the automatic paper feed function cannot be used as is, so the user must first remove the staples, then use the automatic paper feed function to read the sheets, and then re-bind them with a stapler or the like, or place the sheets face down in front of the reading unit of a scanner device or the like and have each sheet read one by one without using the automatic paper feed function, which is time-consuming for the user.

[0005] An object of the present invention is to make the operation of digitizing images formed on each of a plurality of bound original documents less time-consuming for the user than in the past. [Means for solving the problem]

[0006] The invention described in claim 1 includes a processor, wherein the processor acquires data of a moving image obtained by capturing an image of a plurality of sheets of originals in a bound state, and from the acquired data of the moving image, determines characteristics of each of the plurality of sheets of originals that are caused by the bound state. The images formed on the front and back of each of the plurality of originals are and managing the extracted still image data of the plurality of originals as a group. Claim 2 The invention described in is characterized in that the processor extracts data of the still images based on the characteristics of characters included in the images formed on the front and back of each of the plurality of originals. Claim 1 The information processing device is described in the above. Claim 3 The invention described in is characterized in that the processor extracts data of the still image based on the orientation of the characters as a feature of the characters. Claim 2 The information processing device is described in the above. Claim 4 The invention described in is characterized in that the processor extracts the still image data based on the binding position of the plurality of originals, which is specified by the orientation of the characters. Claim 3 The information processing device is described in the above. Claim 5 The invention described in a processor; The processor: acquiring moving image data obtained by capturing an image of a plurality of sheets of original documents in a bound state, extracting still image data from the acquired moving image data based on characteristics of each of the plurality of sheets of original documents that are attributable to the bound state, and managing the extracted still image data of the plurality of sheets of original documents as a group; The plurality of originals The aforementioned The information processing device is characterized in that it extracts features of still image data, and manages the extracted features as features of the group in association with the group. Claim 6 The invention described in claim 1 is characterized in that the processor manages, as the characteristics of the group, one or more of information regarding the orientation of the multiple originals, the side on which the image is formed, the binding position, and reduced printing, in association with the group. 5 The information processing device is described in the above. Claim 7The invention described in the above includes an imaging means for imaging a plurality of sheets of originals in a bound state, an acquisition means for acquiring data of a moving image of the plurality of sheets of originals that have been imaged, and a method for extracting still image data from the acquired data of the moving image based on the characteristics of each of the plurality of sheets of originals that are caused by the bound state. still image an extracting means; and a managing means for managing the extracted still image data as a group; a feature extraction means for extracting features of the still image data of the plurality of originals; a display means for displaying the still image data managed as the group; The management means manages the extracted features as features of the group in association with the group. The information processing system is characterized by: Claim 8 The invention described in the above provides a computer with a function of acquiring moving image data obtained by capturing images of multiple sheets of original documents in a bound state, a function of extracting still image data from the acquired moving image data based on the characteristics of each of the multiple sheets of original documents that are caused by the bound state, and a function of managing the extracted still image data of the multiple sheets of original documents as a group. a function of extracting features of the still image data of the plurality of originals, and a function of managing the extracted features as features of the group in association with the group; This is a program to achieve this. [Effects of the Invention]

[0007] According to the present invention of claim 1, It is possible to extract still image data based on the images formed on the front and back of each of the multiple originals that are the subject of the moving image. It is possible to provide an information processing apparatus that makes the operation of digitizing images formed on each of a plurality of bound originals less time-consuming for the user than before. Claim 2 According to the present invention, it is possible to extract still image data based on the characteristics of characters contained in images formed on the front and back of each of a plurality of originals that are the subject of a moving image. Claim 3 According to the present invention, it is possible to extract still image data based on the orientation of characters contained in images formed on the front and back of each of a plurality of originals that are the subject of a moving image. Claim 4 According to the present invention, it is possible to extract still image data based on the binding position of a plurality of originals that are the subject of a moving image. Claim 5According to the present invention, it is possible to handle a group of data of still images of a plurality of originals extracted from data of a moving image in association with the features of the data. As a result, it is possible to provide an information processing apparatus that makes it possible for the user to digitize the images formed on each of a plurality of bound originals with less effort than in the past. . Claim 6 According to the present invention, it is possible to handle a group of still image data of multiple originals extracted from moving image data in association with information regarding the orientation of the originals, the side on which the image is formed, the binding position, and reduced printing, which are characteristics of the originals. Claim 7 According to the present invention, It is now possible to handle a group of still image data of multiple originals extracted from moving image data in association with their features. It is possible to provide an information processing system that makes the operation of digitizing images formed on each of a plurality of bound originals less time-consuming for the user than in the past. Claim 8 According to the present invention, It is now possible to handle a group of still image data of multiple originals extracted from moving image data in association with their features. It is possible to provide a program that reduces the user's time and effort in digitizing the images formed on each of a plurality of bound originals. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram illustrating an example of the overall configuration of an information processing system to which the present embodiment is applied. [Figure 2] FIG. 1 is a diagram illustrating a hardware configuration of a user terminal as an information processing device to which the present embodiment is applied. [Figure 3] FIG. 1 is a diagram illustrating a hardware configuration of a management server as an information processing apparatus to which the present embodiment is applied. [Figure 4] FIG. 10 is a diagram illustrating the functional configuration of the control unit of a user terminal when the extraction of parameter information, estimation of document placement, extraction of still image data, and management of the still image data as a group are performed on the user terminal side. [Figure 5] This figure shows the functional configuration of the control unit of a user terminal when the extraction of parameter information, estimation of document placement, extraction of still image data, and management of the still image data as a group are performed on the management server side. [Figure 6]This figure shows the functional configuration of the control unit of the management server when the extraction of parameter information, estimation of document placement, extraction of still image data, and management of the still image data as a group are performed on the management server side. [Figure 7] 10 is a flowchart showing the flow of processing of a user terminal. [Figure 8] 10 is a flowchart showing a flow of processing by a management server when, after a user terminal has captured images of a plurality of documents, the management server subsequently analyzes moving image data. [Figure 9] 1A and 1B are diagrams showing a specific example of a technique for capturing moving images of a document containing multiple unbound documents, in which (A) shows the state of the multiple documents before being captured, and (B) shows the state of the multiple documents during the process of being captured. [Figure 10] 1A and 1B are diagrams showing a specific example of a technique for capturing moving images of a document that is bound at one location on the upper left of the long side. FIG. 1A shows the state of the document before imaging. FIG. 1B shows the state of the document during imaging. [Figure 11] 1A and 1B are diagrams showing a specific example of a technique for capturing moving images of a document that is bound at two positions on the left long side of the document. (A) shows the state of the document before imaging. (B) shows the state of the document during imaging. [Figure 12] 1A and 1B are diagrams showing a specific example of a technique for capturing moving images of a document that is bound at two points on the short sides. FIG. 1A shows the state of the document before imaging. FIG. 1B shows the state of the document during imaging. [Figure 13] FIG. 10 is a diagram showing a specific example of a user interface displayed on a display unit of a user terminal. [Figure 14] FIG. 10 is a diagram showing a specific example of a user interface displayed on a display unit of a user terminal. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. (Configuration of information processing system) FIG. 1 is a diagram showing an example of the overall configuration of an information processing system 1 to which the present embodiment is applied. The information processing system 1 is configured by connecting a user terminal 10 and a management server 30 via a network 90. ​​The network 90 is, for example, a LAN (Local Area Network) or the Internet.

[0010] The user terminal 10 is an information processing device such as a smartphone or tablet terminal operated by a user. For example, the user terminal 10 captures moving images of multiple sheets of original documents in a bound or unbound state based on the user's imaging operation. The user's imaging operation is composed of a first action and a second action of the user. Of these, the "first action" is an action for capturing moving images of multiple sheets of original documents as subjects. The "second action" is an action for capturing each of the multiple sheets of original documents as subjects.

[0011] Specifically, as a "first action," the user holds user terminal 10 in one hand and captures a moving image while displaying multiple documents on display unit 16 (see FIG. 2). As a "second action," the user uses the other hand to turn over the multiple documents one by one, starting from the top. For example, the user holds user terminal 10 in the left hand and captures a moving image of the documents as subjects while turning over the multiple documents one by one, starting from the top, with the right hand.

[0012] Furthermore, for example, the user terminal 10 extracts parameter information indicating characteristics of the multiple sheets of originals from the captured moving image data. The parameter information includes characteristics of the multiple sheets of originals that result from the sheets being bound. It may also include characteristics of the multiple sheets of originals that result from the sheets being unbound.

[0013] The extracted parameter information may include, for example, the size of the paper, the orientation of the paper, the side on which an image such as a character or a graphic is formed, the position and number of staples or the like, and the so-called N-up, which is a method of merging images formed on multiple pages onto one page. Among these, the size of the paper can be identified based on information input by the user as information regarding multiple originals. The orientation of the paper can be identified based on the orientation of the formed image, etc. The side on which an image is formed and the position and number of staples or the like can be identified based on whether an image is formed on the back side of the turned original. The N-up can be identified based on the number of rows or columns of formed characters, etc.

[0014] Furthermore, for example, the user terminal 10 estimates the arrangement of multiple documents based on parameter information of the multiple documents. Furthermore, for example, the user terminal 10 extracts still image data for each document and manages the extracted still image data as a group. Details of these processes performed by the user terminal 10 will be described later.

[0015] Here, "manuscript" refers to a sheet of paper with letters or graphics formed on the front or both sides. "Formed" refers to "printed" on the printing surface of the paper that will become the manuscript. "Bound" refers to a state in which some of the multiple sheets of manuscript have been bound together with a stapler or the like. The method of binding multiple sheets of manuscript is not limited to staples, and any method can be used to secure the multiple sheets of manuscript at one or more locations so that they do not become separated. "Unbound" refers to a state in which some of the multiple sheets of manuscript have not been bound together with a stapler or the like, and the individual sheets are independent but are aligned as a whole.

[0016] Furthermore, the "characteristics of the multiple originals due to the bound state" refers to the external characteristics of the multiple originals due to the bound state, such as the outer shape, the length and width, the appearance of the images such as letters and figures printed on the front and back, and the position of the staples. Among these, the external characteristics include, for example, the aspect ratio of the size of the originals and the length of the outer edge. Furthermore, the characteristics of the images printed on the front and back include, for example, the orientation of the letters and figures printed on the front and back. Furthermore, the characteristics of the position of the staples include, for example, single staples on the upper left of the long side of the originals, single staples on the upper right of the long side, double staples on the left of the long side, and double staples on the left of the short side. Specific examples of the "characteristics of the multiple originals due to the bound state" will be described later with reference to FIGS. 9 to 12.

[0017] In addition to the above configuration, the user terminal 10 may be configured to perform only the capturing of moving images among the above processes, with the management server 30 performing the other processes. The management server 30 is an information processing device that serves as a server that manages the entire information processing system 1. In this case, the management server 30 performs, for example, extraction of parameter information indicating the characteristics of multiple documents, estimation of the arrangement of multiple documents, extraction of still image data for each document, and management of the still image data as a group. Details of these processes performed by the management server 30 will be described later.

[0018] Note that the functions of each of the user terminal 10 and management server 30 constituting the information processing system 1 described above are merely examples, and it is sufficient that the information processing system 1 as a whole has the functions to realize the above-described processes. Therefore, some or all of the functions to realize the above-described processes may be shared or cooperated within the information processing system 1. That is, some or all of the functions of the user terminal 10 may be functions of the management server 30, and some or all of the functions of the management server 30 may be functions of the user terminal 10. Furthermore, some or all of the functions of the user terminal 10 and management server 30 constituting the information processing system 1 may be transferred to another server, imaging device, or the like (not shown). This facilitates processing by the information processing system 1 as a whole and also enables the processes to complement each other.

[0019] (Hardware configuration of user terminal) FIG. 2 is a diagram showing the hardware configuration of a user terminal 10 as an information processing device to which this embodiment is applied. The user terminal 10 has a control unit 11, a memory 12, a storage unit 13, a communication unit 14, an operation unit 15, a display unit 16, and an imaging unit 17. These units are connected to each other via a data bus, an address bus, a PCI (Peripheral Component Interconnect) bus, etc.

[0020] The control unit 11 is a processor that controls the functions of the user terminal 10 through the execution of various software such as an OS (operating system) and application software. The control unit 11 is configured, for example, by a CPU (Central Processing Unit). The memory 12 is a storage area that stores various software and data used for executing the software, and is used as a working area for calculations. The memory 12 is configured, for example, by a RAM (Random Access Memory).

[0021] The storage unit 13 is a storage area that stores input data for various software programs, output data from various software programs, etc. The storage unit 13 is configured with, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), a semiconductor memory, etc. that are used to store programs, various setting data, etc. The storage unit 13 stores, as a database that stores various information, an image DB 801 that stores various image data, such as data on moving images of multiple sheets of originals captured by the imaging unit 17 and data on generated still images.

[0022] The communication unit 14 transmits and receives data to and from the management server 30 and external devices via the network 90. ​​The operation unit 15 is configured with, for example, a keyboard, a mouse, mechanical buttons, and switches, and accepts input operations. The operation unit 15 also includes a touch sensor that forms a touch panel integrally with the display unit 16.

[0023] Display unit 16 is configured, for example, by a liquid crystal display or organic EL (Electro Luminescence) display used to display information, and displays images, text data, etc. Display unit 16 also functions as a camera viewfinder for imaging unit 17. Imaging unit 17 is configured by a camera or the like, and captures an image of a subject displayed on display unit 16, which functions as a camera viewfinder, and acquires the image as moving image or still image data.

[0024] (Management server hardware configuration) FIG. 3 is a diagram showing the hardware configuration of a management server 30 serving as an information processing device to which this embodiment is applied. The hardware configuration of management server 30 is similar to the hardware configuration of user terminal 10 shown in Fig. 2 except for image capture unit 17. That is, management server 30 includes control unit 31, memory 32, storage unit 33, communication unit 34, operation unit 35, and display unit 36, which correspond to control unit 11, memory 12, storage unit 13, communication unit 14, operation unit 15, and display unit 16, respectively, of which storage unit 33 stores image DB 901, which corresponds to image DB 801 in Fig. 2.

[0025] (Functional configuration of the control unit of the user terminal) Figure 4 is a diagram showing the functional configuration of the control unit 11 of the user terminal 10 when the extraction of parameter information, estimation of the document layout, extraction of still image data, and management of the still image data as a group are performed on the user terminal 10 side. The control unit 11 of the user terminal 10 functions as a display control unit 101, an input information receiving unit 102, an imaging control unit 103, an image analysis unit 104, a parameter extraction unit 105, an arrangement estimation unit 106, an image extraction unit 107, and an image group management unit 108.

[0026] The display control unit 101, as a display unit, controls the display unit 16 (see FIG. 2) to display various types of information. For example, the display control unit 101 controls the display unit 16 to display a predetermined user interface. This user interface is provided with input fields for inputting various types of information, a display area for displaying various images, and the like. Specifically, it is provided with input fields for inputting information about multiple documents to be imaged (e.g., document size), a display area for displaying images of icons indicating groups of still image data managed by the management server, and the like. Specific examples of the user interface displayed on the user terminal 10 will be described later with reference to FIGS. 13 and 14.

[0027] The input information receiving unit 102 receives information input by a user's input operation. For example, the input information receiving unit 102 receives information input into an input field of a user interface displayed on the display unit 16 (see FIG. 2).

[0028] The imaging control unit 103, as an imaging means, controls the imaging unit 17 (see FIG. 2) to capture moving images of a plurality of original documents as subjects. Specifically, the imaging control unit 103 controls the imaging unit 17 to continuously capture the entire process of a plurality of original documents bound with a stapler or the like being turned over one by one. Data of the moving images captured by the imaging unit 17 is stored and managed in the image DB 801 (see FIG. 2) of the storage unit 13.

[0029] The image analysis unit 104 analyzes the data of the moving images captured by the imaging unit 17. Specifically, the image analysis unit 104 analyzes the state of each stage, starting from the state before the first page of the document is turned over, to the state as the documents are turned over one by one, and to the state after the last page of the document is turned over. The analysis of the moving image data by the image analysis unit 104 may be performed simultaneously with the image capture, or may be performed after the fact on the moving image data generated when the image capture is completed.

[0030] The analysis by the image analysis unit 104 includes an analysis of the trajectory of movement of a predetermined portion of each of the multiple original documents. The "predetermined portion" refers to the portion that moves when the user turns over the original document, and the trajectory of movement of this portion makes it possible to estimate the timing at which the original document changes. Specific examples of the trajectory of movement of a predetermined portion of each of the multiple original documents will be described later with reference to FIGS. 10 to 12.

[0031] For example, if multiple sheets of paper are stapled together at one point on the top left of the long side of a portrait-oriented document, image analysis unit 104 analyzes the relationship between time and the trajectory of movement of the part of the paper that the user holds with their fingers when turning the page (for example, near the bottom right corner of the document). Also, if the paper is stapled together at two points on the left long side, image analysis unit 104 analyzes the relationship between time and the trajectory of movement of the part of the paper that the user holds with their fingers when turning the page (for example, the right long side of the document).

[0032] As a result of the analysis by the image analysis unit 104, for example, it is possible to estimate the timing when the part comes to a standstill as the timing when one page of the document has been turned over. Also, for example, it is possible to estimate the timing when the part comes to a standstill for the last time as the timing when the last page of the document has been turned over. Also, for example, if the analysis by the image analysis unit 104 shows that the trajectory of the part represents a U-turn, it is possible to estimate that the user has failed in the action of turning over the documents one page at a time (for example, they have turned over two pages at once).

[0033] The parameter extraction unit 105 extracts parameter information of the multiple sheets of original document from the moving image data acquired by the image acquisition unit 301. As described above, the features extracted by the parameter extraction unit 105 include features of the multiple sheets of original document that result from the bound state. The parameter extraction unit 105 also extracts parameter information of the data of each still image of the multiple sheets of original document extracted by the image extraction unit 107.

[0034] The arrangement estimation unit 106 estimates the arrangement of the multiple documents based on the parameter information extracted by the parameter extraction unit 105. Specifically, the arrangement estimation unit 106 estimates the arrangement of the multiple documents based on the characteristics of the multiple documents resulting from the state in which the documents are bound by a stapler or the like, among the parameter information extracted by the parameter extraction unit 105. For example, the arrangement estimation unit 106 estimates the arrangement of the multiple documents based on the position where the documents are bound by a stapler or the like, the length and length ratio of each of the two adjacent documents, the outer size, the style of the printed characters, and the like, as the characteristics of the multiple documents resulting from the state in which the documents are bound.

[0035] Furthermore, for example, if the documents are stapled at a single location on the upper left of the long side, the aspect ratios and orientations of the outlines of the two adjacent documents will be reversed, so the arrangement estimation unit 106 estimates the arrangement of the multiple documents based on the aspect ratios of the outlines of the documents calculated from the length and width of the outer edges of the documents, the orientations of the characters and figures printed on the front and back of the documents, and the like.

[0036] Furthermore, for example, if the documents are bound with a stapler or the like at two positions on the left long side, as in bookbinding, when the documents are turned over, one of the two documents on the left and right sides of the spread will be folded, so the length of the outer edge of the document on the folded side will be shorter than the length of the outer edge of the document on the opposite side. In this case, the arrangement estimation unit 106 estimates the arrangement of the multiple documents based on the lengths of the outer edges of the documents.

[0037] Similarly, if the documents are stapled at two positions on the left long side, even if characters are printed on the back side of the first document, the orientation of the characters printed on the back side of the first document will be the same as the orientation of the characters printed on the front side of the second document. In this case, the layout estimation unit 106 estimates the layout of the multiple documents based on the orientation of the characters printed on the front and back sides of the documents. Specific examples of the layout of multiple documents will be described later with reference to FIGS. 9 to 12.

[0038] Image extraction unit 107 functions as an image extraction means that extracts still image data of each of the multiple documents from the moving image data, based on the document layout estimation result by layout estimation unit 106. Image extraction unit 107 also identifies the timing at which the subject document switches from the analysis result of the movement trajectory of a predetermined part by image analysis unit 104, and extracts still image data of each of the multiple documents based on that timing.

[0039] Image group management unit 108 functions as a management unit that manages, as a group, data of still images of multiple sheets of original document extracted by image extraction unit 107. Specifically, image group management unit 108 associates parameter information of the data of still images of the original document extracted by parameter extraction unit 105 with the group of still image data, and stores and manages the parameter information in image DB 801 (see FIG. 2) of storage unit 13. For example, image group management unit 108 manages information such as the orientation, the side on which the image is printed, the position where the image is bound by a stapler or the like, and reduced printing as parameter information of the data of still images of the original document.

[0040] Figure 5 is a diagram showing the functional configuration of the control unit 11 of the user terminal 10 when the extraction of parameter information, estimation of the document layout, extraction of still image data, and management of the still image data as a group are performed on the management server 30 side. The control unit 11 of the user terminal 10 functions as a display control unit 101, an input information receiving unit 102, an imaging control unit 103, a transmission control unit 109, and an information acquisition unit 110. Below, the functional configuration that does not overlap with that in FIG. 4 will be explained.

[0041] The transmission control unit 109 controls the transmission of various information to the management server 30 or to the outside via the communication unit 14 (see FIG. 2). For example, the transmission control unit 109 controls the transmission of a group of still image data managed by the image group management unit 108 to the management server 30. The information acquisition unit 110 acquires various information via the communication unit 14 (see FIG. 2). For example, the information acquisition unit 110 acquires a group of still image data transmitted from the management server 30.

[0042] (Functional configuration of the control unit of the management server) Figure 6 is a diagram showing the functional configuration of the control unit 31 of the management server 30 when the extraction of parameter information, estimation of the document layout, extraction of still image data, and management of the still image data as a group are performed on the management server 30 side. The control unit 31 of the management server 30 functions as an image acquisition unit 301, an image analysis unit 302, a parameter extraction unit 303, an arrangement estimation unit 304, an image extraction unit 305, an image group management unit 306, and a transmission control unit 307. Of these, the image analysis unit 302, the parameter extraction unit 303, the arrangement estimation unit 304, the image extraction unit 305, and the image group management unit 306 are similar to the image analysis unit 104, the parameter extraction unit 105, the arrangement estimation unit 106, the image extraction unit 107, and the image group management unit 108 in Fig. 4, respectively, and therefore description thereof will be omitted.

[0043] The image acquisition unit 301 functions as an acquisition unit that acquires image data via the communication unit 34 (see FIG. 2). For example, the image acquisition unit 301 acquires moving image data of a plurality of sheets of original paper that are bound together with a stapler or the like, captured by the user terminal 10. Specifically, the image acquisition unit 301 acquires moving image data that captures the entire process of multiple sheets of original paper being turned over one by one.

[0044] The moving image data acquired by the image acquisition unit 301 includes the overall appearance of the multiple sheets of originals and the appearance of the images formed on each of the multiple sheets of originals. Here, "images formed on the originals" refers to "images of characters, figures, etc. printed on the printing surface of the paper that serves as the original," and the "printed surface" may be the front surface only or both surfaces (front and back). The moving image data acquired by the image acquisition unit 301 is stored and managed in the image DB 901 (see FIG. 3) of the storage unit 33.

[0045] The transmission control unit 307 controls the transmission of various information to the user terminal 10 or to the outside via the communication unit 34 (see FIG. 3). For example, the transmission control unit 307 controls the transmission of a group of data on still images of multiple pages of documents extracted from the moving image data by the image extraction unit 305 to the user terminal 10.

[0046] (User terminal processing) Fig. 7 is a flowchart showing the flow of processing by the user terminal 10. In the example of Fig. 7, it is assumed that the imaging of multiple sheets of original document and the analysis of moving image data are performed simultaneously, and a series of processes such as the analysis of the moving images are performed by the user terminal 10. The user terminal 10 displays a user interface on the display unit 16 based on an input operation by the user (step 601), and when information about the multiple documents to be imaged is entered in a predetermined input field (YES in step 602), the user terminal 10 acquires the entered information (step 603). On the other hand, if information about the multiple documents to be imaged has not been entered (NO in step 602), the user terminal 10 repeats step 602 until information about the multiple documents to be imaged is entered.

[0047] Here, the user places multiple documents on a desk or the like, and while turning over the documents one by one in order from the top with one hand, captures images using management server 30 held in the other hand. In response to this, user terminal 10 captures images of the multiple documents being turned over one by one based on the user's input operation (step 604). User terminal 10 analyzes the video data of the multiple documents while simultaneously capturing the images, while referring to the information about the multiple documents to be captured acquired in step 603 (step 605), and extracts still image data for each of the multiple documents from the video data based on the results of the analysis (step 606).

[0048] When extraction of the still image data of the last document is completed (YES in step 607), user terminal 10 ends imaging based on the user's input operation (step 608). On the other hand, if extraction of the still image data of the last document is not completed (NO in step 607), step 607 is repeated until extraction of the still image data of the last document is completed.

[0049] The user terminal 10 extracts parameter information for each of the multiple still image data extracted in step 606 (step 609), associates the parameter information with the group of still image data (step 610), and displays an icon representing the group of still image data on the user interface (step 611).

[0050] (Management server processing) FIG. 8 is a flowchart showing the flow of processing by the management server 30 when, after the user terminal 10 has captured images of a plurality of documents, the management server 30 subsequently analyzes the moving image data. When management server 30 receives moving image data of multiple pages of a document from user terminal 10 (YES in step 701), management server 30 acquires the received data (step 702). On the other hand, if moving image data has not been received (NO in step 701), management server 30 repeats the process of step 701 until moving image data is received.

[0051] Furthermore, when information about multiple documents to be imaged is transmitted from the user terminal 10 (YES in step 703), the management server 30 acquires the transmitted information (step 704) and proceeds to step 705. On the other hand, when information about multiple documents to be imaged is not transmitted (NO in step 703), the management server 30 proceeds to step 705.

[0052] Management server 30 analyzes the moving image data acquired in step 702 (step 705), and extracts still image data for each of the multiple documents from the moving image data based on the results of the analysis (step 706). Note that if information about the multiple documents to be imaged has been acquired in step 704, management server 30 analyzes the moving image data while referring to that information, and extracts still image data for each of the multiple documents from the moving image data based on the results of the analysis.

[0053] The management server 30 extracts parameter information for each of the data of the plurality of still images extracted in step 706 (step 707), and manages the parameter information in association with a group of still image data (step 708). Specifically, the parameter information is managed in a manner that allows transmission at any time in response to a request for a group of still image data from the user terminal 10.

[0054] (Example) 9A and 9B are diagrams showing a specific example of a technique for capturing moving images of a plurality of unbound originals. Fig. 9A shows the state of the plurality of originals G before they are captured. Fig. 9B shows the state of the plurality of originals G while they are being captured. As shown in Fig. 9(A), the multiple documents G to be imaged are multiple documents with characters formed on portrait-oriented paper. The first document G1, placed on top of the multiple documents G, has the alphabet characters "ABC" written horizontally near the center on its surface. The user holds user terminal 10 in one hand and starts capturing a moving image from the state before turning over the first document G1.

[0055] The user uses his other hand to pick up any part of the right edge of the first document G1, turn it over while turning it to the left, and place it on the left side of the multiple documents G. FIG. 9(B) shows the state when the first document G1 has been turned over. At this time, the alphabet letters "DEF" are written horizontally near the center on the back side of the first document G1, but the paper orientation is portrait, just like before it was turned over. Also, the alphabet letters "GHI" are written horizontally near the center on the front side of the second document G2 placed on top of the multiple documents G.

[0056] Next, the user picks up any part of the right edge of the second document G2 in the same manner as when turning over the first document G1, turns it over while turning it to the left, and places it on the left side of the multiple documents G. Although not shown, the user repeats the same operation for the third and subsequent documents, turns over the last document placed at the bottom, and places it on top of the multiple documents G with its back side visible, thereby completing the user's turning operation. The user also captures the state of finishing turning the documents.

[0057] As described above, the user's actions of capturing moving images and turning over the documents cause all of the images (such as characters and figures) formed on the front and back of each of the multiple documents G to be stored as moving image data. Furthermore, the layout of the documents is estimated by the above-described analysis of the moving image data. For example, the layout of the documents is estimated from characteristics such as the position and orientation of the paper and the position and orientation of the characters. In the specific example of FIG. 9, the layout of the documents is estimated from characteristics such as the orientation of the paper and characters not changing before and after the documents are turned over. Then, still image data for each of the multiple documents is extracted and managed as a group.

[0058] 10A and 10B are diagrams showing a specific example of a technique for capturing moving images of a document G bound at a single location on the upper left of the long side. Fig. 10A shows the state of the document G before it is captured. Fig. 10B shows the state of the document G during the process of being captured. As shown in Fig. 10(A), the multiple documents G to be imaged are multiple documents with characters formed on portrait-oriented paper, similar to Fig. 9(A), and the alphabet characters "ABC" are written horizontally near the center on the front surface of the topmost first document G1. However, the multiple documents G in Fig. 10(A) have a staple H fastened to one location on the upper left of the long side. Therefore, the multiple documents G in Fig. 10(A) are bound documents. The user holds the user terminal 10 in one hand and starts capturing a moving image from the state before turning over the first document G1.

[0059] The user then uses the other hand to grasp the lower right corner C of the first document G1, and turns it over while turning it toward the upper left, placing the paper in a landscape orientation. Since there is a risk that the turned document G1 will return to its original state, a valley fold may be made in advance at the position of stapler H before imaging begins. Similarly, in the examples of Figures 11 and 12 described below, a valley fold may be made in advance on the line connecting staples H1 and H2 and its extension.

[0060] FIG. 10(B) shows the state after the first document G1 has been turned over. At this time, the back side of the first document G1 has the alphabet characters "DEF" written horizontally near the center, as in FIG. 9(B). However, because the paper is placed landscape, the alphabet characters also appear landscape. Similarly, the front side of the second document G2, placed on top of the multiple documents G, has the alphabet characters "GHI" written horizontally near the center. At this time, because the second document G2 is in a state before being turned over, the alphabet characters "GHI" appear portrait. In this way, when multiple documents G are bound at a single point on the top left of the long side, the characters written on the back side of each document appear landscape, and are captured and stored in that state.

[0061] Next, the user picks up area C at the bottom right corner of the second document G2 in the same manner as when turning over the first document G1, turns it over toward the top left, and places it in a landscape orientation. Although not shown, the user repeats the same procedure for the third and subsequent documents, turns over the last document placed at the bottom, and places it on top of the multiple documents G with its back side visible, thereby completing the user's turning operation. The user also captures the state of finishing turning the documents.

[0062] The layout of the original is estimated by the above-described analysis of the moving image data. In the specific example of FIG. 10, the layout of the original is estimated from characteristics such as the orientation of the paper and characters changing to landscape before and after the original is turned over, and characteristics of the trajectory of the movement of the part that moves when the user turns the original. FIG. 10(B) shows an example of this, with an arrow L indicating the trajectory of the movement of part C. The arrow L is shaped to connect the parts C before and after the first original G1 is turned over. The starting point of the arrow L is the position of part C before the first original G1 is turned over, and the end point of the arrow L is the position of part C after the first original G1 is turned over. As a result, for example, when the part C of the first original G1 is located at the end point of the arrow L, the second original G2 is being imaged as the subject. Therefore, by analyzing the trajectory of part C, it is possible to estimate the timing when the original that is the subject is switched.

[0063] 11A and 11B show a specific example of a technique for capturing moving images of a document G bound at two positions on the left long side of the document. Fig. 11A shows the state of the document G before it is captured. Fig. 11B shows the state of the document G during the process of being captured. As shown in FIG. 11(A), the multiple documents G to be imaged are, similarly to FIGS. 9 and 10, multiple documents with characters formed on portrait-oriented paper, and the alphabet characters "ABC" are written horizontally near the center on the front surface of the topmost first document G1. However, the multiple documents G in FIG. 11(A) have staples H fastened to them in two places on the left side of the long side. Therefore, the multiple documents G in FIG. 11(A) are bound documents. The user holds the user terminal 10 in one hand and starts capturing moving images from the state before turning over the first document G1.

[0064] The user then uses his other hand to grasp the lower right corner C of the first document G1, turn it over to the left, and place it down so that the paper remains in portrait orientation. Figure 11(B) shows the state after the first document G1 has been turned over. At this point, the back side of the first document G1 has the letters "DEF" written horizontally near the center, as in Figures 9 and 10, but the paper remains in portrait orientation, just as it was before it was turned over. Furthermore, the front side of the second document G2, placed on top of the multiple documents G, has the letters "GHI" written horizontally near the center. In this way, when multiple documents G are bound together at two points on the left side of the long edge, the apparent orientation of the characters written on the back side of each document remains unchanged, and the images are captured and stored in that state.

[0065] Next, the user picks up the second document G2 by pinching the area C at the bottom right corner, and turns it over to the left while placing it down, in the same manner as when turning over the first document G1. Although not shown, the user repeats the same operation for the third and subsequent documents, and when the last document placed at the bottom has been turned over, the user's turning operation is complete. The user also captures the state of finishing turning the documents.

[0066] Furthermore, the layout of the document is estimated by the above-described analysis of the moving image data. In the specific example of Fig. 11, the layout of the document is estimated from characteristics such as the orientation of the paper and characters not changing before and after the document is turned over, and characteristics of the trajectory of the movement of the part that moves when the user turns over the document. Fig. 11(B) shows an arrow L indicating the trajectory of the movement of part C as a specific example of the analysis of the trajectory of the movement of the part that moves when the user turns over the document. This makes it possible to estimate the timing at which the document that is the subject is switched.

[0067] As shown in FIGS. 11A and 11B, the size and area of ​​the outer edge of the first document G1 are larger before it is turned than after it is turned, due to the area that is hidden due to the bend. The size and area of ​​the outer edge of the last document are larger after it is turned than before it is turned, due to the area that is hidden due to the bend. In contrast, the size and area of ​​the outer edge of a document that is neither the first nor the last document remains the same before and after it is turned. This is because there are areas that are hidden due to the bend both before and after it is turned. This makes it possible to estimate that a document whose outer edge size and area differ before and after it is turned is the first document or the last document. Furthermore, if the size and area of ​​the outer edge become smaller after it is turned, it is the first document, and if the size and area of ​​the outer edge become larger after it is turned, it is the last document.

[0068] 12A and 12B are diagrams showing a specific example of a technique for capturing moving images of a document G bound at two points on the short sides. Fig. 12A shows the state of the document G before it is captured. Fig. 12B shows the state of the document G while it is being captured. As shown in Fig. 12(A), the multiple documents G to be imaged are multiple documents with characters formed on portrait-oriented paper, similar to Figs. 9 to 11, and the alphabet characters "ABC" are written horizontally near the center on the front surface of the topmost first document G1. However, the multiple documents G in Fig. 12(A) have staples H fastened in two places on the top long side. Therefore, the multiple documents G in Fig. 12(A) are bound documents. The user holds the user terminal 10 in one hand and starts capturing moving images from the state before turning over the first document G1.

[0069] The user then uses his other hand to grasp the lower right corner C of the first document G1, flip it over, and place it upside down while still in portrait orientation. Figure 12(B) shows the state after the first document G1 has been flipped over. At this time, the back side of the first document G1 has the letters "DEF" written horizontally near the center, as in Figures 9 to 11. However, because the paper is placed upside down, the letters also appear upside down. Similarly, the front side of the second document G2, placed on top of the multiple documents G, has the letters "GHI" written horizontally near the center. In this way, when multiple documents G are bound at two points on the top of the long sides, the letters written on the back side of each document appear upside down, and are captured and stored in that state.

[0070] Next, the user picks up area C at the bottom right corner of the second document G2 in the same manner as when turning over the first document G1, turns it over while turning it upward, and places the paper upside down. Although not shown, the user repeats the same operation for the third and subsequent documents, turns over the last document placed at the bottom, and places it on top of the multiple documents G with its back side visible, thereby completing the user's turning operation. The user also captures the state of finishing turning the documents.

[0071] Furthermore, the layout of the document is estimated by the above-described analysis of the moving image data. In the specific example of FIG. 12, the layout of the document is estimated from characteristics such as the orientation of the paper and characters being upside down before and after the document is turned over, and characteristics of the trajectory of the movement of the part that moves when the user turns over the document. In FIG. 12(B), an arrow L indicating the trajectory of the movement of part C is shown as a specific example of the analysis of the trajectory of the movement of the part that moves when the user turns over the document. This makes it possible to estimate the timing at which the document that is the subject is switched, similar to FIGS. 10 and 11.

[0072] 11, in the specific example of FIG. 12, the size and area of ​​the outer edge of the first document G1 are larger before it is turned over than after it is turned over by the amount of the area that is hidden by the bend, and the size and area of ​​the outer edge of the last document are larger after it is turned over than before it is turned over by the amount of the area that is hidden by the bend. Furthermore, the size and area of ​​the outer edge of a document that is neither the first nor the last document remains the same before and after it is turned over. This makes it possible to estimate that a document whose outer edge size or area differs before and after it is turned over is the first document or the last document. Furthermore, if the size and area of ​​the outer edge decreases after it is turned over, it is the first document, and if the size and area of ​​the outer edge increases after it is turned over, it is the last document.

[0073] 13 and 14 are diagrams showing specific examples of the user interface displayed on the display unit 16 of the user terminal 10. FIG. The user interfaces shown in Figures 13 and 14 can be displayed by launching dedicated application software for users that has been pre-installed on the user terminal 10, or by accessing a dedicated website for users.

[0074] 13 has input fields for inputting information about the multiple sheets of original document to be imaged. Specifically, input fields are provided for inputting information about the size of the original document, the orientation of the original document, the print side, the binding side, the binding position, the number of bindings, and N-up.

[0075] 13, the document size is "A4," the document orientation is "portrait," the print side is "mixed," the binding side is "long side," the binding position is "top left," the number of bindings is "1," and N-up is "none." Among these, the "mixed" print side refers to the presence of a mixture of documents with images formed only on the front side and documents with images formed on both the front and back sides among multiple documents.

[0076] 13 also has a button B11 for starting to capture moving images. The user presses the button B11 after or without entering information about multiple documents to be captured in the input fields. Then, although not shown, part of the user interface becomes a camera viewfinder, allowing the user to start capturing images.

[0077] 14, icons D1 and D2 representing groups of still image data extracted from moving image data of multiple sheets of document to be imaged are displayed in a selectable manner as a "list of extracted images." By selecting a displayed icon (e.g., icon D1 or icon D2), the user can display the contents or send them to another information processing device (e.g., a printer or a multifunction peripheral) connected to the network 90.

[0078] Additionally, the characteristics of the still image data group are displayed next to each of the icons D1 and D2. Specifically, the file name, size, orientation, number of sheets, and creation date are displayed. Of these, "file name" indicates the name of the still image data group, "size" indicates the paper size, "orientation" indicates the paper orientation, "number of sheets" indicates the number of still image data that make up the group, and "creation date" indicates the date on which the still image data was extracted and began to be managed as a group. Note that in the specific example of FIG. 14, icons D1 and D2 indicating groups of still image data are displayed, but if three or more groups of still image data are saved, icons indicating all saved groups of still image data and their information can be displayed by scrolling the screen.

[0079] For example, in the specific example of FIG. 14, the name (file name) of the group of still image data represented by icon D1 is "XX Conference Materials," the paper size is "A4," the paper orientation is "portrait," the number of pages is "7," and the date on which the still image data was extracted and started to be managed as a group is "MM / DD / 20XX." Now, for example, suppose a user wishes to obtain more detailed information about the group of still image data represented by icon D1. In this case, the user presses button B21 labeled "View Details," which is displayed superimposed on icon D1. Then, although not shown, detailed information about the group of still image data represented by icon D1 is displayed in the user interface.

[0080] Although the present embodiment has been described above, the present invention is not limited to the above-described embodiment. Furthermore, the effects of the present invention are not limited to those described in the above-described embodiment. For example, the configuration of the information processing system 1 shown in FIG. 1 and the hardware configurations of the user terminal 10 and management server 30 shown in FIGS. 2 and 3 are merely examples for achieving the object of the present invention and are not particularly limited. Furthermore, the functional configurations of the user terminal 10 shown in FIGS. 4 and 5 and the functional configuration of the management server 30 shown in FIG. 6 are also merely examples and are not particularly limited. It is sufficient for the information processing system 1 of FIG. 1 to be provided with the functionality to execute the above-described processing as a whole, and the functional configuration used to realize this functionality is not limited to the examples of FIGS. 4 to 6.

[0081] 7 and 8 are merely examples and are not particularly limited. The steps may not necessarily be performed in chronological order, but may be performed in parallel or individually. The specific examples shown in FIGS. 9 to 14 are also merely examples and are not particularly limited.

[0082] Furthermore, in the above-described embodiment, the user terminal 10 has been described on the assumption that it is an information processing device such as a smartphone or tablet terminal, but since it is sufficient to be able to capture moving images of multiple documents, it may also be, for example, a so-called document camera that is capable of capturing images using an imaging device whose position is fixed by an arm, etc. In this case, processes such as extraction of parameter information indicating the characteristics of the multiple documents, estimation of the arrangement of the multiple documents, extraction of still image data for each document, and management of the still image data as a group are performed on the management server 30 side. [Explanation of symbols]

[0083] 1...information processing system, 10...user terminal, 11...control unit, 30...management server, 90...network, 101...display control unit, 102...input information receiving unit, 103...imaging control unit, 104...image analysis unit, 105...parameter extraction unit, 106...arrangement estimation unit, 107...image extraction unit, 108...image group management unit, 109...transmission control unit, 110...information acquisition unit, 301...image acquisition unit, 302...image analysis unit, 303...parameter extraction unit, 304...arrangement estimation unit, 305...image extraction unit, 306...image group management unit, 307...transmission control unit

Claims

1. a processor; The processor: Acquires moving image data of a plurality of bound manuscript sheets, extracting still image data from the acquired moving image data based on images formed on the front and back of each of the plurality of originals, which are characteristics of each of the plurality of originals due to the bound state; The extracted still image data of the plurality of originals is managed as a group. Information processing device.

2. The processor extracts data of the still images based on characteristics of characters included in images formed on the front and back of each of the plurality of original documents. The information processing device according to claim 1 .

3. The processor extracts data of the still image based on the orientation of the character as the feature of the character. The information processing device according to claim 2 .

4. The processor extracts data of the still image based on a binding position of the plurality of originals, which is specified by the orientation of the characters. The information processing device according to claim 3 .

5. A processor is provided, The processor: Acquires moving image data of a plurality of bound manuscript sheets, extracting still image data from the acquired moving image data based on characteristics of each of the plurality of originals due to the bound state; managing the extracted still image data of the plurality of originals as a group; extracting features of the still image data of the plurality of originals; The extracted features are managed as features of the group in association with the group. Information processing device.

6. The processor manages, as the characteristics of the group, one or more of information regarding the orientation of the plurality of originals, the side on which the image is formed, the binding position, and reduced printing, in association with the group. The information processing device according to claim 5 .

7. an imaging means for imaging a plurality of bound originals; an acquisition means for acquiring data of moving images of the plurality of captured originals; a still image extraction means for extracting still image data from the acquired moving image data based on characteristics of each of the plurality of originals due to the bound state; a management means for managing the extracted still image data as a group; a feature extraction means for extracting features of the still image data of the plurality of originals; a display means for displaying the still image data managed as the group; and the management means manages the extracted features as features of the group in association with the group. Information processing system.

8. On the computer, A function to acquire moving image data of a bound document containing multiple pages; a function of extracting still image data from the acquired moving image data based on characteristics of each of the plurality of originals due to the bound state; a function of managing the extracted still image data of the plurality of originals as a group; a function of extracting characteristics of the still image data of the plurality of originals; a function of managing the extracted features as features of the group in association with the group; A program to achieve this.

Citation Information

Patent Citations

  • Overhead scanner apparatus, image acquisition method, and program

    JP2011254366A

  • Image processor, write control program and write control method

    JP2016004403A

  • Information processing device, and data structure of data to be obtained by imaging with respect to paper medium

    JP2016177364A