A document service system, the methods that the document service system performs, and a program that causes a computer to perform the methods that the document service system performs.
The document service system addresses the mismatch in information provision by extracting multiple candidates and priorities, personalizing the model based on user feedback and document type, ensuring information relevance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2026-04-09
AI Technical Summary
Existing document processing systems fail to provide users with information tailored to their specific needs due to variations in user operations and document data, despite using the same document data, leading to mismatched information provision.
A document service system that includes a model storage unit and processing unit to extract multiple candidates and priorities from document data, allowing users to select the most relevant information, with a learning device that personalizes the model based on user feedback and document type.
Enables the provision of information that accurately reflects user intentions by presenting multiple candidates and priorities, adapting to individual user needs and document types, thereby improving information relevance.
Smart Images

Figure 0007843384000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a document service system configured to communicate with a user's terminal, a method performed by the document service system, and a program for causing a computer to execute the method.
Background Art
[0002] Japanese Unexamined Patent Application Publication No. 2020-13281 (Patent Document 1) or Japanese Patent No. 7561378 (Patent Document 2) describes a system that automatically reads character information in document data such as a claim by OCR (Optical Character Recognition / Reader) and automatically extracts one piece of character information as the content of a predetermined item by AI (Artificial Intelligence) and provides it to a user.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the systems described in Patent Documents 1 and 2, when the document data to be processed is the same, the character information extracted as the content of each item is specified to be one.
[0005] However, even if the document data to be processed is the same, the information to be extracted as content for each item will differ for each user due to differences in the operation and position of the user who uses the information in that document data. Despite this, if the text information to be extracted as content for each item is specified to be unique, the user will only be provided with information that does not match their intentions. Patent documents 1 and 2 do not mention this problem or any countermeasures.
[0006] This disclosure was made to solve the aforementioned problems, and its purpose is to facilitate the automatic reading of characters in document data and the provision of information intended by the user to the user. [Means for solving the problem]
[0007] (Section 1) The document service system provided by this disclosure is a document service system configured to communicate with a user's terminal and includes a model storage unit that stores a model configured to automatically read characters in document data when document data acquired from the user's terminal is input, and to extract and output a plurality of candidates that can be selected as the content of a predetermined item from the read characters, and a processing unit that performs processing using the model. The processing unit acquires a plurality of candidates by inputting document data acquired from the user's terminal into the model, and displays the acquired plurality of candidates on the screen of the user's terminal in a manner that allows the user to select one.
[0008] (Paragraph 2) In the document service system described in Paragraph 1, the processing device displays multiple candidates and a layout image of the document data side by side on the user's terminal screen, and highlights the location of the candidate selected by the user from among the multiple candidates on the layout image.
[0009] (Section 3) In the document service system described in Section 1, the model is configured to extract multiple candidates and output the priority of the multiple candidates when document data is input. The processing unit inputs document data obtained from the user's terminal into the model to obtain multiple candidates and the priority of the multiple candidates, and displays the obtained priority on the screen of the user's terminal in a manner that the user can recognize.
[0010] (Article 4) In the document service system described in Article 3, the processing device displays a list of multiple candidates arranged in order of priority on the screen of the user's terminal.
[0011] (Clause 5) In the document service system described in paragraph 4, the processing device defaults to displaying the candidate with the highest priority among multiple candidates on the user's terminal screen, and displays a list of multiple candidates arranged in descending order of priority on the user's terminal screen in response to user operations.
[0012] (Paragraph 6) In the document service system described in Paragraph 1, the processing device displays on the user's terminal screen information that represents the meaning of multiple candidates in addition to the multiple candidates.
[0013] (Clause 7) The document service system described in paragraph 1 further comprises a database that acquires and stores as feedback information the user has selected from among multiple candidates from the user's terminal, and a learning device that learns a model using the feedback information stored in the database as training data. The processing device acquires multiple candidates by inputting document data into the model learned by the learning device and displays them on the screen of the user's terminal.
[0014] (Clause 8) In the document service system described in Clause 7, the database stores feedback information in association with user transaction information to identify the user's transaction trends. The learning device learns a model based on the user transaction information. The processing device inputs document data obtained from the user's terminal and user transaction information into the model to obtain multiple candidates suitable for the user's operations and displays them on the user's terminal screen.
[0015] (Section 9) In the document service system described in Section 7, the database stores document types indicating the type of document data in addition to feedback information. The learning device learns a model based on the document types. The processing device inputs document data acquired from the user's terminal into the model, thereby acquiring multiple candidates suitable for the type of document data and displaying them on the user's terminal screen.
[0016] (Clause 10) In the document service system described in Clause 7, the feedback information includes semantic information representing the meaning of the user-selected candidate, in addition to the candidate selected by the user. The learning device uses the semantic information contained in the feedback information to train a model.
[0017] (Section 11) The method provided herein is a method performed by a document service system configured to communicate with a user's terminal. The document service system stores a model configured to automatically read characters in document data when document data acquired from the user's terminal is input, and to extract and output a plurality of candidates that can be selected as the content of a predetermined item from the read characters. The method performed by the document service system includes the steps of acquiring a plurality of candidates by inputting document data acquired from the user's terminal into the model, and displaying the acquired plurality of candidates on the screen of the user's terminal in a manner that allows the user to select one.
[0018] (Item 12) The program according to the present disclosure is a program that causes a computer to execute a method performed by a document service system configured to be communicable with a user's terminal. The document service system is configured to automatically read characters in document data when the document data acquired from the user's terminal is input, and to extract and output a plurality of candidates selectable as the content of a predetermined item from the read characters, and stores a model. The program causes the computer to execute steps of acquiring a plurality of candidates by inputting the document data acquired from the user's terminal into the model, and displaying the acquired plurality of candidates on the screen of the user's terminal in a manner that allows the user to select them.
Advantages of the Invention
[0019] According to the present disclosure, it is possible to easily provide the user with the information intended by the user by automatically reading the characters in the document data.
Brief Description of the Drawings
[0020] [Figure 1] It is a diagram schematically showing the overall configuration of a system including a document service system. [Figure 2] It is a flowchart showing an example of a processing procedure executed by the document service system. [Figure 3] It is a diagram (Part 1) schematically showing an example of a user terminal image. [Figure 4] It is a diagram (Part 2) schematically showing an example of a user terminal image. [Figure 5] It is a diagram (Part 3) schematically showing an example of a user terminal image. [Figure 6] It is a diagram schematically showing an example of a data structure stored in a database. [Figure 7] It is a diagram schematically showing the AI-OCR processing results when the document service system receives the same document data from Company A, Company B, and Company C, respectively. [Figure 8]A diagram schematically showing the AI-OCR processing results when the document type is "claim form" and the AI-OCR processing results when the document type is "receipt".
Embodiments for Carrying Out the Invention
[0021] Embodiments of the present disclosure will be described in detail while referring to the drawings. For the same or corresponding parts in the drawings, the same reference numerals are given and their descriptions will not be repeated.
[0022] [System Configuration] FIG. 1 is a diagram schematically showing the overall configuration of a system 1 including a document service system 2 according to the present embodiment. This system 1 includes a document service system 2 and a user terminal 3.
[0023] The document service system 2 is configured to be communicable with the user terminal 3. The user terminal 3 is a terminal operated by a user who uses the document service system 2. The user terminal 3 may be a computer equipped with peripheral devices such as a display, a keyboard, and a mouse, or may be a smartphone. Although one user terminal 3 is shown in FIG. 1, the document service system 2 can communicate with a plurality of user terminals 3 used by a plurality of users respectively.
[0024] When the document service system 2 receives document data 4 and user transaction information from the user terminal 3, it automatically reads the characters in the received document data 4 by OCR and performs a process of automatically extracting character information suitable for the content of predetermined display items by AI (hereinafter also referred to as "AI-OCR process"). The user transaction information is information for specifying the transaction tendency of the user. The user transaction information includes a user ID for specifying the user, a document ID for specifying the document data, a destination ID for specifying the user's business partner, a tenant ID for specifying the tenant, and the like. The document service system 2 displays the result of the AI-OCR process on the screen of the user terminal 3 specified by the user transaction information.
[0025] Specifically, the document service system 2 includes a model storage unit 10, a processing unit 20, a database 30, and a learning device 40.
[0026] The model memory unit 10 includes a first memory unit 12 and a second memory unit 14. The first memory unit 12 stores the first model. The second memory unit 14 stores the second model.
[0027] The first model is a model (candidate output model) that, upon input of document data 4, automatically reads the characters in document data 4 using OCR, extracts multiple candidates from the read characters that can be selected as content for predetermined display items, and outputs them. The predetermined display items include multiple items such as "document type" indicating the type of document read (invoice, receipt, etc.), "customer name" indicating the name of the business partner, "amount" indicating the transaction amount, and "transaction date" indicating the date of the transaction.
[0028] The first model is configured to output multiple candidates for each of the multiple items that are suitable for the content of that item. For example, if there are three dates listed in document data 4, the first model is configured to output all three dates listed in document data 4 as candidates for the item "Transaction Date," rather than outputting just one of the three dates listed in document data 4.
[0029] The second model is a priority output model that takes multiple candidates for each item output from the first model as input and outputs the priority of the multiple candidates for each item. For example, if three candidates for a certain item are input to the second model, the second model will stratify the three candidates for that item into the candidate with the highest priority, the candidate with the second highest priority, and the candidate with the lowest priority, and output them.
[0030] The processing unit 20 is an arithmetic circuit (arithmetic unit) that performs various processes by executing various programs. The processing unit 20 includes a processor such as a CPU (Central Processing Unit), memory, and input / output buffers. The processing unit 20 may be implemented in, for example, a general-purpose computer, or in a dedicated computer for document services.
[0031] The processing unit 20 performs the AI-OCR processing described above using the models (first model and second model) stored in the model storage unit 10, and displays the results of the AI-OCR processing on the screen of the user terminal 3. Specifically, the processing unit 20 inputs the document data 4 acquired from the user terminal 3 into the models stored in the model storage unit 10, thereby acquiring multiple candidates for each item and the priority of those candidates.
[0032] The processing unit 20 displays the acquired candidates on the screen of the user terminal 3 in a manner that allows the user to select one (see Figure 4 below). The processing unit 20 also displays the priority of the acquired candidates on the screen of the user terminal 3 in a manner that allows the user to recognize it (see Figure 4 below).
[0033] Database 30 stores AI-OCR processing results and feedback information from user terminal 3. The feedback information from user terminal 3 includes information that identifies the candidate selected by the user from among multiple candidates obtained by AI-OCR processing (hereinafter also referred to as "user-selected candidate"). Database 30 stores combinations of AI-OCR processing results and feedback information for those AI-OCR processing results, associated with user transaction information (see Figure 6 below).
[0034] The learning device 40, like the processing unit 20, is configured to include a processor such as a CPU, memory, and input / output buffers. The learning device 40 may be implemented as, for example, a general-purpose computer or as a dedicated computer for machine learning. The model learning method used by the learning device 40 is not limited to a specific method, and known methods can be adopted.
[0035] The learning device 40 uses the feedback information stored in the database 30 as training data to perform machine learning on the models (first model and second model) stored in the model storage unit 10. In this embodiment, the learning device 40 uses the feedback information stored in association with user transaction information to learn the models stored in the model storage unit 10 for each user transaction. As a result, the models stored in the model storage unit 10 are personalized for each user (more specifically, for each user's transaction content). The processing unit 20 inputs the document data 4 acquired from the user terminal 3 and the user transaction information into the models stored in the model storage unit 10, thereby acquiring multiple candidates for each item that are suitable for each user's use of the document data.
[0036] Furthermore, the learning device 40 learns the model stored in the model storage unit 10 based on the document type included in the AI-OCR processing result. This ensures that the model stored in the model storage unit 10 is adapted to each user's operation for each document type. The processing unit 20 inputs the document data 4 acquired from the user terminal 3 into the model stored in the model storage unit 10, thereby acquiring multiple candidates for each item that are suitable for each user's operation for each document type.
[0037] Figure 2 is a flowchart showing an example of a processing procedure performed by the document service system 2. Of the steps S10 to S60 shown in Figure 2, steps S10 to S30 are the "utilization phase" in which the processing unit 20 utilizes the model to perform AI-OCR processing, and steps S40 to S60 are the "learning phase" in which the learning device 40 learns the model. Each of the processes shown in Figure 2 is expected to be implemented by software (programs), but may also be implemented by dedicated hardware (electronic circuits).
[0038] <Utilization Phase> Referring to Figure 2, the processing of the utilization phase by the document service system 2 (steps S10 to S30) will be explained.
[0039] The processing unit 20 receives document data 4 and user transaction information from the user terminal 3 (step S10).
[0040] Next, the processing unit 20 performs the AI-OCR processing described above using the document data 4 and user transaction information received in step S10 (step S20).
[0041] Specifically, the processing unit 20 acquires multiple candidates for each display item by inputting the document data 4 and user transaction information received in step S10 into the first model, and sets priorities for the multiple candidates for each display item by inputting the acquired multiple candidates and user transaction information into the second model. For example, if three candidates are acquired for a certain display item, the three candidates for that display item are stratified into the candidate with the highest priority, the candidate with the second highest priority, and the candidate with the lowest priority.
[0042] Next, the processing unit 20 transmits the results of the AI-OCR processing in step S20 to the user terminal 3, thereby displaying the AI-OCR processing results on the screen of the user terminal 3 in a predetermined layout (step S30). At this time, the processing unit 20 displays multiple candidates for each display item on the screen of the user terminal 3 in a manner that allows the user to select one.
[0043] Figure 3 schematically shows an example of a user terminal image 3a that is displayed by default on the screen of the user terminal 3 after processing in step S30 by the processing unit 20. In Figure 3, an example is shown where the document data 4 received by the processing unit 20 is a receipt from "AAA Corporation" to "XXX Corporation".
[0044] On the user terminal 3's screen, the result image 70 showing the result of AI-OCR processing by the document service system 2 and the layout image 80 of the document data 4 are displayed side by side.
[0045] The result image 70 displays several predetermined display items. Figure 3 shows an example where multiple items are displayed, including "Document Type," "Customer Name," "Qualified Business Operator Number," "Amount," and "Transaction Date."
[0046] Directly below each display item, there is a display field that shows the content of that item. The display field for each display item defaults to displaying the highest priority candidate. Figure 3 shows an example where "Receipt" is displayed in the document type display field 71, "AAA Corporation" is displayed in the customer name display field 72, "aaaaaa" is displayed in the qualified business number display field 73, "7,630" is displayed in the amount display field 74, and "2024-05-30 (Purchase Date)" is displayed in the transaction date display field 75. Note that the display field for a display item for which no candidates exist is left blank.
[0047] The layout image 80 highlights the locations of the text displayed in the display fields of the result image 70. In the example shown in Figure 3, the locations of "AAA Corporation" and "aaaaaa" displayed in display fields 72 and 73 of the result image 70 are highlighted in the layout image 80 with a frame 82. The location of "7,630" displayed in display field 74 of the result image 70 is highlighted in the layout image 80 with a frame 84, and the location of "2024-05-30 (Purchase Date)" displayed in display field 75 of the result image 70 is highlighted in the layout image 80 with a frame 85. Through this highlighting, users can easily recognize where in the document data 4 the text displayed in the display fields of the result image 70 was located by looking at the layout image 80.
[0048] In the transaction date display field 75, instead of just displaying the date "2024-05-30," information indicating the meaning of that date, such as "(Purchase Date)," is added after the date. This allows users to easily understand the meaning of the date displayed in the transaction date display field 75 simply by looking at it. In other words, if we assume that only the date "2024-05-30" is displayed in the transaction date display field 75, users cannot understand the meaning of May 30, 2024 (the action taken on that date) simply by looking at the transaction date display field 75. In particular, in the example shown in Figure 3, the date "May 30, 2024" is listed in two places in the document data 4: the purchase date field and the boarding date field. Therefore, simply by looking at the transaction date display field 75, it is not possible to determine whether the "2024-05-30" displayed in the transaction date display field 75 is extracted from the purchase date or the boarding date in the document data 4. In contrast, in this embodiment, the transaction date display field 75 displays "2024-05-30 (purchase date)," so the user can easily recognize that the date displayed in the transaction date display field 75 is the "purchase date" simply by looking at it.
[0049] When a user performs a predetermined list display operation on the display field of each display item (for example, selecting the display field of a display item with cursor 76), a list is displayed containing multiple options that the user can select as the content of that display field.
[0050] Figure 4 schematically shows an example of a user terminal image 3b when a user performs a list display operation on the transaction date display field 75. In this case, as shown in Figure 4, the candidate list 75a is displayed below the transaction date display field 75.
[0051] The candidate list 75a displays candidates in descending order of priority from top to bottom. In the example shown in Figure 4, the candidate list 75a displays, from top to bottom, the three candidates: "2024-05-30 (Purchase Date)," which has the highest priority and was displayed by default; "2024-05-30 (Date of Travel)," which has the next highest priority; and "2024-05-31 (Display Date)," which has the lowest priority. By checking the order of the candidates in the candidate list 75a, the user can recognize the multiple candidates obtained by the AI-OCR processing by the document service system 2 and their priorities.
[0052] By performing a predetermined selection operation on the candidate list 75a (for example, selecting one of the candidates within the candidate list 75a using the cursor 76), the date displayed in the transaction date display field 75 can be changed to the date of the candidate selected by the user.
[0053] Below the list of suggested trading dates 75a, a calendar 75b is displayed. Users can also change the date displayed in the trading date display field 75 to any date they wish by selecting a date from the calendar 75b.
[0054] Figure 5 schematically shows an example of a user terminal image 3c when the user selects "2024-05-30 (Date of Travel)," which was the second item from the top in the candidate list 75a in Figure 4. In this case, the date displayed in the transaction date display field 75 changes from the default "2024-05-30 (Date of Purchase)" (see Figure 4) to "2024-05-30 (Date of Travel)" (see Figure 5), which the user selected from the candidate list 75a.
[0055] Furthermore, in conjunction with the change in the date displayed in the transaction date display field 75 to "2024-05-30 (date of travel)", the location of the highlighted frame 85 on the layout image 80 changes from the location where the purchase date is written (see Figure 4) to the location where the travel date is written (see Figure 5). This allows the user to easily recognize where their selected candidate is located in the document data 4 by looking at the layout image 80.
[0056] While Figures 4 and 5 show examples of displaying and selecting candidate transaction dates, it is also possible to display and select candidates for other display items. For example, if a user wants to change the default "AAA Corporation" displayed in the customer name display field 72, the user can perform a list display operation on the customer name display field 72 to display a list of candidates for that field directly below it, and then change the content displayed in the customer name display field 72 by selecting the desired candidate from the displayed list.
[0057] A confirmation button 90 is displayed below the result image 70. When the user selects and presses the confirmation button 90, the content displayed in the display area of the result image 70 is confirmed as the content requested by the user. The confirmed content is sent back from the user terminal 3 to the document service system 2 as feedback information indicating the content requested by the user. The feedback information includes not only the user selection candidates, but also semantic information indicating the meaning of the user selection candidates (for example, if the user selection candidate is a date, information indicating the action performed on that date), and user transaction information.
[0058] <Learning Phase> Returning to Figure 2, we will explain the processing of the learning phase by the document service system 2 (steps S40-S60).
[0059] After the AI-OCR processing result is displayed on the user terminal 3 screen in step S30, the learning device 40 determines whether or not feedback information has been returned from the user terminal 3 (step S40). If no feedback information has been returned (NO in step S40), the learning device 40 repeats the process in step S40 and waits for feedback information to be returned.
[0060] If feedback information is received from the user terminal 3 (YES in step S40), the learning device 40 saves the feedback information to the database 30 (step S50).
[0061] Figure 6 schematically shows an example of the data structure stored in database 30. Database 30 stores multiple pieces of feedback information. Each piece of feedback information is associated with user transaction information and AI-OCR processing results. The AI-OCR processing results include the document type and multiple candidates for each display item.
[0062] Furthermore, the feedback information includes not only user selection options but also semantic information that indicates the meaning of each user selection option.
[0063] The learning device 40 uses the feedback information stored in the database 30 as training data to learn the models (first model and second model) stored in the model storage unit 10 (step S60).
[0064] In this process, the learning device 40 learns a model based on user transaction information (user ID, etc.) and AI-OCR processing results (document type, etc.). As a result, the model stored in the model storage unit 10 is personalized for each user (for each user's transaction content) and adapted to the specific use cases for each user's document type. Consequently, in subsequent AI-OCR processing, the user can be presented with multiple candidates and their priorities that are suitable for their needs.
[0065] Furthermore, the learning device 40 learns a model using the semantic information contained in the feedback information. Specifically, the learning device 40 grasps the user's needs from the semantic information contained in the feedback information and learns a model to satisfy the user's needs. As a result, in subsequent AI-OCR processing, the user can be presented with multiple candidates that are more suitable to the user's needs and their priorities.
[0066] As described above, in the document service system 2 according to this embodiment, the utilization phase processing (steps S10 to S30), in which the model is used to present multiple candidates for each item to the user, and the learning phase processing (steps S40 to S60), in which the model is trained using feedback information containing information on the candidate selected by the user from the multiple candidates as training data, are repeated each time document data 4 is received. As a result, the model stored in the model storage unit 10 is personalized for each user, and content that reflects each user's needs is preferentially displayed on the screen of the user terminal 3. As a result, the user can be presented with AI-OCR processing results that appropriately reflect the user's needs.
[0067] In particular, in the document service system 2 according to this embodiment, during the AI-OCR processing in the utilization phase, instead of outputting one candidate for each item, multiple candidates and their priorities are output. This avoids providing the user with information that does not match the user's intent, compared to the case where only one candidate is output for each item. As a result, it becomes easier to provide the user with the information they intended.
[0068] Furthermore, in the document service system 2 according to this embodiment, feedback information is stored in association with user transaction information, and the model is trained based on the user transaction information. This allows the process of training the model based on the user's transaction information to be performed accurately and quickly. The results of the AI-OCR processing using the model trained for each user are then presented to the user. Therefore, information that appropriately reflects the needs of each user can be presented to each user.
[0069] Thus, the document service system 2 according to this embodiment can present each user with AI-OCR processing results 5B that are suitable for each user's needs, even with the same document data 4. For example, even with the same receipt, the user's needs may differ in whether they want the company name or the official name extracted as the "business partner" item. Similarly, the user's needs may differ in whether they want the amount excluding tax or the amount including tax extracted as the "amount" item. The document service system 2 according to this embodiment can appropriately respond to such differences in user needs.
[0070] Figure 7 schematically shows the AI-OCR processing results 5A to 5C when the document service system 2 receives the same document data 4 from companies A, B, and C, respectively.
[0071] Even if the document data 4 to be processed is the same, the information to be extracted as content for each item will differ depending on the user's needs. As shown in Figure 7, the document service system 2 according to this embodiment can, when it receives the same document data 4 from companies A, B, and C, present company A with an AI-OCR processing result 5A suitable for company A, company B with an AI-OCR processing result 5B suitable for company B, and company C with an AI-OCR processing result 5C suitable for company C.
[0072] Furthermore, in the document service system 2 according to this embodiment, the model is trained based on the document type. This allows the system to learn the detailed use cases for each user's document type. The results of AI-OCR processing using the model trained based on the document type are then presented to the user. Therefore, even if the user's needs differ for each document type, the system can present the user with AI-OCR processing results that appropriately reflect each user's needs.
[0073] Figure 8 schematically shows the AI-OCR processing result 6A when the document type is "invoice" and the AI-OCR processing result 6B when the document type is "receipt". Even if the user's needs differ between invoices and receipts, as shown in Figure 8, if document data 4 is an invoice, the user can be presented with AI-OCR processing result 6A suitable for the invoice use case, and if document data 4 is a receipt, the user can be presented with AI-OCR processing result 6b suitable for the receipt use case.
[0074] Furthermore, in the document service system 2 according to this embodiment, the feedback information includes not only user selection candidates but also semantic information indicating the meaning of each user selection candidate. This allows the learning device 40 to learn the model while taking into account the user's intention in selecting each user selection candidate. Therefore, the process of learning the model for each user can be performed more accurately and more quickly.
[0075] Furthermore, in the document service system 2 according to this embodiment, the models stored in the model storage unit 10 are divided into a first model (candidate output model) that outputs multiple candidates for each item, and a second model (priority output model) that outputs the priority of the multiple candidates for each item. This allows the first model to be a general-purpose model that can be shared by many users, while the second model can be a model customized for each user. This makes it possible to clearly stratify the learning objectives of each model. As a result, it is possible to suppress the complexity of the learning content of the entire model compared to the case where there is only one model stored in the model storage unit 10.
[0076] <Example 1> Figure 1 above shows an example where the models stored in the model storage unit 10 are divided into two models (first model and second model). However, the number of models stored in the model storage unit 10 is not necessarily limited to two; it may be three or more, or even just one.
[0077] <Modification 2> In Figure 1 above, an example is shown in which the learning device 40 is located outside the processing device 20. However, the learning device 40 may also be located inside the processing device 20 (i.e., the processing device 20 also functions as the learning device 40).
[0078] The embodiments disclosed herein should be considered in all respects to be illustrative and not restrictive. The scope of this disclosure is indicated by the claims rather than the foregoing description, and all modifications within the meaning and scope equivalent to the claims are intended. [Explanation of Symbols]
[0079] 1 System, 2 Document service system, 3 User terminal, 3a, 3b, 3c User terminal image, 4 Document data, 5A, 5B, 5C, 6A, 6B AI-OCR processing result, 10 Model storage unit, 12 First storage unit, 14 Second storage unit, 20 Processing unit, 30 Database, 40 Learning device, 70 Result image, 71-75 Display area, 75a Candidate list, 75b Calendar, 76 Cursor, 80 Layout image, 82, 84, 85 Frame, 90 Confirm button.
Claims
1. A document service system configured to communicate with a user's terminal, A model storage unit stores a model configured to automatically read characters in document data obtained from the user's terminal, extract multiple candidates that can be selected as content for a predetermined item from the read characters, and output them. The apparatus comprises a processing device that performs processing using the aforementioned model, The aforementioned processing apparatus is The multiple candidates are obtained by inputting the document data acquired from the user's terminal into the model. The acquired plurality of candidates are displayed on the user's terminal screen in a manner that allows the user to select one. The processing device is a document service system that displays information representing the meaning of the multiple candidates, in addition to the multiple candidates, on the screen of the user's terminal.
2. A document service system configured to communicate with a user's terminal, A model storage unit stores a model configured to automatically read characters in document data obtained from the user's terminal, extract multiple candidates that can be selected as content for a predetermined item from the read characters, and output them. The apparatus comprises a processing device that performs processing using the aforementioned model, The aforementioned processing apparatus is The multiple candidates are obtained by inputting the document data acquired from the user's terminal into the model. The acquired plurality of candidates are displayed on the user's terminal screen in a manner that allows the user to select one. A database which acquires and stores the candidate selected by the user from among the aforementioned multiple candidates as feedback information from the user's terminal, The system further comprises a learning device that learns the model using the feedback information stored in the database as training data, The processing device obtains the multiple candidates by inputting the document data into the model learned by the learning device and displays them on the screen of the user's terminal. The database stores the feedback information in association with user transaction information for identifying the user's transaction trends. The learning device learns the model based on the user transaction information, The aforementioned processing apparatus is A document service system that inputs document data and user transaction information obtained from the user's terminal into the model, thereby obtaining a plurality of candidates suitable for the user's operations and displaying them on the screen of the user's terminal.
3. A document service system configured to communicate with a user's terminal, A model storage unit stores a model configured to automatically read characters in document data obtained from the user's terminal, extract multiple candidates that can be selected as content for a predetermined item from the read characters, and output them. The apparatus comprises a processing device that performs processing using the aforementioned model, The aforementioned processing apparatus is The multiple candidates are obtained by inputting the document data acquired from the user's terminal into the model. The acquired plurality of candidates are displayed on the user's terminal screen in a manner that allows the user to select one. A database which acquires and stores the candidate selected by the user from among the aforementioned multiple candidates as feedback information from the user's terminal, The system further comprises a learning device that learns the model using the feedback information stored in the database as training data, The processing device obtains the multiple candidates by inputting the document data into the model learned by the learning device and displays them on the screen of the user's terminal. The aforementioned database stores document types indicating the type of document data, in addition to the feedback information. The learning device learns the model based on the document type, The aforementioned processing apparatus is A document service system that inputs document data acquired from the user's terminal into the model, thereby acquiring a plurality of candidates suitable for the type of document data and displaying them on the screen of the user's terminal.
4. A document service system configured to communicate with a user's terminal, A model storage unit stores a model configured to automatically read characters in document data obtained from the user's terminal, extract multiple candidates that can be selected as content for a predetermined item from the read characters, and output them. The apparatus comprises a processing device that performs processing using the aforementioned model, The aforementioned processing apparatus is The multiple candidates are obtained by inputting the document data acquired from the user's terminal into the model. The acquired plurality of candidates are displayed on the user's terminal screen in a manner that allows the user to select one. A database which acquires and stores the candidate selected by the user from among the aforementioned multiple candidates as feedback information from the user's terminal, The system further comprises a learning device that learns the model using the feedback information stored in the database as training data, The processing device obtains the multiple candidates by inputting the document data into the model learned by the learning device and displays them on the screen of the user's terminal. The feedback information includes, in addition to the candidates selected by the user, semantic information representing the meaning of the candidates selected by the user. The learning device is a document service system that learns the model using the semantic information contained in the feedback information.
5. The aforementioned processing apparatus is The aforementioned multiple candidates and the layout image of the document data are displayed side by side on the screen of the user's terminal. A document service system according to any one of claims 1 to 4, wherein the location of the candidate selected by the user from among the plurality of candidates is highlighted on the layout image.
6. The model is configured to extract the multiple candidates when the document data is input and to output the priority of the multiple candidates. The aforementioned processing apparatus is By inputting the document data obtained from the user's terminal into the model, the plurality of candidates and the priority of the plurality of candidates are obtained. A document service system according to any one of claims 1 to 4, wherein the acquired priority is displayed on the screen of the user's terminal in a manner that the user can recognize.
7. The document service system according to claim 6, wherein the processing device displays a list of the plurality of candidates arranged in order of priority on the screen of the user's terminal.
8. The aforementioned processing apparatus is The candidate with the highest priority among the aforementioned multiple candidates is displayed by default on the user's terminal screen. The document service system according to claim 7, which displays a list of the multiple candidates arranged in descending order of priority on the screen of the user's terminal in response to the user's operation.
9. A method performed by a document service system configured to communicate with a user's terminal, The document service system stores a model configured to automatically read characters in document data obtained from the user's terminal, extract multiple candidates that can be selected as content for a predetermined item from the read characters, and output them. The method performed by the aforementioned document service system is: The steps include obtaining the multiple candidates by inputting document data obtained from the user's terminal into the model, The process includes the step of displaying the acquired plurality of candidates on the screen of the user's terminal in a manner that allows the user to select them, A method performed by a document service system, wherein the step of displaying the plurality of candidates on the screen includes the step of displaying information representing the meaning of the plurality of candidates on the screen of the user's terminal, in addition to the plurality of candidates.
10. A method performed by a document service system configured to communicate with a user's terminal, The document service system stores a model configured to automatically read characters in document data obtained from the user's terminal, extract multiple candidates that can be selected as content for a predetermined item from the read characters, and output them. The method performed by the aforementioned document service system is: The steps include obtaining the multiple candidates by inputting document data obtained from the user's terminal into the model, The steps include displaying the acquired plurality of candidates on the screen of the user's terminal in a manner that allows the user to select one, The steps include: acquiring and storing the candidate selected by the user from among the aforementioned multiple candidates as feedback information from the user's terminal; The process includes the step of training the model using the accumulated feedback information as training data, The step of displaying the plurality of candidates on the screen includes the step of obtaining the plurality of candidates by inputting the document data into the trained model and displaying them on the screen of the user's terminal, The step of accumulating the candidates selected by the user as feedback information includes the step of accumulating the feedback information in association with user transaction information for identifying the user's transaction trends, The step of training the model includes the step of training the model based on the user transaction information, A method performed by a document service system, wherein the step of displaying the aforementioned plurality of candidates on the screen includes the step of inputting document data and user transaction information obtained from the user's terminal into the model to obtain the plurality of candidates suitable for the user's operation and display them on the screen of the user's terminal.
11. A method performed by a document service system configured to communicate with a user's terminal, The document service system stores a model configured to automatically read characters in document data obtained from the user's terminal, extract multiple candidates that can be selected as content for a predetermined item from the read characters, and output them. The method performed by the aforementioned document service system is: The steps include obtaining the multiple candidates by inputting document data obtained from the user's terminal into the model, The steps include displaying the acquired plurality of candidates on the screen of the user's terminal in a manner that allows the user to select one, The steps include: acquiring and storing the candidate selected by the user from among the aforementioned multiple candidates as feedback information from the user's terminal; The process includes the step of training the model using the accumulated feedback information as training data, The step of displaying the plurality of candidates on the screen includes the step of obtaining the plurality of candidates by inputting the document data into the trained model and displaying them on the screen of the user's terminal, The step of storing the candidate selected by the user as feedback information includes, in addition to the feedback information, a step of storing a document type indicating the type of document data, The step of training the aforementioned model includes the step of training the aforementioned model based on the aforementioned document type, A method performed by a document service system, wherein the step of displaying the plurality of candidates on the screen includes the step of inputting document data obtained from the user's terminal into the model to obtain the plurality of candidates suitable for the type of document data and displaying them on the screen of the user's terminal.
12. A method performed by a document service system configured to communicate with a user's terminal, The document service system stores a model configured to automatically read characters in document data obtained from the user's terminal, extract multiple candidates that can be selected as content for a predetermined item from the read characters, and output them. The method performed by the aforementioned document service system is: The steps include obtaining the multiple candidates by inputting document data obtained from the user's terminal into the model, The steps include displaying the acquired plurality of candidates on the screen of the user's terminal in a manner that allows the user to select one, The steps include: acquiring and storing the candidate selected by the user from among the aforementioned multiple candidates as feedback information from the user's terminal; The process includes the step of training the model using the accumulated feedback information as training data, The step of displaying the plurality of candidates on the screen includes the step of obtaining the plurality of candidates by inputting the document data into the trained model and displaying them on the screen of the user's terminal, The feedback information includes, in addition to the candidates selected by the user, semantic information representing the meaning of the candidates selected by the user. A method performed by a document service system, wherein the step of learning the model includes the step of learning the model using the semantic information contained in the feedback information.
13. A program that causes a computer to perform the same actions as a document service system configured to communicate with a user's terminal, The document service system stores a model configured to automatically read characters in document data obtained from the user's terminal, extract multiple candidates that can be selected as content for a predetermined item from the read characters, and output them. The aforementioned program is installed on the computer. The steps include obtaining the multiple candidates by inputting document data obtained from the user's terminal into the model, The process involves the user performing the steps of displaying the acquired plurality of candidates on the user's terminal screen in a manner that allows the user to select one, A program to be executed by a computer, wherein the step of displaying the plurality of candidates on the screen includes the step of displaying information representing the meaning of the plurality of candidates, in addition to the plurality of candidates, on the screen of the user's terminal.
14. A program that causes a computer to perform a method performed by a document service system configured to communicate with a user's terminal, The document service system stores a model configured to automatically read characters in document data obtained from the user's terminal, extract multiple candidates that can be selected as content for a predetermined item from the read characters, and output them. The aforementioned program is installed on the computer. The steps include obtaining the multiple candidates by inputting document data obtained from the user's terminal into the model, The steps include displaying the acquired plurality of candidates on the screen of the user's terminal in a manner that allows the user to select one, The steps include: acquiring and storing the candidate selected by the user from among the aforementioned multiple candidates as feedback information from the user's terminal; The accumulated feedback information is used as training data to train the model, and the following steps are performed: The step of displaying the plurality of candidates on the screen includes the step of obtaining the plurality of candidates by inputting the document data into the trained model and displaying them on the screen of the user's terminal, The step of accumulating the candidates selected by the user as feedback information includes the step of accumulating the feedback information in association with user transaction information for identifying the user's transaction trends, The step of training the model includes the step of training the model based on the user transaction information, The step of displaying the aforementioned multiple candidates on the screen includes a program to be executed by a computer, which includes inputting document data and user transaction information obtained from the user's terminal into the model to obtain the aforementioned multiple candidates suitable for the user's operation and display them on the screen of the user's terminal.
15. A program that causes a computer to perform a method performed by a document service system configured to communicate with a user's terminal, The document service system stores a model configured to automatically read characters in document data obtained from the user's terminal, extract multiple candidates that can be selected as content for a predetermined item from the read characters, and output them. The aforementioned program is installed on the computer. The steps include obtaining the multiple candidates by inputting document data obtained from the user's terminal into the model, The steps include displaying the acquired plurality of candidates on the screen of the user's terminal in a manner that allows the user to select one, The steps include: acquiring and storing the candidate selected by the user from among the aforementioned multiple candidates as feedback information from the user's terminal; The accumulated feedback information is used as training data to train the model, and the following steps are performed: The step of displaying the plurality of candidates on the screen includes the step of obtaining the plurality of candidates by inputting the document data into the trained model and displaying them on the screen of the user's terminal, The step of storing the candidate selected by the user as feedback information includes, in addition to the feedback information, a step of storing a document type indicating the type of document data, The step of training the aforementioned model includes the step of training the aforementioned model based on the aforementioned document type, The step of displaying the plurality of candidates on the screen is a program to be executed by a computer, which includes the step of inputting document data obtained from the user's terminal into the model to obtain the plurality of candidates suitable for the type of document data and display them on the screen of the user's terminal.
16. A program that causes a computer to perform a method performed by a document service system configured to communicate with a user's terminal, The document service system stores a model configured to automatically read characters in document data obtained from the user's terminal, extract multiple candidates that can be selected as content for a predetermined item from the read characters, and output them. The aforementioned program is installed on the computer. The steps include obtaining the multiple candidates by inputting document data obtained from the user's terminal into the model, The steps include displaying the acquired plurality of candidates on the screen of the user's terminal in a manner that allows the user to select one, The steps include: acquiring and storing the candidate selected by the user from among the aforementioned multiple candidates as feedback information from the user's terminal; The accumulated feedback information is used as training data to train the model, and the following steps are performed: The step of displaying the plurality of candidates on the screen includes the step of obtaining the plurality of candidates by inputting the document data into the trained model and displaying them on the screen of the user's terminal, The feedback information includes, in addition to the candidates selected by the user, semantic information representing the meaning of the candidates selected by the user. A program to be executed by a computer, the step of learning the model includes a step of learning the model using the semantic information contained in the feedback information.
Citation Information
Patent Citations
Image processor, method for processing image, and program
JP2021012741A
Information processing device, program, and information processing method
JP2021144617A
Information processing system, and program
JP2024033878A
Document information processing device, document information structuring processing method, and document information structuring processing program
JP2020013281A
Document image processing system, document image processing method, and document image processing program
JP7561378B2